<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2025.1527980</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>MedAlmighty: enhancing disease diagnosis with large vision model distillation</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Ren</surname> <given-names>Yajing</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2894433/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Gu</surname> <given-names>Zheng</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/3153221/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Liu</surname> <given-names>Wen</given-names></name>
<role content-type="https://credit.niso.org/contributor-roles/funding-acquisition/"/>
<role content-type="https://credit.niso.org/contributor-roles/project-administration/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff><institution>Artificial Intelligence and Smart Mine Engineering Technology Center, Xinjiang Institute of Engineering</institution>, <addr-line>Urumqi</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Liping Zhang, Harvard Medical School, United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Moiz Khan Sherwani, University of Copenhagen, Denmark</p>
<p>Hugo Vega-Huerta, National University of San Marcos, Peru</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Yajing Ren <email>1441518764&#x00040;qq.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>12</day>
<month>08</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>8</volume>
<elocation-id>1527980</elocation-id>
<history>
<date date-type="received">
<day>14</day>
<month>11</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>18</day>
<month>07</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2025 Ren, Gu and Liu.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Ren, Gu and Liu</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<sec>
<title>Introduction</title>
<p>Accurate disease diagnosis is critical in the medical field, yet it remains a challenging task due to the limited, heterogeneous, and complex nature of medical data. These challenges are particularly pronounced in multimodal tasks requiring the integration of diverse data sources. While lightweight models offer computational efficiency, they often lack the comprehensive understanding necessary for reliable clinical predictions. Conversely, large vision models, trained on extensive general-domain datasets, provide strong generalization but fall short in specialized medical applications due to domain mismatch and limited medical data availability.</p></sec>
<sec>
<title>Methods</title>
<p>To bridge the gap between general and specialized performance, we propose MedAlmighty, a knowledge distillation-based framework that synergizes the strengths of both large and small models. In this approach, we utilize DINOv2&#x02014;a pre-trained large vision model&#x02014;as a frozen teacher, and a lightweight convolutional neural network (CNN) as the trainable student. The student model is trained using both hard labels from the ground truth and soft targets generated by the teacher model. We adopt a hybrid loss function that combines cross-entropy loss (for classification accuracy) and Kullback-Leibler divergence (for distillation), enabling the student model to capture rich semantic features while remaining efficient and domain-aware.</p></sec>
<sec>
<title>Results</title>
<p>Experimental evaluations reveal that MedAlmighty significantly improves disease diagnosis performance across datasets characterized by sparse and diverse medical data. The proposed model outperforms baselines by effectively integrating the generalizable representations of large models with the specialized knowledge from smaller models. The results confirm improved robustness and accuracy in complex diagnostic scenarios.</p></sec>
<sec>
<title>Discussion</title>
<p>The MedAlmighty framework demonstrates that incorporating general-domain representations via frozen large vision models&#x02014;when guided by task-specific distillation strategies&#x02014;can enhance the performance of lightweight medical models. This approach offers a promising solution to data scarcity and domain gap issues in medical imaging. Future work may explore extending this distillation strategy to other medical modalities and incorporating multimodal alignment for even richer representation learning.</p></sec></abstract>
<kwd-group>
<kwd>disease diagnosis</kwd>
<kwd>large vision model</kwd>
<kwd>knowledge distillation</kwd>
<kwd>model capacity</kwd>
<kwd>domain generalization</kwd>
</kwd-group>
<counts>
<fig-count count="6"/>
<table-count count="4"/>
<equation-count count="9"/>
<ref-count count="58"/>
<page-count count="13"/>
<word-count count="9091"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Pattern Recognition</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1 Introduction</title>
<p>Artificial intelligence (AI) has driven transformation progress in imaging and vision, empowering applications across fields such as autonomous driving, robotics, and healthcare. In the medical domain, computer-aided diagnosis (CAD) has become pivotal for enhancing disease prognosis, enabling early detection, guiding risk stratification, and supporting personalized treatment planning (<xref ref-type="bibr" rid="B7">Edupuganti et al., 2024</xref>). However, despite remarkable advances, a persistent bottleneck remains: disparities in medical resource distribution and limited access to high-quality annotated data often hinder the deployment of robust diagnostic systems, especially in under-resourced regions. To address this, medical professionals increasingly rely on diverse imaging modalities&#x02014;including X-rays, MRI, CT scans, PET, SPECT, and ultrasound&#x02014;which provide critical insights into patient conditions (Zhu Y. et al., <xref ref-type="bibr" rid="B58">2024</xref>; <xref ref-type="bibr" rid="B2">Arumugam et al., 2024</xref>; <xref ref-type="bibr" rid="B28">Miller et al., 2024</xref>; Wang X. et al., <xref ref-type="bibr" rid="B43">2024</xref>).</p>
<p>Within the broader landscape of multimodal vision AI (<xref ref-type="bibr" rid="B12">Huo et al., 2025</xref>; <xref ref-type="bibr" rid="B46">Yan et al., 2024</xref>; <xref ref-type="bibr" rid="B22">Lyu et al., 2024</xref>), large-scale models have demonstrated unparalleled generalization capabilities, learning powerful feature representations from vast non-medical datasets (<xref ref-type="bibr" rid="B55">Zheng et al., 2025</xref>). Yet, their application to medical imaging remains underexplored, largely due to two key challenges: the scarcity of diverse, high-quality medical data and the complex multimodal nature of diagnostic tasks (<xref ref-type="bibr" rid="B54">Zheng et al., 2022</xref>; <xref ref-type="bibr" rid="B5">Chen et al., 2022</xref>). While successful classification models often build upon enhanced CNN variants, lightweight models face inherent limitations in capturing sufficient knowledge because of their restricted capacity. Consequently, there has been a shift toward large vision models with deeper architectures and richer parameters, which promise more expressive and generalized features (Zhang W. et al., <xref ref-type="bibr" rid="B51">2024</xref>; Zhu C. et al., <xref ref-type="bibr" rid="B57">2024</xref>; <xref ref-type="bibr" rid="B18">Liao et al., 2025</xref>; <xref ref-type="bibr" rid="B53">Zhao et al., 2025</xref>; <xref ref-type="bibr" rid="B56">Zhong et al., 2025</xref>). However, applying these models directly to medical imaging introduces domain gaps; they are typically pre-trained on natural images, which lack the biological and pathological features essential for accurate medical interpretation.</p>
<p>In the realm of medical image classification, researchers have sought to address data scarcity through three main strategies: image generation and enhancement, transfer learning, and knowledge distillation. For example, GAN-based methods (<xref ref-type="bibr" rid="B8">Feng et al., 2024</xref>; <xref ref-type="bibr" rid="B26">MeenaPrakash et al., 2025</xref>) synthesize realistic medical images to augment small datasets, improving downstream performance. Pre-training approaches such as RadImageNet (<xref ref-type="bibr" rid="B27">Mei et al., 2022</xref>) leverage large radiology datasets to enhance generalizability, while transfer learning techniques (Wang W. et al., <xref ref-type="bibr" rid="B42">2024</xref>) adapt models to specific diagnostic tasks. Knowledge distillation methods (<xref ref-type="bibr" rid="B39">Song et al., 2025</xref>) further compress large models into lightweight versions suitable for resource-constrained environments. Despite their individual successes, many of these methods require custom model designs tailored to specific datasets and tasks, incurring significant development costs and limiting scalability.</p>
<p>In the area of transfer learning, Wang W. et al. (<xref ref-type="bibr" rid="B42">2024</xref>) propose a DenseNet-based breast cancer classification model that incorporates attention mechanisms and multi-level transfer learning. This model achieves an accuracy of over 84.0%, demonstrating improved efficiency for pathological image analysis. In knowledge distillation, (<xref ref-type="bibr" rid="B39">Song et al. 2025</xref>) present a lightweight Shift-MLP-based student model with multi-teacher distillation. Additionally, they introduce a two-stage diagnostic framework that fuses multimodal data and transfers privileged knowledge from teacher to student models, outperforming existing methods in glioma grading and skin lesion classification.</p>
<p>While the aforementioned approaches effectively address the limited sample sizes in medical image datasets, they often rely on custom-designed models tailored to specific datasets and tasks, resulting in high development and deployment costs. In contrast, our goal is to explore <bold><italic>a more generalizable solution&#x02014;one that performs consistently across diverse</italic></bold> <bold><italic>medical modalities despite data scarcity</italic></bold>. Large vision models, with their substantial parameter capacity and complex architectures, have shown strong generalization and robust feature representation in computer vision, making them promising candidates for this purpose.</p>
<p>Accordingly, we adopt DINOv2 as our base model. However, experiments reveal inconsistent performance across different medical tasks. This may stem from a fundamental domain gap: DINOv2 is pre-trained on natural images, which lack the concept of biological tissue, whereas medical image analysis often relies on distinguishing between normal and abnormal tissues. For example, pneumonia diagnosis in X-rays depends on detecting diffuse pathological changes in lung tissue, while conditions like cardiac tumors or edema are more associated with localized boundary changes&#x02014;features that align better with edge-sensitive representations learned from natural images. These observations suggest that large vision models alone may not fully capture the nuances of medical data. Therefore, we propose combining the strengths of large vision models with the domain-specific expertise of smaller models to achieve more reliable and adaptable performance.</p>
<p>To <bold><italic>bridge the gap between generalizable representation learning and domain-specific</italic></bold> <bold><italic>efficiency</italic></bold>, we introduce <bold><italic>MedAlmighty</italic></bold>, a distillation framework that transfers knowledge from a large vision model (DINOv2) to a compact student model (ResNet), which demonstrated strong performance among baselines on the MedMNISTv2 dataset. As illustrated in <xref ref-type="fig" rid="F1">Figure 1</xref>, the distillation process enables the student model to inherit robust and general features from the teacher while maintaining high classification accuracy on limited medical data. We evaluate MedAlmighty across all 12 modalities in the MedMNISTv2 dataset and compare it against existing lightweight enhancement approaches to validate its effectiveness.</p>
<fig position="float" id="F1">
<label>Figure 1</label>
<caption><p>Comparison of generalization and training efficiency between CNNs and DINOv2. This figure provides a comprehensive comparison of CNNs and DINOv2 in terms of generalization and training efficiency. <bold>(a)</bold> Generalization Performance: CNNs struggle with robustness and accuracy on unseen data, while DINOv2 exhibits stronger generalization across diverse tasks due to self-supervised learning. <bold>(b)</bold> Training Efficiency: DINOv2 requires significantly more computational resources and training time, limiting its practicality. <bold>(c)</bold> Synergy Potential: The figure also underscores the advantages of combining CNNs&#x00027; efficiency with DINOv2&#x00027;s generalization, motivating the integration of both in a unified framework.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1527980-g0001.tif">
<alt-text>Diagram showing three models: (a) a lightweight CNN model with about twenty million parameters, (b) a heavyweight DINOV2 model with about two hundred million parameters, and (c) a hybrid model combining CNNs and DINOV2. The hybrid model uses a student-teacher framework where CNNs (student) learn from DINOV2 (teacher) using cross-entropy and distillation loss. The hybrid model aims to be lightweight yet knowledge-rich.</alt-text>
</graphic>
</fig>
<p>Our key contributions are as follows:</p>
<list list-type="bullet">
<list-item><p>We present MedAlmighty, a knowledge distillation framework that integrates the robust feature representation of a large vision model (DINOv2) with the lightweight efficiency of a small model (ResNet), enabling effective classification across 12 diverse medical imaging modalities.</p></list-item>
<list-item><p>We investigate the use of a frozen DINOv2 backbone with a trainable linear classifier and observe its limitations in domain-specific tasks, motivating the need for knowledge transfer to a smaller, specialized model.</p></list-item>
<list-item><p>MedAlmighty leverages the complementary strengths of both large and small models, offering a generalizable and scalable solution for medical image classification under data-scarce and multi-modality conditions.</p></list-item>
<list-item><p>We conduct extensive experiments on the MedMNISTv2 dataset and benchmark MedAlmighty against multiple lightweight model enhancement baselines, demonstrating its superior performance and robustness.</p></list-item>
</list></sec>
<sec id="s2">
<title>2 Related work</title>
<sec>
<title>2.1 Supervised learning in medical image analysis</title>
<p>Supervised learning has been widely adopted in medical image analysis, achieving strong performance across various tasks. Recent efforts have focused on enhancing the accuracy and robustness of supervised models, particularly through deep learning approaches. Convolutional neural networks (CNNs) have proven effective in extracting discriminative features from medical images (<xref ref-type="bibr" rid="B21">Liu et al., 2024</xref>; <xref ref-type="bibr" rid="B49">Yang et al., 2025</xref>; <xref ref-type="bibr" rid="B29">Mishra et al., 2025</xref>; <xref ref-type="bibr" rid="B25">Maree et al., 2024</xref>; <xref ref-type="bibr" rid="B32">Oyelade et al., 2024</xref>). Attention mechanisms have also been incorporated into CNNs to improve classification by emphasizing salient image regions (<xref ref-type="bibr" rid="B40">Takahashi et al., 2024</xref>). Supervised learning offers key advantages, including high accuracy with clear labels, the ability to learn complex feature representations from large annotated datasets, and adaptability to diverse clinical tasks such as disease diagnosis and lesion detection. To further boost performance, transfer learning has been widely explored, enabling models to generalize better with limited labeled data. For instance, CNN-based transfer learning approaches have been used to classify benign and malignant breast masses in X-ray images, with enhanced results achieved through model selection, ensemble averaging, and feature concatenation (<xref ref-type="bibr" rid="B9">Han et al., 2024</xref>).</p>
<p>More recently, Vision Transformers (ViTs) (<xref ref-type="bibr" rid="B6">Dosovitskiy et al., 2020</xref>), originally developed for natural image recognition, have gained attention in medical imaging due to their ability to model long-range dependencies via self-attention. ViT-based models have shown promise in handling complex and heterogeneous data. For example, MedViT (<xref ref-type="bibr" rid="B23">Manzari et al., 2023</xref>) integrates local feature extraction with global context modeling and efficient attention mechanisms, demonstrating strong performance on datasets like MedMNIST-2D (<xref ref-type="bibr" rid="B47">Yang et al., 2020</xref>). Despite these advances, generalizability remains a significant challenge (<xref ref-type="bibr" rid="B19">Liu et al., 2025</xref>; <xref ref-type="bibr" rid="B33">Pacal et al., 2025</xref>). Many supervised methods are designed for single-modality inputs and struggle with the multimodal nature of real-world medical data. Moreover, the need for large volumes of expert-labeled data poses practical constraints, as annotation is time-consuming, costly, and may introduce subjectivity. Addressing these limitations is crucial for developing robust, scalable models suited for clinical deployment.</p></sec>
<sec>
<title>2.2 Self-supervised learning in medical image analysis</title>
<p>Self-supervised learning has emerged as a powerful approach in medical image analysis, particularly for addressing the scarcity and cost of annotated data. By leveraging the inherent structure of medical images to generate supervisory signals, it enables the learning of robust feature representations without extensive manual labeling. Many self-supervised methods draw inspiration from Masked Autoencoders (MAE) (<xref ref-type="bibr" rid="B10">He et al., 2021</xref>). For instance, ChA-MAEViT (<xref ref-type="bibr" rid="B36">Pham et al., 2025</xref>) enhances medical image classification by explicitly modeling cross-channel dependencies via dynamic masking and memory tokens, achieving up to 21.5% higher accuracy than existing MCI-ViTs on microscopy datasets. Similarly,MSMAE (<xref ref-type="bibr" rid="B24">Mao et al., 2025</xref>) introduces a supervised attention-driven masking strategy (SAM) to precisely localize and learn lesion-related regions in medical images, achieving SOTA classification accuracy 68.41%&#x02013;99.60% while reducing FLOPs by 74.08% and inference time by 11.2% compared to MAE.</p>
<p>Autoencoder-based architectures have also been explored, with design improvements addressing class imbalance and enhancing diagnostic accuracy (<xref ref-type="bibr" rid="B44">Xing et al., 2023</xref>; <xref ref-type="bibr" rid="B1">Arafa et al., 2023</xref>). A lightweight self-supervised learning (SSL) model (<xref ref-type="bibr" rid="B14">Karagoz and Nalbantoglu, 2024</xref>) for mammography classification, combining a VAE-based pretext task for feature learning and a 3-layer CNN downstream network, achieving 0.94&#x02013;0.99 AUC with only 228 parameters and 204.95K FLOPs on INbreast/MIAS datasets. A notable example, ViT-AE&#x0002B;&#x0002B;, introduces novel loss functions for self-reconstruction and contrastive learning to enhance representation quality (<xref ref-type="bibr" rid="B37">Prabhakar et al., 2023</xref>). More broadly, contrastive learning has proven effective in distinguishing disease features from normal structures without labeled data (Zhang X. et al., <xref ref-type="bibr" rid="B52">2024</xref>; <xref ref-type="bibr" rid="B13">Jiao et al., 2024</xref>; <xref ref-type="bibr" rid="B15">Kumar and Marttinen, 2024</xref>; Li M. et al., <xref ref-type="bibr" rid="B16">2025</xref>; Li Q. et al., <xref ref-type="bibr" rid="B17">2025</xref>; <xref ref-type="bibr" rid="B45">Xu and Wong, 2025</xref>; <xref ref-type="bibr" rid="B20">Liu et al., 2022</xref>; <xref ref-type="bibr" rid="B30">Nguyen et al., 2024</xref>). Despite its advantages, self-supervised learning faces limitations. Without task-specific labels, models may miss subtle features essential for accurate diagnosis. Moreover, many approaches are tailored to specific diseases or modalities, often requiring significant computation and long training times. Their adaptability to heterogeneous datasets&#x02014;common in medical imaging due to varying acquisition protocols&#x02014;also remains limited.</p></sec>
<sec>
<title>2.3 Large vision models</title>
<p>Large vision models have advanced significantly in computer vision, drawing inspiration from architectures such as CLIP (<xref ref-type="bibr" rid="B38">Radford et al., 2021</xref>), MAE, ViT, and Beit (<xref ref-type="bibr" rid="B3">Bao et al., 2021</xref>; <xref ref-type="bibr" rid="B35">Peng et al., 2022</xref>). Most adopt the Transformer architecture&#x02014;particularly Vision Transformer (ViT)&#x02014;and are trained using diverse paradigms, including supervised learning [e.g., DeiT-III (<xref ref-type="bibr" rid="B41">Touvron et al., 2022</xref>)], text-image contrastive learning (e.g., OpenCLIP), and self-supervised learning [e.g., DINO (<xref ref-type="bibr" rid="B4">Caron et al., 2021</xref>), DINOv2 (<xref ref-type="bibr" rid="B31">Oquab et al., 2023</xref>)]. Numerous variants have been proposed to improve performance and efficiency. For example, Stream-ViT (<xref ref-type="bibr" rid="B34">Pan et al., 2025</xref>) dynamically integrates streamlined high-to-low resolution convolutions with self-attention, enhancing model capacity and efficiency . These models have set benchmarks across vision tasks and motivated the scaling of model size and data volume. Following their success on natural image datasets like ImageNet, interest is growing in adapting these models to medical imaging. Researchers are exploring whether representations learned during large-scale pre-training can benefit medical tasks such as classification, segmentation, and diagnosis&#x02014;especially where labeled data is limited (<xref ref-type="bibr" rid="B11">Huix et al., 2024</xref>). Early results suggest that zero-shot and few-shot capabilities of vision-language models (e.g., CLIP variants) show promise for lesion recognition and medical report generation, enabling more flexible and scalable diagnostic tools. Emerging frameworks support alignment between medical tasks and large models. Domain-specific pre-training on datasets such as MIMIC-CXR and MedMNIST has been explored to better adapt general models to medical data. Additionally, methods like lightweight adapters, prompt tuning, and hybrid architectures aim to integrate pre-trained models into clinical pipelines with reduced computational overhead. Despite these advances, challenges remain. Medical images differ markedly from natural images in modality, appearance, and annotation granularity. Moreover, clinical deployment demands greater interpretability, robustness, and domain adaptation. Addressing these issues is essential to fully realize the potential of large vision models in medical image analysis.</p></sec></sec>
<sec sec-type="methods" id="s3">
<title>3 Methods</title>
<sec>
<title>3.1 Problem definition</title>
<p>To address the challenges of medical image classification under data-scarce conditions, we propose a knowledge distillation framework that integrates the robust, generalizable feature extraction capabilities of large-scale vision models with the fine-tuned, domain-specific expertise of smaller models trained on medical datasets. Formally, we define the medical dataset as <inline-formula><mml:math id="M1"><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, where <italic>n</italic> is the number of labeled samples, <inline-formula><mml:math id="M2"><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represents a 2D medical image with <italic>C</italic><sub><italic>m</italic></sub> channels and spatial dimensions <italic>H</italic>&#x000D7;<italic>W</italic>, and <italic>y</italic><sub><italic>i</italic></sub> is the corresponding ground-truth label.</p>
<p>In this framework, the teacher model <italic>T</italic> is a large, pre-trained vision network (e.g., DINOv2) with a vast parameter space, capable of learning generalizable visual features. The student model <italic>S</italic> is a compact, lightweight network (e.g., CNN-based) specifically adapted to medical data. The distillation process aims to guide the student model <italic>S</italic> to accurately predict labels &#x00177;<sub><italic>i</italic></sub> by learning from both the ground-truth labels and the rich feature representations <inline-formula><mml:math id="M3"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> provided by the teacher model. By bridging the gap between general visual knowledge and domain-specific patterns, the proposed approach enables the student model to combine broad, transferable features with detailed, task-relevant refinements, ultimately improving efficiency and accuracy in medical image classification tasks.</p></sec>
<sec>
<title>3.2 Architecture</title>
<p>To leverage the advantages of large vision models while incorporating domain-specific knowledge from smaller medical datasets, we have designed a distillation framework that combines the powerful feature extraction capabilities of large models with the fine-grained domain-specific features learned by smaller models.</p>
<p>For an input sample <italic>X</italic><sub><italic>i</italic></sub>, it is simultaneously passed through both the teacher model <italic>T</italic><sub><italic>DINOv</italic>2</sub> and the student model <italic>S</italic><sub><italic>CNNs</italic></sub>, generating output features <inline-formula><mml:math id="M4"><mml:mrow><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="script">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>N</mml:mi><mml:mi>O</mml:mi><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M5"><mml:mrow><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="script">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>. These output features are then processed through a linear layer <italic>f</italic><sub><italic>Linear</italic></sub>(&#x000B7;) to obtain the output vectors:</p>
<disp-formula id="E1"><mml:math id="M6"><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msubsup><mml:mi mathvariant='script'>Z</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>N</mml:mi><mml:mi>O</mml:mi><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mn>11</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>N</mml:mi><mml:mi>O</mml:mi><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mn>12</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>N</mml:mi><mml:mi>O</mml:mi><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mn>...</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>N</mml:mi><mml:mi>O</mml:mi><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:msubsup><mml:mi mathvariant='script'>Z</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mn>11</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mn>12</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mn>...</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>The dimensions of <inline-formula><mml:math id="M8"><mml:mrow><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="script">Z</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>N</mml:mi><mml:mi>O</mml:mi><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M9"><mml:mrow><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="script">Z</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> are both <italic>m</italic>, where <italic>m</italic> represents the number of classes. We then apply the softmax function <italic>f</italic><sub><italic>softmax</italic></sub>(&#x000B7;) to convert the output vectors <inline-formula><mml:math id="M10"><mml:mrow><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="script">Z</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>N</mml:mi><mml:mi>O</mml:mi><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M11"><mml:mrow><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="script">Z</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> into probability distribution vectors <inline-formula><mml:math id="M12"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M13"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>, respectively. Each element <inline-formula><mml:math id="M14"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> represents the probability that the input belongs to class <italic>m</italic>:</p>
<disp-formula id="E2"><mml:math id="M15"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>/</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:mstyle><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>/</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>t</italic> is the distillation temperature, controlling the softness of the label distribution between the teacher and student networks. A higher temperature value <italic>t</italic> smooths the soft labels from the teacher model, providing more informative learning signals to the student model, which helps improve its generalization ability. By increasing <italic>t</italic>, the student model can more easily learn from the teacher&#x00027;s knowledge, leading to improved performance.</p>
<p>To enhance the generalization ability of the CNN student model in medical image classification, we minimize the Kullback-Leibler (KL) divergence between the soft labels <inline-formula><mml:math id="M16"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> of the student and teacher models. KL divergence measures the difference between two probability distributions and is defined as:</p>
<disp-formula id="E3"><mml:math id="M17"><mml:mrow><mml:msub><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mo>||</mml:mo><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">log</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>Here, <italic>t</italic> represents the distillation temperature. By minimizing this KL divergence, the student model optimizes its prediction while acquiring richer knowledge from the teacher model, thereby improving both performance and generalization. Thus, the distillation loss <italic>L</italic><sub>distill</sub> is introduced to minimize the KL divergence and enable the CNN model to better emulate the knowledge distribution of the teacher model <italic>T</italic><sub><italic>DINOv</italic>2</sub> for medical image classification. The distillation loss is represented as:</p>
<disp-formula id="E4"><mml:math id="M18"><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">distill</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>To effectively combine the strengths of both the teacher and student models, the same training dataset is used. The student model <italic>S</italic><sub><italic>CNNs</italic></sub> is trained on the true labels of the dataset <inline-formula><mml:math id="M19"><mml:mrow><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>. For each input sample <italic>X</italic><sub><italic>i</italic></sub>, the student model generates the output feature <inline-formula><mml:math id="M20"><mml:mrow><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="script">F</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>, which is then passed through a linear layer <italic>f</italic><sub><italic>Linear</italic></sub>(&#x000B7;) to produce the vector:</p>
<disp-formula id="E5"><mml:math id="M21"><mml:mrow><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="script">Z</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mn>21</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mn>22</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <italic>m</italic> is the number of classes. The predicted probability distribution of the student network, with the distillation temperature <italic>t</italic> &#x0003D; 1, is given by:</p>
<disp-formula id="E6"><mml:math id="M22"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:mstyle><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where each element <inline-formula><mml:math id="M23"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> represents the probability of the input belonging to class <italic>m</italic>. The student network minimizes the cross-entropy loss between the predicted probabilities and the true labels <italic>y</italic><sub><italic>i</italic></sub> in the medical dataset <inline-formula><mml:math id="M24"><mml:mrow><mml:msubsup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>. The cross-entropy loss is computed as:</p>
<disp-formula id="E7"><mml:math id="M25"><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">cross_entropy</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo class="qopname">log</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>y</italic><sub><italic>i</italic></sub> is the true label of the <italic>i</italic>-th sample.</p>
<p>To integrate knowledge from the teacher network and align with the student network&#x00027;s training objective, we define the total loss <italic>L</italic><sub>total</sub> as a weighted sum of the cross-entropy loss <italic>L</italic><sub>cross_entropy</sub> and the distillation loss <italic>L</italic><sub>distill</sub>. This total loss is expressed as:</p>
<disp-formula id="E8"><mml:math id="M26"><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">total</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">cross_entropy</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">distill</mml:mtext></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where &#x003B1; and &#x003B2; are hyperparameters that control the relative importance of the two loss functions, with &#x003B2; &#x0003D; 1&#x02212;&#x003B1;. Specifically, &#x003B1; controls the weight of the cross-entropy loss, while &#x003B2; determines the weight of the distillation loss. Adjusting &#x003B1; and &#x003B2; allows for balancing the incorporation of teacher network knowledge and the maintenance of classification accuracy.</p>
<p>When &#x003B1; is larger, the student model focuses more on matching the true labels, thereby improving classification accuracy. Conversely, increasing &#x003B2; emphasizes the distillation loss, enabling the student model to learn more effectively from the teacher network&#x00027;s knowledge. This facilitates the transfer of prior knowledge from the teacher, enhancing the generalization ability of the student model. By carefully tuning &#x003B1; and &#x003B2;, an optimal balance can be achieved that maximizes the use of the teacher network&#x00027;s knowledge while maintaining strong classification performance. The specific values of these hyperparameters should be determined through experimentation and tuning based on the task and dataset.</p></sec></sec>
<sec id="s4">
<title>4 Experiments</title>
<sec>
<title>4.1 Dataset</title>
<p>MedMNIST v2 is a large-scale collection of standardized biomedical images, designed similarly to the widely used MNIST dataset. It includes 12 different datasets for 2D images, with detailed descriptions provided in <xref ref-type="table" rid="T1">Table 1</xref>. All images in MedMNIST v2 have been pre-processed to a uniform size of 28 &#x000D7; 28 pixels (2D). The dataset spans a wide range of primary data modalities in biomedical imaging and is specifically designed for lightweight classification tasks on 2D images. MedMNIST v2 accommodates datasets of varying scales, ranging from as few as 100 samples to as many as 100,000 samples. It supports various classification tasks, including binary/multi-class, ordinal regression, and multi-label classification. In total, MedMNIST v2 contains an impressive 708,069 2D images, providing a rich and diverse set of data for experimental analysis.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Detailed information of the medical image dataset.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>MedMNIST2D</bold></th>
<th valign="top" align="center"><bold>PathMNIST</bold></th>
<th valign="top" align="center"><bold>ChestMNIST</bold></th>
<th valign="top" align="center"><bold>DermaMNIST</bold></th>
<th valign="top" align="center"><bold>OCTMNIST</bold></th>
<th valign="top" align="center"><bold>PneumoniaMNIST</bold></th>
<th valign="top" align="center"><bold>RetinaMNIST</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Data Modality</td>
<td valign="top" align="center">Colon Pathology</td>
<td valign="top" align="center">Chest X-Ray</td>
<td valign="top" align="center">Dermatoscope</td>
<td valign="top" align="center">Retinal OCT</td>
<td valign="top" align="center">Chest X-Ray</td>
<td valign="top" align="center">Fundus Camera</td>
</tr> <tr style="background-color:#919498;color:#ffffff">
<td valign="top" align="left"><bold>Tasks(Labels)</bold></td>
<td valign="top" align="center"><bold>Multi-class (9)</bold></td>
<td valign="top" align="center"><bold>Multi-label (14) Binary-class (2)</bold></td>
<td valign="top" align="center"><bold>Multi-class (7)</bold></td>
<td valign="top" align="center"><bold>Multi-class (4)</bold></td>
<td valign="top" align="center"><bold>Binary-class (2)</bold></td>
<td valign="top" align="center"><bold>Ordinal regression (5)</bold></td>
</tr> <tr>
<td valign="top" align="left">Samples</td>
<td valign="top" align="center">107,180</td>
<td valign="top" align="center">112,120</td>
<td valign="top" align="center">10,015</td>
<td valign="top" align="center">109,309</td>
<td valign="top" align="center">5,856</td>
<td valign="top" align="center">1,600</td>
</tr> <tr>
<td valign="top" align="left">Training</td>
<td valign="top" align="center">89,996</td>
<td valign="top" align="center">78,468</td>
<td valign="top" align="center">7,007</td>
<td valign="top" align="center">97,477</td>
<td valign="top" align="center">4,708</td>
<td valign="top" align="center">1,080</td>
</tr> <tr>
<td valign="top" align="left">Validation</td>
<td valign="top" align="center">10,004</td>
<td valign="top" align="center">11,219</td>
<td valign="top" align="center">1,003</td>
<td valign="top" align="center">10,832</td>
<td valign="top" align="center">524</td>
<td valign="top" align="center">120</td>
</tr> <tr>
<td valign="top" align="left">Test</td>
<td valign="top" align="center">7,180</td>
<td valign="top" align="center">22,433</td>
<td valign="top" align="center">2,005</td>
<td valign="top" align="center">1,000</td>
<td valign="top" align="center">624</td>
<td valign="top" align="center">400</td>
</tr> <tr style="background-color:#919498;color:#ffffff">
<td valign="top" align="left"><bold>MedMNIST2D</bold></td>
<td valign="top" align="center"><bold>BreastMNIST</bold></td>
<td valign="top" align="center"><bold>BloodMNIST</bold></td>
<td valign="top" align="center"><bold>TissueMNIST</bold></td>
<td valign="top" align="center"><bold>OrganAMNIST</bold></td>
<td valign="top" align="center"><bold>OrganCMNIST</bold></td>
<td valign="top" align="center"><bold>OrganSMNIST</bold></td>
</tr> <tr>
<td valign="top" align="left">Data modality</td>
<td valign="top" align="center">Breastultrasound</td>
<td valign="top" align="center">Blood cellmicroscope</td>
<td valign="top" align="center">Kidney cortex Microscope</td>
<td valign="top" align="center">Abdominal CT</td>
<td valign="top" align="center">Abdominal CT</td>
<td valign="top" align="center">Abdominal CT</td>
</tr> <tr style="background-color:#919498;color:#ffffff">
<td valign="top" align="left"><bold>Tasks (Labels)</bold></td>
<td valign="top" align="center"><bold>Binary-class (2)</bold></td>
<td valign="top" align="center"><bold>Multi-class (8)</bold></td>
<td valign="top" align="center"><bold>Multi-class (8)</bold></td>
<td valign="top" align="center"><bold>Multi-class (11)</bold></td>
<td valign="top" align="center"><bold>Multi-class (11)</bold></td>
<td valign="top" align="center"><bold>Multi-class (11)</bold></td>
</tr> <tr>
<td valign="top" align="left">Samples</td>
<td valign="top" align="center">780</td>
<td valign="top" align="center">17,092</td>
<td valign="top" align="center">236,386</td>
<td valign="top" align="center">58,830</td>
<td valign="top" align="center">23,583</td>
<td valign="top" align="center">25,211</td>
</tr> <tr>
<td valign="top" align="left">Training</td>
<td valign="top" align="center">546</td>
<td valign="top" align="center">11,959</td>
<td valign="top" align="center">165,466</td>
<td valign="top" align="center">34,561</td>
<td valign="top" align="center">12,975</td>
<td valign="top" align="center">13,932</td>
</tr> <tr>
<td valign="top" align="left">Validation</td>
<td valign="top" align="center">78</td>
<td valign="top" align="center">1,712</td>
<td valign="top" align="center">23,640</td>
<td valign="top" align="center">6,491</td>
<td valign="top" align="center">2,392</td>
<td valign="top" align="center">2,452</td>
</tr> <tr>
<td valign="top" align="left">Test</td>
<td valign="top" align="center">156</td>
<td valign="top" align="center">3,421</td>
<td valign="top" align="center">47,280</td>
<td valign="top" align="center">17,778</td>
<td valign="top" align="center">8,216</td>
<td valign="top" align="center">8,827</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>The table presents various medical image modalities, including 2D medical images, pathology, chest X-ray, dermoscopy, retinal OCT, pneumonia X-ray, and fundus camera. Each modality has different tasks and labels associated with it. The table also provides the sample count for each dataset, as well as the division into training, validation, and test sets.</p>
</table-wrap-foot>
</table-wrap></sec>
<sec>
<title>4.2 Implementation details</title>
<sec>
<title>4.2.1 Training process</title>
<p>All experiments are conducted using NVIDIA RTX 3090 GPUs within a PyTorch framework. Input images are uniformly resized to 224 &#x000D7; 224 pixels, and the number of input channels is fixed at 3. For grayscale images, we replicate the single channel to form RGB inputs. To identify the best-performing models on the MedMNIST v2 dataset, we adopt an early stopping strategy based on validation performance. We evaluate three convolutional neural networks: ResNet50, SENet50, and SKNet50. During training, we employ a multi-step learning rate schedule, starting with an initial learning rate of 0.001. The learning rate is reduced by a factor of 0.1 at epochs 50 and 75. The models are optimized using the Adam optimizer with a batch size of 128 and trained for 100 epochs. For multi-label and binary classification tasks, we use the binary cross-entropy loss function; for all other tasks, standard cross-entropy loss is applied.</p>
<p>For experiments involving the DINOv2 framework, we assess the performance of three backbone variants: DINOv2-ViT-S/14, DINOv2-ViT-B/14, and DINOv2-ViT-L/14, across all 12 datasets in MedMNIST v2. In these experiments, the DINOv2 backbone remains frozen, and only the classification head is fine-tuned. No data augmentation is applied. The models are trained using the Adam optimizer with an initial learning rate of 0.001 and a batch size of 32. The learning rate schedule mirrors that of the CNN models, with reductions at epochs 50 and 75. Cross-entropy loss is used for all classification tasks. For the MedAlmighty experiments, we adopt the same training settings as used in the DINOv2 experiments. Specifically, the best-performing DINOv2 model with a ViT-B/14 backbone is selected as the teacher, and ResNet50 is used as the student model. We explore the impact of knowledge distillation on classification performance across the MedMNIST v2 datasets.</p></sec>
<sec>
<title>4.2.2 Evaluation metrics</title>
<p>To enable a fair comparison with the baseline methods used in MedMNIST v2, we adopt the same evaluation metrics: Area Under the Curve (AUC) and Accuracy (ACC).</p>
<p><bold>AUC</bold> serves as a comprehensive performance metric that captures the trade-off between the true positive rate and false positive rate across various threshold settings. It is particularly useful for evaluating classification models on imbalanced datasets, where traditional accuracy metrics may be misleading. A higher AUC indicates stronger discriminatory ability.</p>
<p><bold>ACC</bold> measures the proportion of correctly classified samples across the entire dataset. As a straightforward and intuitive metric, it provides a general assessment of a model&#x00027;s overall classification accuracy. Higher ACC values denote better performance.</p></sec></sec>
<sec>
<title>4.3 Experimental results</title>
<sec>
<title>4.3.1 Inconsistent improvement of DINOv2 compared to CNNs</title>
<p>We evaluate three CNN-based architectures&#x02014;ResNet50, SENet50, and SKNet50&#x02014;and compare them against transformer-based models from the DINOv2 family: DINOv2-ViT-S/14, DINOv2-ViT-B/14, and DINOv2-ViT-L/14, as well as our proposed MedAlmighty framework. Their performances in terms of AUC and ACC across the 12 MedMNIST v2 datasets are summarized in <xref ref-type="table" rid="T2">Table 2</xref>.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Performance comparison of CNN, DINOv2 and CA-MKD models on 12 MedMNIST v2 datasets.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Methods</bold></th>
<th valign="top" align="center" colspan="2"><bold>PathMNIST</bold></th>
<th valign="top" align="center" colspan="2"><bold>ChestMNIST</bold></th>
<th valign="top" align="center" colspan="2"><bold>DermaMNIST</bold></th>
<th valign="top" align="center" colspan="2"><bold>OCTMNIST</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919498;color:#ffffff">
<td/>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
</tr> <tr>
<td valign="top" align="left">ResNet-50</td>
<td valign="top" align="center">0.978 &#x000B1; 0.007</td>
<td valign="top" align="center">0.854 &#x000B1; 0.028</td>
<td valign="top" align="center">0.768 &#x000B1; 0.004</td>
<td valign="top" align="center">0.947 &#x000B1; 0.001</td>
<td valign="top" align="center">0.901 &#x000B1; 0.011</td>
<td valign="top" align="center">0.721 &#x000B1; 0.011</td>
<td valign="top" align="center">0.948 &#x000B1; 0.016</td>
<td valign="top" align="center">0.776 &#x000B1; 0.002</td>
</tr> <tr>
<td valign="top" align="left">SENet-50</td>
<td valign="top" align="center">0.982 &#x000B1; 0.002</td>
<td valign="top" align="center">0.863 &#x000B1; 0.015</td>
<td valign="top" align="center"><bold>0.774</bold> <bold>&#x000B1;0.01</bold></td>
<td valign="top" align="center">0.948 &#x000B1; 0.003</td>
<td valign="top" align="center">0.989 &#x000B1; 0.004</td>
<td valign="top" align="center">0.731 &#x000B1; 0.007</td>
<td valign="top" align="center">0.949 &#x000B1; 0.014</td>
<td valign="top" align="center">0.780 &#x000B1; 0.058</td>
</tr> <tr>
<td valign="top" align="left">SKNet-50</td>
<td valign="top" align="center"><bold>0.986</bold> <bold>&#x000B1;0.004</bold></td>
<td valign="top" align="center">0.844 &#x000B1; 0.014</td>
<td valign="top" align="center">0.768 &#x000B1; 0.002</td>
<td valign="top" align="center">0.947 &#x000B1; 0.001</td>
<td valign="top" align="center">0.986 &#x000B1; 0.003</td>
<td valign="top" align="center">0.719 &#x000B1; 0.015</td>
<td valign="top" align="center">0.942 &#x000B1; 0.017</td>
<td valign="top" align="center">0.770 &#x000B1; 0.006</td>
</tr> <tr>
<td valign="top" align="left">DINOV2-vits14</td>
<td valign="top" align="center">0.986 &#x000B1; 0.005</td>
<td valign="top" align="center">0.857 &#x000B1; 0.009</td>
<td valign="top" align="center">0.666 &#x000B1; 0.016</td>
<td valign="top" align="center">0.946 &#x000B1; 0.007</td>
<td valign="top" align="center">0.903 &#x000B1; 0.021</td>
<td valign="top" align="center">0.727 &#x000B1; 0.037</td>
<td valign="top" align="center">0.924 &#x000B1; 0.010</td>
<td valign="top" align="center">0.659 &#x000B1; 0.026</td>
</tr> <tr>
<td valign="top" align="left">DINOV2-vitb14</td>
<td valign="top" align="center">0.981 &#x000B1; 0.005</td>
<td valign="top" align="center">0.870 &#x000B1; 0.051</td>
<td valign="top" align="center">0.654 &#x000B1; 0.015</td>
<td valign="top" align="center">0.943 &#x000B1; 0.005</td>
<td valign="top" align="center">0.905 &#x000B1; 0.025</td>
<td valign="top" align="center">0.725 &#x000B1; 0.034</td>
<td valign="top" align="center">0.929 &#x000B1; 0.015</td>
<td valign="top" align="center">0.629 &#x000B1; 0.017</td>
</tr> <tr>
<td valign="top" align="left">DINOV2-vitl14</td>
<td valign="top" align="center">0.979 &#x000B1; 0.002</td>
<td valign="top" align="center">0.862 &#x000B1; 0.032</td>
<td valign="top" align="center">0.649 &#x000B1; 0.022</td>
<td valign="top" align="center"><bold>0.962</bold> <bold>&#x000B1;0.007</bold></td>
<td valign="top" align="center">0.901 &#x000B1; 0.029</td>
<td valign="top" align="center">0.732 &#x000B1; 0.042</td>
<td valign="top" align="center">0.933 &#x000B1; 0.011</td>
<td valign="top" align="center">0.663 &#x000B1; 0.021</td>
</tr> <tr>
<td valign="top" align="left">CA-MKD</td>
<td valign="top" align="center">0.966 &#x000B1; 0.002</td>
<td valign="top" align="center">0.832 &#x000B1; 0.022</td>
<td valign="top" align="center">0.762 &#x000B1; 0.004</td>
<td valign="top" align="center">0.947 &#x000B1; 0.001</td>
<td valign="top" align="center">0.898 &#x000B1; 0.001</td>
<td valign="top" align="center">0.700 &#x000B1; 0.008</td>
<td valign="top" align="center">0.933 &#x000B1; 0.01</td>
<td valign="top" align="center">0.766 &#x000B1; 0.002</td>
</tr> <tr>
<td valign="top" align="left">MedAlmighty</td>
<td valign="top" align="center">0.955 &#x000B1; 0.03</td>
<td valign="top" align="center"><bold>0.871</bold> <bold>&#x000B1;0.02</bold></td>
<td valign="top" align="center">0.643 &#x000B1; 0.05</td>
<td valign="top" align="center">0.927 &#x000B1; 0.02</td>
<td valign="top" align="center"><bold>0.905</bold> <bold>&#x000B1;0.007</bold></td>
<td valign="top" align="center"><bold>0.735</bold> <bold>&#x000B1;0.005</bold></td>
<td valign="top" align="center"><bold>0.951</bold> <bold>&#x000B1;0.008</bold></td>
<td valign="top" align="center"><bold>0.781</bold> <bold>&#x000B1;0.003</bold></td>
</tr> <tr style="background-color:#919498;color:#ffffff">
<td valign="top" align="left"><bold>Methods</bold></td>
<td valign="top" align="center" colspan="2"><bold>BreastMNIST</bold></td>
<td valign="top" align="center" colspan="2"><bold>BloodMNIST</bold></td>
<td valign="top" align="center" colspan="2"><bold>TissueMNIST</bold></td>
<td valign="top" align="center" colspan="2"><bold>PneumoniaMNIST</bold></td>
</tr>
 <tr style="background-color:#919498;color:#ffffff">
<td/>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
</tr> <tr>
<td valign="top" align="left">ResNet-50</td>
<td valign="top" align="center">0.883 &#x000B1; 0.015</td>
<td valign="top" align="center"><bold>0.843</bold> <bold>&#x000B1;0.004</bold></td>
<td valign="top" align="center">0.994 &#x000B1; 0.001</td>
<td valign="top" align="center">0.950 &#x000B1; 0.012</td>
<td valign="top" align="center">0.931 &#x000B1; 0.005</td>
<td valign="top" align="center">0.680 &#x000B1; 0.011</td>
<td valign="top" align="center">0.962 &#x000B1; 0.005</td>
<td valign="top" align="center">0.884 &#x000B1; 0.007</td>
</tr> <tr>
<td valign="top" align="left">SENet-50</td>
<td valign="top" align="center">0.876 &#x000B1; 0.001</td>
<td valign="top" align="center">0.827 &#x000B1; 0.006</td>
<td valign="top" align="center"><bold>0.995</bold> <bold>&#x000B1;0.002</bold></td>
<td valign="top" align="center">0.960 &#x000B1; 0.005</td>
<td valign="top" align="center">0.926 &#x000B1; 0.004</td>
<td valign="top" align="center">0.687 &#x000B1; 0.013</td>
<td valign="top" align="center">0.949 &#x000B1; 0.016</td>
<td valign="top" align="center">0.816 &#x000B1; 0.058</td>
</tr> <tr>
<td valign="top" align="left">SKNet-50</td>
<td valign="top" align="center">0.859 &#x000B1; 0.031</td>
<td valign="top" align="center">0.818 &#x000B1; 0.009</td>
<td valign="top" align="center">0.993 &#x000B1; 0.002</td>
<td valign="top" align="center">0.956 &#x000B1; 0.014</td>
<td valign="top" align="center">0.925 &#x000B1; 0.003</td>
<td valign="top" align="center">0.687 &#x000B1; 0.005</td>
<td valign="top" align="center">0.958 &#x000B1; 0.006</td>
<td valign="top" align="center">0.839 &#x000B1; 0.018</td>
</tr> <tr>
<td valign="top" align="left">DINOV2-vits14</td>
<td valign="top" align="center">0.871 &#x000B1; 0.015</td>
<td valign="top" align="center">0.853 &#x000B1; 0.012</td>
<td valign="top" align="center">0.992 &#x000B1; 0.004</td>
<td valign="top" align="center">0.926 &#x000B1; 0.016</td>
<td valign="top" align="center">0.902 &#x000B1; 0.007</td>
<td valign="top" align="center">0.610 &#x000B1; 0.015</td>
<td valign="top" align="center">0.963 &#x000B1; 0.02</td>
<td valign="top" align="center">0.893 &#x000B1; 0.016</td>
</tr> <tr>
<td valign="top" align="left">DINOV2-vitb14</td>
<td valign="top" align="center"><bold>0.893</bold> <bold>&#x000B1;0.024</bold></td>
<td valign="top" align="center">0.827 &#x000B1; 0.008</td>
<td valign="top" align="center">0.993 &#x000B1; 0.006</td>
<td valign="top" align="center">0.930 &#x000B1; 0.007</td>
<td valign="top" align="center">0.916 &#x000B1; 0.004</td>
<td valign="top" align="center">0.647 &#x000B1; 0.018</td>
<td valign="top" align="center">0.962 &#x000B1; 0.017</td>
<td valign="top" align="center">0.890 &#x000B1; 0.021</td>
</tr> <tr>
<td valign="top" align="left">DINOV2-vitl14</td>
<td valign="top" align="center">0.853 &#x000B1; 0.017</td>
<td valign="top" align="center">0.842 &#x000B1; 0.022</td>
<td valign="top" align="center">0.994 &#x000B1; 0.009</td>
<td valign="top" align="center">0.928 &#x000B1; 0.010</td>
<td valign="top" align="center">0.911 &#x000B1; 0.002</td>
<td valign="top" align="center">0.632 &#x000B1; 0.021</td>
<td valign="top" align="center"><bold>0.963</bold> <bold>&#x000B1;0.009</bold></td>
<td valign="top" align="center">0.872 &#x000B1; 0.017</td>
</tr> <tr>
<td valign="top" align="left">CA-MKD</td>
<td valign="top" align="center">0.877 &#x000B1; 0.01</td>
<td valign="top" align="center">0.830 &#x000B1; 0.002</td>
<td valign="top" align="center">0.991 &#x000B1; 0.001</td>
<td valign="top" align="center">0.949 &#x000B1; 0.011</td>
<td valign="top" align="center">0.928 &#x000B1; 0.002</td>
<td valign="top" align="center">0.662 &#x000B1; 0.001</td>
<td valign="top" align="center">0.962 &#x000B1; 0.001</td>
<td valign="top" align="center">0.882 &#x000B1; 0.003</td>
</tr> <tr>
<td valign="top" align="left">MedAlmighty</td>
<td valign="top" align="center">0.837 &#x000B1; 0.03</td>
<td valign="top" align="center">0.833 &#x000B1; 0.01</td>
<td valign="top" align="center"><bold>0.995</bold> <bold>&#x000B1;0.002</bold></td>
<td valign="top" align="center"><bold>0.960</bold> <bold>&#x000B1;0.004</bold></td>
<td valign="top" align="center"><bold>0.932</bold> <bold>&#x000B1;0.002</bold></td>
<td valign="top" align="center"><bold>0.693</bold> <bold>&#x000B1;0.005</bold></td>
<td valign="top" align="center">0.946 &#x000B1; 0.003</td>
<td valign="top" align="center"><bold>0.915</bold> <bold>&#x000B1;0.003</bold></td>
</tr> <tr style="background-color:#919498;color:#ffffff">
<td valign="top" align="left"><bold>Methods</bold></td>
<td valign="top" align="center" colspan="2"><bold>RetinaMNIST</bold></td>
<td valign="top" align="center" colspan="2"><bold>OrganAMNIST</bold></td>
<td valign="top" align="center" colspan="2"><bold>OrganCMNIST</bold></td>
<td valign="top" align="center" colspan="2"><bold>OrganSMNIST</bold></td>
</tr>
 <tr style="background-color:#919498;color:#ffffff">
<td/>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
</tr> <tr>
<td valign="top" align="left">ResNet-50</td>
<td valign="top" align="center">0.713 &#x000B1; 0.008</td>
<td valign="top" align="center">0.500 &#x000B1; 0.010</td>
<td valign="top" align="center">0.997 &#x000B1; 0.001</td>
<td valign="top" align="center">0.944 &#x000B1; 0.002</td>
<td valign="top" align="center">0.991 &#x000B1; 0.001</td>
<td valign="top" align="center">0.910 &#x000B1; 0.001</td>
<td valign="top" align="center">0.963 &#x000B1; 0.003</td>
<td valign="top" align="center">0.766 &#x000B1; 0.017</td>
</tr> <tr>
<td valign="top" align="left">SENet-50</td>
<td valign="top" align="center">0.729 &#x000B1; 0.014</td>
<td valign="top" align="center"><bold>0.524</bold> <bold>&#x000B1;0.018</bold></td>
<td valign="top" align="center">0.997 &#x000B1; 0.001</td>
<td valign="top" align="center">0.941 &#x000B1; 0.003</td>
<td valign="top" align="center">0.992 &#x000B1; 0.001</td>
<td valign="top" align="center">0.905 &#x000B1; 0.002</td>
<td valign="top" align="center">0.963 &#x000B1; 0.001</td>
<td valign="top" align="center">0.765 &#x000B1; 0.019</td>
</tr> <tr>
<td valign="top" align="left">SKNet-50</td>
<td valign="top" align="center">0.728 &#x000B1; 0.007</td>
<td valign="top" align="center">0.520 &#x000B1; 0.006</td>
<td valign="top" align="center">0.997 &#x000B1; 0.002</td>
<td valign="top" align="center">0.935 &#x000B1; 0.005</td>
<td valign="top" align="center">0.991 &#x000B1; 0.002</td>
<td valign="top" align="center">0.907 &#x000B1; 0.001</td>
<td valign="top" align="center">0.954 &#x000B1; 0.032</td>
<td valign="top" align="center">0.778 &#x000B1; 0.006</td>
</tr> <tr>
<td valign="top" align="left">DINOV2-vits14</td>
<td valign="top" align="center"><bold>0.738</bold> <bold>&#x000B1;0.022</bold></td>
<td valign="top" align="center">0.514 &#x000B1; 0.020</td>
<td valign="top" align="center">0.992 &#x000B1; 0.005</td>
<td valign="top" align="center">0.880 &#x000B1; 0.008</td>
<td valign="top" align="center">0.982 &#x000B1; 0.010</td>
<td valign="top" align="center">0.863 &#x000B1; 0.025</td>
<td valign="top" align="center">0.961 &#x000B1; 0.031</td>
<td valign="top" align="center">0.754 &#x000B1; 0.006</td>
</tr> <tr>
<td valign="top" align="left">DINOV2-vitb14</td>
<td valign="top" align="center">0.735 &#x000B1; 0.038</td>
<td valign="top" align="center">0.489 &#x000B1; 0.020</td>
<td valign="top" align="center">0.989 &#x000B1; 0.001</td>
<td valign="top" align="center">0.889 &#x000B1; 0.002</td>
<td valign="top" align="center">0.981 &#x000B1; 0.005</td>
<td valign="top" align="center">0.851 &#x000B1; 0.001</td>
<td valign="top" align="center">0.962 &#x000B1; 0.029</td>
<td valign="top" align="center">0.750 &#x000B1; 0.012</td>
</tr> <tr>
<td valign="top" align="left">DINOV2-vitl14</td>
<td valign="top" align="center">0.730 &#x000B1; 0.047</td>
<td valign="top" align="center">0.450 &#x000B1; 0.025</td>
<td valign="top" align="center">0.991 &#x000B1; 0.002</td>
<td valign="top" align="center">0.892 &#x000B1; 0.002</td>
<td valign="top" align="center">0.978 &#x000B1; 0.007</td>
<td valign="top" align="center">0.842 &#x000B1; 0.001</td>
<td valign="top" align="center">0.960 &#x000B1; 0.017</td>
<td valign="top" align="center">0.748 &#x000B1; 0.011</td>
</tr> <tr>
<td valign="top" align="left">CA-MKD</td>
<td valign="top" align="center">0.701 &#x000B1; 0.003</td>
<td valign="top" align="center">0.489 &#x000B1; 0.006</td>
<td valign="top" align="center">0.997 &#x000B1; 0.004</td>
<td valign="top" align="center">0.931 &#x000B1; 0.002</td>
<td valign="top" align="center">0.991 &#x000B1; 0.001</td>
<td valign="top" align="center">0.908 &#x000B1; 0.001</td>
<td valign="top" align="center">0.954 &#x000B1; 0.001</td>
<td valign="top" align="center">0.744 &#x000B1; 0.011</td>
</tr> <tr>
<td valign="top" align="left">MedAlmighty</td>
<td valign="top" align="center">0.645 &#x000B1; 0.1</td>
<td valign="top" align="center">0.510 &#x000B1; 0.025</td>
<td valign="top" align="center"><bold>0.998</bold> <bold>&#x000B1;0.001</bold></td>
<td valign="top" align="center"><bold>0.952</bold> <bold>&#x000B1;0.002</bold></td>
<td valign="top" align="center"><bold>0.994</bold> <bold>&#x000B1;0.001</bold></td>
<td valign="top" align="center"><bold>0.915</bold> <bold>&#x000B1;0.002</bold></td>
<td valign="top" align="center"><bold>0.975</bold> <bold>&#x000B1;0.001</bold></td>
<td valign="top" align="center"><bold>0.782</bold> <bold>&#x000B1;0.001</bold></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>Values represent mean &#x000B1; standard deviation over 3 independent runs (highest metric per dataset in bold).</p>
</table-wrap-foot>
</table-wrap>
<p>The results reveal that both CNN and DINOv2 models exhibit varying performance across different medical image classification tasks. Among the CNN models, ResNet50 consistently outperforms SENet50 and SKNet50 on most datasets. However, DINOv2 only shows marginal improvements on a subset of datasets&#x02014;namely, DermaMNIST, RetinaMNIST, and PathMNIST. Although DINOv2 achieves competitive results in some cases, its overall performance is inconsistent and does not consistently surpass that of the CNN models.</p>
<p>To further explore this, we compare DINOv2-ViT-B/14 directly with ResNet50 using AUC and ACC metrics, as shown in <xref ref-type="fig" rid="F2">Figures 2</xref>, <xref ref-type="fig" rid="F3">3</xref>, respectively. These comparisons indicate that DINOv2-ViT-B/14 offers improvements in only a limited number of datasets, and in several cases, its performance falls short of that achieved by ResNet50.</p>
<fig position="float" id="F2">
<label>Figure 2</label>
<caption><p>Comparing AUC values of DINOv2-ViTs14 with ResNet18, DINOv2-ViTb14 with ResNet50, and DINOv2-ViTl14 with ResNet50 on 12 MedMNIST datasets. Results are based on experiments using MedMNISTV2, where all models were evaluated on 224 &#x000D7; 224 images.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1527980-g0002.tif">
<alt-text>Bar chart comparing error rates across various medical datasets using three models: DINOv2-ViT-S14, DINOv2-ViT-B14, and DINOv2-ViT-I14. Each dataset is represented on the x-axis with error rates on the y-axis, indicating performance differences.</alt-text>
</graphic>
</fig>
<fig position="float" id="F3">
<label>Figure 3</label>
<caption><p>Comparing ACC values of DINOv2-ViTs14 with ResNet18, DINOv2-ViTb14 with ResNet50, and DINOv2-ViTl14 with ResNet50 on 12 MedMNIST datasets. Results are based on experiments using MedMNISTV2, where all models were evaluated on 224 &#x000D7; 224 images.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1527980-g0003.tif">
<alt-text>Bar chart comparing the performance of three models&#x02014;DINOv2-ViTs14, DINOv2-ViTb14, and DINOv2-ViTi14&#x02014;across different datasets. The datasets are listed along the x-axis, including BreastMNIST, RetinaMNIST, and more. The y-axis represents performance metrics, with some values above and some below zero, indicating varying levels of performance across datasets. Each model is represented by different colors: blue, teal, and green.</alt-text>
</graphic>
</fig>
<p>In summary, our findings demonstrate that while DINOv2 can provide benefits on certain datasets, it fails to deliver consistent improvements over CNN models. ResNet50 remains a strong baseline, outperforming both SENet50 and SKNet50, and often matching or exceeding the performance of DINOv2 models in terms of both AUC and ACC across the MedMNIST v2 classification tasks.</p></sec>
<sec>
<title>4.3.2 CA-MKD vs. MedAlmighty comparison</title>
<p>Additionally, we benchmark against CA-MKD (<xref ref-type="bibr" rid="B50">Zhang et al., 2022</xref>) (Confidence-Aware Multi-Teacher Knowledge Distillation), which employs three ResNet32x4 teachers (138M total params) to distill knowledge into a lightweight MobileNetV2 (3.4M params). Their performances in terms of AUC and ACC across the 12 MedMNIST v2 datasets are summarized in <xref ref-type="table" rid="T2">Table 2</xref>. While multi-teacher distillation incorporates features from multiple teacher models, it merely expands feature quantity without ensuring comprehensive coverage. In contrast, our MedAlmighty method incorporates more comprehensive feature considerations during classification. Experimental results on this dataset demonstrate unsatisfactory performance of the selected multi-teacher distillation model&#x02014;likely attributable to the identical ResNet32x4 architecture shared by all three teacher models, which constrained its effectiveness. This raises the essential consideration of model selection when applying multi-teacher distillation to target datasets for classification tasks. Furthermore, multi-teacher distillation imposes significantly higher computational demands during training. Compared to this approach, our method demonstrates superior performance and efficiency.</p></sec>
<sec>
<title>4.3.3 Enhancing classification performance with MedAlmighty</title>
<p>To investigate how to improve the classification performance of CNNs and DINOv2 in medical imaging, we adopt a knowledge distillation approach by transferring learned representations from DINOv2 to ResNet50. This results in our proposed framework, <italic>MedAlmighty</italic>, whose performance across the MedMNIST v2 datasets is presented in <xref ref-type="table" rid="T2">Table 2</xref>. Experimental results demonstrate that MedAlmighty consistently outperforms both standalone CNNs and DINOv2 on the majority of datasets. This performance gain can be attributed to the integration of DINOv2&#x00027;s rich visual representation capabilities and the efficiency of CNNs, which typically have lower parameter counts than large-scale vision models. The parameter counts for CNN models, DINOv2, and MedAlmighty are summarized in <xref ref-type="table" rid="T3">Table 3</xref>. These findings highlight the effectiveness of leveraging knowledge distillation from large vision models to enhance lightweight CNNs for medical image classification. Notably, MedAlmighty achieves these improvements without the use of additional data augmentation techniques, which are often computationally intensive. This demonstrates that our distillation-based approach can significantly improve classification performance while maintaining training efficiency. Overall, MedAlmighty offers a promising direction for applying large vision models like DINOv2 in the medical domain, particularly in scenarios with limited labeled data. By efficiently transferring generalizable knowledge, it enhances classification accuracy while preserving the compactness and interpretability of CNN architectures.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Model parameters are measured in millions.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>ResNet50</bold></th>
<th valign="top" align="center"><bold>SENet50</bold></th>
<th valign="top" align="center"><bold>SKNet50</bold></th>
<th valign="top" align="center"><bold>DINOV2-ViTs14</bold></th>
<th valign="top" align="center"><bold>DINOV2-ViTb14</bold></th>
<th valign="top" align="center"><bold>DINOV2-ViTl14</bold></th>
<th valign="top" align="center"><bold>MedAlmighty</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Model parameter count</td>
<td valign="top" align="center">23M</td>
<td valign="top" align="center">26M</td>
<td valign="top" align="center">23M</td>
<td valign="top" align="center">22M</td>
<td valign="top" align="center">86M</td>
<td valign="top" align="center">304M</td>
<td valign="top" align="center">23M</td>
</tr></tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>4.3.4 Parameter ablation experiments</title>
<p>In the parameter ablation experiments for MedAlmighty, we explored the effects of varying the distillation temperature <italic>t</italic> and the hyperparameter &#x003B1; across three different combinations. In the first group, we set the distillation temperature <italic>t</italic> &#x0003D; 2.0 and &#x003B1; &#x0003D; 0.2. In the second group, <italic>t</italic> &#x0003D; 5.0 and &#x003B1; &#x0003D; 0.5, and in the third group, <italic>t</italic> &#x0003D; 8.0 and &#x003B1; &#x0003D; 0.8. The results, as shown in <xref ref-type="table" rid="T4">Table 4</xref>, revealed that the model performed best when both the distillation temperature <italic>t</italic> and &#x003B1; were set to lower values. In further experiments, we fixed the distillation temperature <italic>t</italic> at 2.0 and varied &#x003B1; from 0.1 to 0.9 and we then fixed &#x003B1; at 0.2 and varied <italic>t</italic> from 1 to 9, as shown in <xref ref-type="fig" rid="F4">Figure 4</xref>. The analysis of these experiments indicated that MedAlmighty exhibited superior performance when the distillation temperature <italic>t</italic> was set to a lower value and &#x003B1; was within the range of 0 to 0.5. Furthermore, varying the distillation temperature had little impact on the results when &#x003B1; was kept at a smaller value.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>AUC and ACC results of distilling ResNet50 with DINOv2 ViTb14 on 12 MedMNIST datasets, with varying distillation temperature <italic>t</italic> and &#x003B1; parameters( <italic>t</italic>/&#x003B1;).</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Methods</bold></th>
<th valign="top" align="center" colspan="2"><bold>PathMNIST</bold></th>
<th valign="top" align="center" colspan="2"><bold>ChestMNIST</bold></th>
<th valign="top" align="center" colspan="2"><bold>DermaMNIST</bold></th>
<th valign="top" align="center" colspan="2"><bold>OCTMNIST</bold></th>
<th valign="top" align="center" colspan="2"><bold>PneumoniaMNIST</bold></th>
<th valign="top" align="center" colspan="2"><bold>RetinaMNIST</bold></th>
</tr>
<tr style="background-color:#919498;color:#ffffff">
<th/>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold>ACC</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold>ACC</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold>ACC</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold>ACC</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold>ACC</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold>ACC</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">MedAlmighty (2.0/0.2)</td>
<td valign="top" align="center"><bold>0.975</bold></td>
<td valign="top" align="center">0.883</td>
<td valign="top" align="center"><bold>0.686</bold></td>
<td valign="top" align="center">0.946</td>
<td valign="top" align="center"><bold>0.911</bold></td>
<td valign="top" align="center">0.735</td>
<td valign="top" align="center">0.954</td>
<td valign="top" align="center">0.781</td>
<td valign="top" align="center"><bold>0.946</bold></td>
<td valign="top" align="center"><bold>0.915</bold></td>
<td valign="top" align="center"><bold>0.735</bold></td>
<td valign="top" align="center"><bold>0.535</bold></td>
</tr> <tr>
<td valign="top" align="left">MedAlmighty (5.0/0.5)</td>
<td valign="top" align="center">0.968</td>
<td valign="top" align="center">0.878</td>
<td valign="top" align="center">0.671</td>
<td valign="top" align="center"><bold>0.947</bold></td>
<td valign="top" align="center">0.910</td>
<td valign="top" align="center"><bold>0.738</bold></td>
<td valign="top" align="center"><bold>0.965</bold></td>
<td valign="top" align="center"><bold>0.804</bold></td>
<td valign="top" align="center">0.917</td>
<td valign="top" align="center">0.909</td>
<td valign="top" align="center">0.723</td>
<td valign="top" align="center">0.515</td>
</tr> 
<tr>
<td valign="top" align="left">MedAlmighty (8.0/0.8)</td>
<td valign="top" align="center">0.962</td>
<td valign="top" align="center"><bold>0.913</bold></td>
<td valign="top" align="center">0.642</td>
<td valign="top" align="center"><bold>0.947</bold></td>
<td valign="top" align="center">0.894</td>
<td valign="top" align="center">0.726</td>
<td valign="top" align="center">0.946</td>
<td valign="top" align="center">0.783</td>
<td valign="top" align="center">0.875</td>
<td valign="top" align="center">0.638</td>
<td valign="top" align="center">0.718</td>
<td valign="top" align="center">0.458</td>
</tr> <tr style="background-color:#919498;color:#ffffff">
<td valign="top" align="left"><bold>Methods</bold></td>
<td valign="top" align="center" colspan="2"><bold>BreastMNIST</bold></td>
<td valign="top" align="center" colspan="2"><bold>BloodMNIST</bold></td>
<td valign="top" align="center" colspan="2"><bold>TissueMNIST</bold></td>
<td valign="top" align="center" colspan="2"><bold>OrganAMNIST</bold></td>
<td valign="top" align="center" colspan="2"><bold>OrganCMNIST</bold></td>
<td valign="top" align="center" colspan="2"><bold>OrganSMNIST</bold></td>
</tr>
<tr style="background-color:#919498;color:#ffffff">
<td/>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
<td valign="top" align="center"><bold>AUC</bold></td>
<td valign="top" align="center"><bold>ACC</bold></td>
</tr> <tr>
<td valign="top" align="left">MedAlmighty (2.0/0.2)</td>
<td valign="top" align="center"><bold>0.857</bold></td>
<td valign="top" align="center"><bold>0.833</bold></td>
<td valign="top" align="center">0.995</td>
<td valign="top" align="center"><bold>0.962</bold></td>
<td valign="top" align="center"><bold>0.932</bold></td>
<td valign="top" align="center"><bold>0.693</bold></td>
<td valign="top" align="center"><bold>0.998</bold></td>
<td valign="top" align="center"><bold>0.952</bold></td>
<td valign="top" align="center"><bold>0.994</bold></td>
<td valign="top" align="center">0.915</td>
<td valign="top" align="center"><bold>0.975</bold></td>
<td valign="top" align="center">0.783</td>
</tr> <tr>
<td valign="top" align="left">MedAlmighty (5.0/0.5)</td>
<td valign="top" align="center">0.834</td>
<td valign="top" align="center">0.814</td>
<td valign="top" align="center"><bold>0.997</bold></td>
<td valign="top" align="center">0.953</td>
<td valign="top" align="center">0.931</td>
<td valign="top" align="center">0.690</td>
<td valign="top" align="center">0.996</td>
<td valign="top" align="center">0.926</td>
<td valign="top" align="center">0.989</td>
<td valign="top" align="center"><bold>0.921</bold></td>
<td valign="top" align="center"><bold>0.975</bold></td>
<td valign="top" align="center"><bold>0.789</bold></td>
</tr> <tr>
<td valign="top" align="left">MedAlmighty (8.0/0.8)</td>
<td valign="top" align="center">0.699</td>
<td valign="top" align="center">0.513</td>
<td valign="top" align="center">0.994</td>
<td valign="top" align="center">0.945</td>
<td valign="top" align="center">0.925</td>
<td valign="top" align="center">0.683</td>
<td valign="top" align="center">0.996</td>
<td valign="top" align="center">0.925</td>
<td valign="top" align="center">0.988</td>
<td valign="top" align="center">0.902</td>
<td valign="top" align="center">0.971</td>
<td valign="top" align="center">0.782</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>Bold values represent the better-performing AUC/ACC results among three parameter configurations within the same data subset.</p>
</table-wrap-foot>
</table-wrap>
<fig position="float" id="F4">
<label>Figure 4</label>
<caption><p>Performance evaluation on RetinaMNIST: AUC and ACC performance of ( <italic>t</italic>/&#x003B1;) with <italic>t</italic> = 2; AUC and ACC performance of (<italic>t</italic>/&#x003B1;) with &#x003B1;=0.2.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1527980-g0004.tif">
<alt-text>Two line graphs compare AUC and ACC metrics. The left graph shows AUC scores around 0.70 and ACC scores fluctuating around 0.50 across various ratios. The right graph displays AUC scores nearing 0.75 and ACC scores varying around 0.52, also across different ratios.</alt-text>
</graphic>
</fig>
</sec>
<sec>
<title>4.3.5 Qualitative visualization</title>
<p>To compare the performance differences among three models&#x02014;ResNet50, DINOv2-ViT-B/14, and MedAlmighty&#x02014;we visualize the t-SNE plots of normal and abnormal samples from the RetinaMNIST dataset. The results are presented in <xref ref-type="fig" rid="F5">Figure 5</xref>. For normal samples, ResNet50 demonstrates a better clustering effect, with similar samples grouped closely together, forming tight clusters. DINOv2-ViT-B/14 also exhibits some degree of clustering, but there is noticeable separation between different regions of normal samples. In contrast, MedAlmighty shows the most pronounced clustering effect, with clear separation and well-defined clusters of normal samples. For abnormal samples, ResNet50 displays slightly weaker clustering. Some abnormal samples are grouped with normal samples, although many remain distinguishable. DINOv2-ViT-B/14 shows more separation between normal and abnormal samples, but the clustering effect is less distinct. MedAlmighty, however, excels in distinguishing abnormal samples from normal ones, with abnormal samples forming a separate, clearly defined region, distinct from the normal samples. Through qualitative analysis, it is evident that MedAlmighty outperforms both ResNet50 and DINOv2-ViT-B/14 in terms of t-SNE visualization of normal and abnormal samples on the RetinaMNIST dataset. It effectively separates normal and abnormal samples, demonstrating its higher potential and application value for medical image classification tasks.</p>
<fig position="float" id="F5">
<label>Figure 5</label>
<caption><p>t-SNE visualization of features (ResNet50, DINOV2-ViTb14, MedAlmighty).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1527980-g0005.tif">
<alt-text>Three scatter plots display data categorized as &#x0201C;Normal&#x0201D; in blue and &#x0201C;Abnormal&#x0201D; in orange. The left plot shows blue points concentrated at the bottom, while orange dominates the top. The middle plot presents a more even distribution, with a mix of blue and orange points. The right plot has blue points clustered at the top and scattered among orange points throughout.</alt-text>
</graphic>
</fig>
<p><xref ref-type="fig" rid="F6">Figure 6</xref> showcases heatmap visualizations from a deep learning-based image classification task, demonstrating the model&#x00027;s robust ability to focus on key image features. The top row presents the original input images across different data categories, while the bottom row shows the corresponding heatmaps. The color-coded heatmaps (with red indicating high attention and blue indicating low attention) clearly highlight the areas the model finds most significant for classification. This ability to effectively identify and focus on critical regions underscores the model&#x00027;s strong interpretability and proficiency in understanding complex image patterns.</p>
<fig position="float" id="F6">
<label>Figure 6</label>
<caption><p>Input images <bold>(top)</bold> and heatmaps <bold>(bottom)</bold>. Color intensity reflects the relative importance of image regions for the model&#x00027;s classification.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1527980-g0006.tif">
<alt-text>A collage of medical images is shown. The top row features six diagnostic images, including scans and microscopic views. The bottom row features heatmaps corresponding to each image, highlighting areas of interest with varying color intensities.</alt-text>
</graphic>
</fig>
</sec></sec></sec>
<sec sec-type="conclusions" id="s5">
<title>5 Conclusion</title>
<p>The scarcity of diverse, well-annotated medical data and the inherent complexity of disease diagnosis across various imaging modalities underscore the need for innovative solutions in medical image analysis. In this work, we proposed <italic>MedAlmighty</italic>, a framework that integrates the robust generalization capabilities of large pre-trained vision models with the domain-specific knowledge captured by lightweight CNNs, through the technique of knowledge distillation. By distilling knowledge from DINOv2 into ResNet50, MedAlmighty successfully combines the broad semantic understanding of large vision models with the efficiency and specialization of smaller models. This approach addresses the capacity limitations of CNNs while avoiding the computational burden of fine-tuning large models end-to-end. Although initial results from frozen large models with only a trainable linear head showed limited performance, our method effectively bridges this performance gap. Extensive experiments on the MedMNIST v2 benchmark demonstrate that MedAlmighty consistently outperforms both standalone CNNs and large models across a wide range of medical classification tasks. Notably, it achieves these improvements without relying on computationally intensive data augmentation techniques. In summary, MedAlmighty illustrates the promise of leveraging knowledge distillation to combine the complementary strengths of large and small models in the medical domain. This work lays a foundation for future research into hybrid model architectures, with the goal of improving diagnostic accuracy and enabling practical, scalable deployment of AI systems in real-world healthcare settings.</p></sec>
<sec sec-type="discussion" id="s6">
<title>6 Discussion</title>
<p>The application of large-scale vision models in medical image classification has shown significant promise, but challenges persist, particularly in scenarios with limited labeled data. Although deep learning algorithms have made substantial progress, these models typically require large volumes of annotated data for effective training. Acquiring such data in the medical field is often time-consuming and difficult, making the improvement of model performance in data-scarce environments a crucial area of research. Large-scale vision models excel at feature extraction and representation learning, enabling them to capture intricate textures and complex features in medical images. These models have demonstrated success in discriminative tasks across a variety of disease domains by leveraging hierarchical feature representations. Moreover, through pre-training on extensive general image datasets, large-scale vision models can benefit from transfer learning, thereby enhancing their feature extraction capabilities for medical image tasks.</p>
<p>This study investigates the use of large-scale vision models to distill knowledge into smaller models, focusing specifically on ResNet distillation. Knowledge distillation facilitates the transfer of knowledge from a larger model to a more compact one, improving efficiency and computational performance. By utilizing the strengths of large vision models, our approach enhances disease classification performance while addressing the challenges associated with deploying large models. MedAlmighty, which combines the advantages of large vision models and compact CNNs, provides a versatile solution for medical image classification, especially when dealing with limited data and complex patterns. Importantly, MedAlmighty does not rely on data augmentation techniques. Instead, it employs simple multistep learning rate adjustments. Despite this simplicity, it consistently outperforms both ResNet50 and DINOv2 individually, showcasing its substantial potential. By using a large vision model as the teacher and guiding the student model through knowledge distillation, MedAlmighty effectively transfers valuable knowledge while maintaining the parameter efficiency of the smaller model. This strategy bridges the gap between generalization capabilities and computational efficiency, presenting a highly promising approach for medical image analysis.</p>
<p>Overall, MedAlmighty represents an innovative solution to the challenges posed by limited data and diverse imaging patterns in medical diagnostics. Its performance on the MedMNIST v2 dataset demonstrates the effectiveness of using large vision models as teachers, enhancing classification performance and discrimination capabilities in specific disease categories. These results underscore the potential of large-scale vision models in improving medical image classification, particularly for targeted disease categories within the MedMNIST v2 dataset. While the results are promising, further research is needed to evaluate the effectiveness of this approach across diverse datasets and real-world scenarios. Extending this evaluation will help validate the approach and explore its broader applicability.</p></sec>
</body>
<back>
<sec sec-type="data-availability" id="s7">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: (<xref ref-type="bibr" rid="B48">Yang et al. 2023</xref>).</p>
</sec>
<sec sec-type="author-contributions" id="s8">
<title>Author contributions</title>
<p>YR: Writing &#x02013; original draft, Writing &#x02013; review &#x00026; editing. ZG: Data curation, Visualization, Writing &#x02013; review &#x00026; editing. WL: Funding acquisition, Project administration, Writing &#x02013; review &#x00026; editing.</p>
</sec>
<sec sec-type="funding-information" id="s9">
<title>Funding</title>
<p>The author(s) declare that financial support was received for the research and/or publication of this article. This work was supported by the Key Project Program of Xinjiang Institute of Engineering (Grant No. 2024xgy062605), Tianshan Talent of Xinjiang Uygur Autonomous Region&#x02014;Young Top Talents in Science and Technology (Grant No. 2022TSYCCY0008), NSFC under grant (Grant No. 61962058), Integration of Industry and Education-Joint Laboratory of Data Engineering and Digital Mine (Grant No. 2019QX0035), Bayingolin Mongolian Autonomous Prefecture Science and Technology Research Program (Grant No. 202117), Natural Science Foundation of Xinjiang Uygur Autonomous Region (Grant No. 2019D01A30), and Scientific Research Program of the Higher Education Institution of Xinjiang (Grant Nos. XJEDU2018Y056 and XJEDU2024P081).</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="ai-statement" id="s10">
<title>Generative AI statement</title>
<p>The author(s) declare that no Gen AI was used in the creation of this manuscript.</p></sec><sec sec-type="disclaimer" id="s11">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Arafa</surname> <given-names>A. B.</given-names></name> <name><surname>El-Fishawy</surname> <given-names>N. A.</given-names></name> <name><surname>Badawy</surname> <given-names>M.</given-names></name> <name><surname>Radad</surname> <given-names>M.</given-names></name></person-group> (<year>2023</year>). <article-title>RN-autoencoder: reduced noise autoencoder for classifying imbalanced cancer genomic data</article-title>. <source>J. Biol. Eng</source>. <volume>17</volume>:<fpage>7</fpage>. <pub-id pub-id-type="doi">10.1186/s13036-022-00319-3</pub-id><pub-id pub-id-type="pmid">36717866</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Arumugam</surname> <given-names>M.</given-names></name> <name><surname>Thiyagarajan</surname> <given-names>A.</given-names></name> <name><surname>Adhi</surname> <given-names>L.</given-names></name> <name><surname>Alagar</surname> <given-names>S.</given-names></name></person-group> (<year>2024</year>). <article-title>Crossover smell agent optimized multilayer perceptron for precise brain tumor classification on mri images</article-title>. <source>Expert Syst. Appl</source>. <volume>238</volume>:<fpage>121453</fpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2023.121453</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bao</surname> <given-names>H.</given-names></name> <name><surname>Dong</surname> <given-names>L.</given-names></name> <name><surname>Piao</surname> <given-names>S.</given-names></name> <name><surname>Wei</surname> <given-names>F.</given-names></name></person-group> (<year>2021</year>). <article-title>Beit: bert pre-training of image transformers</article-title>. <source>arXiv preprint arXiv:2106.08254</source>.<pub-id pub-id-type="pmid">39164302</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Caron</surname> <given-names>M.</given-names></name> <name><surname>Touvron</surname> <given-names>H.</given-names></name> <name><surname>Misra</surname> <given-names>I.</given-names></name> <name><surname>J&#x000E9;gou</surname> <given-names>H.</given-names></name> <name><surname>Mairal</surname> <given-names>J.</given-names></name> <name><surname>Bojanowski</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>&#x0201C;Emerging properties in self-supervised vision transformers,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision</source>, 9650&#x02013;9660. <pub-id pub-id-type="doi">10.1109/ICCV48922.2021.00951</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Fu</surname> <given-names>C.</given-names></name> <name><surname>Xie</surname> <given-names>H.</given-names></name> <name><surname>Zheng</surname> <given-names>X.</given-names></name> <name><surname>Geng</surname> <given-names>R.</given-names></name> <name><surname>Sham</surname> <given-names>C.-W.</given-names></name></person-group> (<year>2022</year>). <article-title>Uncertainty teacher with dense focal loss for semi-supervised medical image segmentation</article-title>. <source>Comput. Biol. Med</source>. <volume>149</volume>:<fpage>106034</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2022.106034</pub-id><pub-id pub-id-type="pmid">36058068</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dosovitskiy</surname> <given-names>A.</given-names></name> <name><surname>Beyer</surname> <given-names>L.</given-names></name> <name><surname>Kolesnikov</surname> <given-names>A.</given-names></name> <name><surname>Weissenborn</surname> <given-names>D.</given-names></name> <name><surname>Zhai</surname> <given-names>X.</given-names></name> <name><surname>Unterthiner</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>An image is worth 16x16 words: transformers for image recognition at scale</article-title>. <source>arXiv preprint arXiv:2010.11929</source>.</citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Edupuganti</surname> <given-names>M.</given-names></name> <name><surname>Rathikarani</surname> <given-names>V.</given-names></name> <name><surname>Chaduvula</surname> <given-names>K.</given-names></name></person-group> (<year>2024</year>). <article-title>Classification of heart diseases using fusion based learning approach</article-title>. <source>Int. J. Intell. Syst. Applic. Eng</source>. <volume>12</volume>, <fpage>570</fpage>&#x02013;<lpage>580</lpage>.</citation>
</ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Feng</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>B.</given-names></name> <name><surname>Xiao</surname> <given-names>L.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Gegen</surname> <given-names>T.</given-names></name> <name><surname>Chen</surname> <given-names>Z.</given-names></name></person-group> (<year>2024</year>). <article-title>&#x0201C;Enhancing medical imaging with gans synthesizing realistic images from limited data,&#x0201D;</article-title> in <source>2024 IEEE 4th International Conference on Electronic Technology, Communication and Information (ICETCI)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>1192</fpage>&#x02013;<lpage>1197</lpage>. <pub-id pub-id-type="doi">10.1109/ICETCI61221.2024.10594540</pub-id><pub-id pub-id-type="pmid">38067200</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>Q.</given-names></name> <name><surname>Qian</surname> <given-names>X.</given-names></name> <name><surname>Xu</surname> <given-names>H.</given-names></name> <name><surname>Wu</surname> <given-names>K.</given-names></name> <name><surname>Meng</surname> <given-names>L.</given-names></name> <name><surname>Qiu</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>DM-CNN: dynamic multi-scale convolutional neural network with uncertainty quantification for medical image classification</article-title>. <source>Comput. Biol. Med</source>. <volume>168</volume>:<fpage>107758</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2023.107758</pub-id><pub-id pub-id-type="pmid">38042102</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Xie</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Doll&#x00027;ar</surname> <given-names>P.</given-names></name> <name><surname>Girshick</surname> <given-names>R. B.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Masked autoencoders are scalable vision learners,&#x0201D;</article-title> in <source>2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>, 15979&#x02013;15988. <pub-id pub-id-type="doi">10.1109/CVPR52688.2022.01553</pub-id><pub-id pub-id-type="pmid">38715952</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huix</surname> <given-names>J. P.</given-names></name> <name><surname>Ganeshan</surname> <given-names>A. R.</given-names></name> <name><surname>Haslum</surname> <given-names>J. F.</given-names></name> <name><surname>S&#x000F6;derberg</surname> <given-names>M.</given-names></name> <name><surname>Matsoukas</surname> <given-names>C.</given-names></name> <name><surname>Smith</surname> <given-names>K.</given-names></name></person-group> (<year>2024</year>). <article-title>&#x0201C;Are natural domain foundation models useful for medical image classification?&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision</source>, 7634&#x02013;7643. <pub-id pub-id-type="doi">10.1109/WACV57701.2024.00746</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huo</surname> <given-names>J.</given-names></name> <name><surname>Yan</surname> <given-names>Y.</given-names></name> <name><surname>Zheng</surname> <given-names>X.</given-names></name> <name><surname>Lyu</surname> <given-names>Y.</given-names></name> <name><surname>Zou</surname> <given-names>X.</given-names></name> <name><surname>Wei</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2025</year>). <article-title>Mmunlearner: Reformulating multimodal machine unlearning in the era of multimodal large language models</article-title>. <source>arXiv preprint arXiv:2502.11051</source>.</citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jiao</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Mabu</surname> <given-names>S.</given-names></name> <name><surname>Xia</surname> <given-names>H.</given-names></name> <name><surname>Hara</surname> <given-names>T.</given-names></name></person-group> (<year>2024</year>). <article-title>&#x0201C;Multi-modal contrastive learning for medical image classification with limited training data,&#x0201D;</article-title> in <source>2024 International Conference on Machine Learning and Applications (ICMLA)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>1083</fpage>&#x02013;<lpage>1088</lpage>. <pub-id pub-id-type="doi">10.1109/ICMLA61862.2024.00164</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Karagoz</surname> <given-names>M. A.</given-names></name> <name><surname>Nalbantoglu</surname> <given-names>O. U.</given-names></name></person-group> (<year>2024</year>). <article-title>A self-supervised learning model based on variational autoencoder for limited-sample mammogram classification</article-title>. <source>Appl. Intell</source>. <volume>54</volume>, <fpage>3448</fpage>&#x02013;<lpage>3463</lpage>. <pub-id pub-id-type="doi">10.1007/s10489-024-05358-5</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kumar</surname> <given-names>Y.</given-names></name> <name><surname>Marttinen</surname> <given-names>P.</given-names></name></person-group> (<year>2024</year>). <article-title>&#x0201C;Improving medical multi-modal contrastive learning with expert annotations,&#x0201D;</article-title> in <source>European Conference on Computer Vision</source> (<publisher-loc>Springer</publisher-loc>), <fpage>468</fpage>&#x02013;<lpage>486</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-031-72661-3_27</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>M.</given-names></name> <name><surname>Meng</surname> <given-names>M.</given-names></name> <name><surname>Fulham</surname> <given-names>M.</given-names></name> <name><surname>Feng</surname> <given-names>D. D.</given-names></name> <name><surname>Bi</surname> <given-names>L.</given-names></name> <name><surname>Kim</surname> <given-names>J.</given-names></name></person-group> (<year>2025</year>). <article-title>Enhancing medical vision-language contrastive learning via inter-matching relation modelling</article-title>. <source>IEEE Trans. Med. Imag</source>. <volume>44</volume>, <fpage>2463</fpage>&#x02013;<lpage>2476</lpage>. <pub-id pub-id-type="doi">10.1109/TMI.2025.3534436</pub-id><pub-id pub-id-type="pmid">40031323</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Q.</given-names></name> <name><surname>Qiu</surname> <given-names>C.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Gu</surname> <given-names>J.</given-names></name> <name><surname>Luo</surname> <given-names>D.</given-names></name></person-group> (<year>2025</year>). <article-title>Decoupled contrastive learning for multilingual multimodal medical pre-trained model</article-title>. <source>Neurocomputing</source> <volume>633</volume>:<fpage>129809</fpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2025.129809</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liao</surname> <given-names>C.</given-names></name> <name><surname>Zheng</surname> <given-names>X.</given-names></name> <name><surname>Lyu</surname> <given-names>Y.</given-names></name> <name><surname>Xue</surname> <given-names>H.</given-names></name> <name><surname>Cao</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2025</year>). <article-title>Memorysam: memorize modalities and semantics with segment anything model 2 for multi-modal semantic segmentation</article-title>. <source>arXiv preprint arXiv:2503.06700</source>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Cen</surname> <given-names>X.</given-names></name> <name><surname>Yi</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>F.-,a.</given-names></name> <name><surname>Ding</surname> <given-names>J.</given-names></name> <name><surname>Cheng</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2025</year>). <article-title>Challenges in AI-driven biomedical multimodal data fusion and analysis</article-title>. <source>Genom. Prot. Bioinform</source>. <volume>23</volume>:<fpage>qzaf011</fpage>. <pub-id pub-id-type="doi">10.1093/gpbjnl/qzaf011</pub-id><pub-id pub-id-type="pmid">40036568</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>K.</given-names></name> <name><surname>Zhu</surname> <given-names>W.</given-names></name> <name><surname>Shen</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Razavian</surname> <given-names>N.</given-names></name> <name><surname>Geras</surname> <given-names>K. J.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>&#x0201C;Multiple instance learning via iterative self-paced supervised contrastive learning,&#x0201D;</article-title> in <source>2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>, 3355&#x02013;3365. <pub-id pub-id-type="doi">10.1109/CVPR52729.2023.00327</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Yue</surname> <given-names>W.</given-names></name></person-group> (<year>2024</year>). <article-title>An efficient medical image classification network based on multi-branch cnn, token grouping transformer and mixer MLP</article-title>. <source>Appl. Soft Comput</source>. <volume>153</volume>:<fpage>111323</fpage>. <pub-id pub-id-type="doi">10.1016/j.asoc.2024.111323</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lyu</surname> <given-names>Y.</given-names></name> <name><surname>Zheng</surname> <given-names>X.</given-names></name> <name><surname>Kim</surname> <given-names>D.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name></person-group> (<year>2024</year>). <article-title>Omnibind: teach to build unequal-scale modality interaction for omni-bind of all</article-title>. <source>arXiv preprint arXiv:2405.16108</source>.</citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Manzari</surname> <given-names>O. N.</given-names></name> <name><surname>Ahmadabadi</surname> <given-names>H.</given-names></name> <name><surname>Kashiani</surname> <given-names>H.</given-names></name> <name><surname>Shokouhi</surname> <given-names>S. B.</given-names></name> <name><surname>Ayatollahi</surname> <given-names>A.</given-names></name></person-group> (<year>2023</year>). <article-title>Medvit: a robust vision transformer for generalized medical image classification</article-title>. <source>Comput. Biol. Med</source>. <volume>157</volume>:<fpage>106791</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2023.106791</pub-id><pub-id pub-id-type="pmid">36958234</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mao</surname> <given-names>J.</given-names></name> <name><surname>Guo</surname> <given-names>S.</given-names></name> <name><surname>Yin</surname> <given-names>X.</given-names></name> <name><surname>Chang</surname> <given-names>Y.</given-names></name> <name><surname>Nie</surname> <given-names>B.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name></person-group> (<year>2025</year>). <article-title>Medical supervised masked autoencoder: crafting a better masking strategy and efficient fine-tuning schedule for medical image classification</article-title>. <source>Appl. Soft Comput</source>. <volume>169</volume>:<fpage>112536</fpage>. <pub-id pub-id-type="doi">10.1016/j.asoc.2024.112536</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maree</surname> <given-names>M.</given-names></name> <name><surname>Zanoon</surname> <given-names>T.</given-names></name> <name><surname>Dababat</surname> <given-names>A.</given-names></name> <name><surname>Awwad</surname> <given-names>M.</given-names></name></person-group> (<year>2024</year>). <article-title>Constructing a hybrid activation and parameter-fusion based cnn medical image classifier</article-title>. <source>Int. J. Inf. Technol</source>. <volume>16</volume>, <fpage>3265</fpage>&#x02013;<lpage>3272</lpage>. <pub-id pub-id-type="doi">10.1007/s41870-024-01798-x</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>MeenaPrakash</surname> <given-names>R.</given-names></name> <name><surname>Kamali</surname> <given-names>B.</given-names></name> <name><surname>Vimala</surname> <given-names>M.</given-names></name> <name><surname>Madhuvandhana</surname> <given-names>K.</given-names></name> <name><surname>Krishnaleela</surname> <given-names>P.</given-names></name></person-group> (<year>2025</year>). <article-title>&#x0201C;A densenet-enhanced gan model for classification of medical images into original and fake,&#x0201D;</article-title> in <source>2025 International Conference on Multi-Agent Systems for Collaborative Intelligence (ICMSCI)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>1540</fpage>&#x02013;<lpage>1545</lpage>. <pub-id pub-id-type="doi">10.1109/ICMSCI62561.2025.10894026</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mei</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Robson</surname> <given-names>P. M.</given-names></name> <name><surname>Marinelli</surname> <given-names>B.</given-names></name> <name><surname>Huang</surname> <given-names>M.</given-names></name> <name><surname>Doshi</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Radimagenet: an open radiologic deep learning research dataset for effective transfer learning</article-title>. <source>Radiology</source> <volume>4</volume>:<fpage>e210315</fpage>. <pub-id pub-id-type="doi">10.1148/ryai.210315</pub-id><pub-id pub-id-type="pmid">36204533</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Miller</surname> <given-names>R. J.</given-names></name> <name><surname>Bednarski</surname> <given-names>B. P.</given-names></name> <name><surname>Pieszko</surname> <given-names>K.</given-names></name> <name><surname>Kwiecinski</surname> <given-names>J.</given-names></name> <name><surname>Williams</surname> <given-names>M. C.</given-names></name> <name><surname>Shanbhag</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Clinical phenotypes among patients with normal cardiac perfusion using unsupervised learning: a retrospective observational study</article-title>. <source>EBioMedicine</source> 99. <pub-id pub-id-type="doi">10.1016/j.ebiom.2023.104930</pub-id><pub-id pub-id-type="pmid">38168587</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mishra</surname> <given-names>N. K.</given-names></name> <name><surname>Singh</surname> <given-names>P.</given-names></name> <name><surname>Gupta</surname> <given-names>A.</given-names></name> <name><surname>Joshi</surname> <given-names>S. D.</given-names></name></person-group> (<year>2025</year>). <article-title>Pp-cnn: probabilistic pooling cnn for enhanced image classification</article-title>. <source>Neur. Comput. Applic</source>. <volume>37</volume>, <fpage>4345</fpage>&#x02013;<lpage>4361</lpage>. <pub-id pub-id-type="doi">10.1007/s00521-024-10862-3</pub-id><pub-id pub-id-type="pmid">37362578</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nguyen</surname> <given-names>H.</given-names></name> <name><surname>Nguyen</surname> <given-names>H.</given-names></name> <name><surname>Chang</surname> <given-names>M.</given-names></name> <name><surname>Pham</surname> <given-names>H.</given-names></name> <name><surname>Narayanan</surname> <given-names>S.</given-names></name> <name><surname>Pazzani</surname> <given-names>M.</given-names></name></person-group> (<year>2024</year>). <article-title>&#x0201C;Conpro: learning severity representation for medical images using contrastive learning and preference optimization,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>, 5105&#x02013;5112. <pub-id pub-id-type="doi">10.1109/CVPRW63382.2024.00517</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oquab</surname> <given-names>M.</given-names></name> <name><surname>Darcet</surname> <given-names>T.</given-names></name> <name><surname>Moutakanni</surname> <given-names>T.</given-names></name> <name><surname>Vo</surname> <given-names>H.</given-names></name> <name><surname>Szafraniec</surname> <given-names>M.</given-names></name> <name><surname>Khalidov</surname> <given-names>V.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Dinov2: learning robust visual features without supervision</article-title>. <source>arXiv preprint arXiv:2304.07193</source>.</citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oyelade</surname> <given-names>O. N.</given-names></name> <name><surname>Irunokhai</surname> <given-names>E. A.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name></person-group> (<year>2024</year>). <article-title>A twin convolutional neural network with hybrid binary optimizer for multimodal breast cancer digital image classification</article-title>. <source>Sci. Rep</source>. <volume>14</volume>:<fpage>692</fpage>. <pub-id pub-id-type="doi">10.1038/s41598-024-51329-8</pub-id><pub-id pub-id-type="pmid">38184742</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pacal</surname> <given-names>I.</given-names></name> <name><surname>Ozdemir</surname> <given-names>B.</given-names></name> <name><surname>Zeynalov</surname> <given-names>J.</given-names></name> <name><surname>Gasimov</surname> <given-names>H.</given-names></name> <name><surname>Pacal</surname> <given-names>N.</given-names></name></person-group> (<year>2025</year>). <article-title>A novel cnn-vit-based deep learning model for early skin cancer diagnosis</article-title>. <source>Biomed. Signal Process. Control</source> <volume>104</volume>:<fpage>107627</fpage>. <pub-id pub-id-type="doi">10.1016/j.bspc.2025.107627</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pan</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Yao</surname> <given-names>T.</given-names></name> <name><surname>Ngo</surname> <given-names>C.-W.</given-names></name> <name><surname>Mei</surname> <given-names>T.</given-names></name></person-group> (<year>2025</year>). <article-title>Stream-VIT: learning streamlined convolutions in vision transformer</article-title>. <source>IEEE Trans. Multim</source>. <volume>44</volume>, <fpage>3755</fpage>&#x02013;<lpage>3765</lpage>. <pub-id pub-id-type="doi">10.1109/TMM.2025.3535321</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peng</surname> <given-names>Z.</given-names></name> <name><surname>Dong</surname> <given-names>L.</given-names></name> <name><surname>Bao</surname> <given-names>H.</given-names></name> <name><surname>Ye</surname> <given-names>Q.</given-names></name> <name><surname>Wei</surname> <given-names>F.</given-names></name></person-group> (<year>2022</year>). <article-title>Beit v2: masked image modeling with vector-quantized visual tokenizers</article-title>. <source>arXiv preprint arXiv:2208.06366</source>.</citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pham</surname> <given-names>C.</given-names></name> <name><surname>Caicedo</surname> <given-names>J. C.</given-names></name> <name><surname>Plummer</surname> <given-names>B. A.</given-names></name></person-group> (<year>2025</year>). <article-title>Cha-maevit: unifying channel-aware masked autoencoders and multi-channel vision transformers for improved cross-channel learning</article-title>. <source>arXiv preprint arXiv:2503.19331</source>.</citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Prabhakar</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Shit</surname> <given-names>S.</given-names></name> <name><surname>Wiestler</surname> <given-names>B.</given-names></name> <name><surname>Menze</surname> <given-names>B. H.</given-names></name></person-group> (<year>2023</year>). <article-title>Vit-ae&#x0002B;&#x0002B;: improving vision transformer autoencoder for self-supervised medical image representations</article-title>. <source>ArXiv, abs/2301.07382</source>.</citation>
</ref>
<ref id="B38">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Radford</surname> <given-names>A.</given-names></name> <name><surname>Kim</surname> <given-names>J. W.</given-names></name> <name><surname>Hallacy</surname> <given-names>C.</given-names></name> <name><surname>Ramesh</surname> <given-names>A.</given-names></name> <name><surname>Goh</surname> <given-names>G.</given-names></name> <name><surname>Agarwal</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>&#x0201C;Learning transferable visual models from natural language supervision,&#x0201D;</article-title> in <source>International Conference on Machine Learning</source> (<publisher-loc>PMLR</publisher-loc>), <fpage>8748</fpage>&#x02013;<lpage>8763</lpage>.</citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>Y.</given-names></name> <name><surname>Song</surname> <given-names>A.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Liao</surname> <given-names>Z.</given-names></name></person-group> (<year>2025</year>). <article-title>Multiple teachers are beneficial: a lightweight and noise-resistant student model for point-of-care imaging classification</article-title>. <source>Exp. Syst. Applic</source>. <volume>275</volume>:<fpage>127145</fpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2025.127145</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Takahashi</surname> <given-names>S.</given-names></name> <name><surname>Sakaguchi</surname> <given-names>Y.</given-names></name> <name><surname>Kouno</surname> <given-names>N.</given-names></name> <name><surname>Takasawa</surname> <given-names>K.</given-names></name> <name><surname>Ishizu</surname> <given-names>K.</given-names></name> <name><surname>Akagi</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Comparison of vision transformers and convolutional neural networks in medical image analysis: a systematic review</article-title>. <source>J. Med. Syst</source>. <volume>48</volume>:<fpage>84</fpage>. <pub-id pub-id-type="doi">10.1007/s10916-024-02105-8</pub-id><pub-id pub-id-type="pmid">39264388</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Touvron</surname> <given-names>H.</given-names></name> <name><surname>Cord</surname> <given-names>M.</given-names></name> <name><surname>J&#x000E9;gou</surname> <given-names>H.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Deit III: revenge of the vit,&#x0201D;</article-title> in <source>European Conference on Computer Vision</source> (<publisher-loc>Springer</publisher-loc>), <fpage>516</fpage>&#x02013;<lpage>533</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-031-20053-3_30</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Yan</surname> <given-names>X.</given-names></name> <name><surname>Xiao</surname> <given-names>M.</given-names></name> <name><surname>Gao</surname> <given-names>M.</given-names></name></person-group> (<year>2024</year>). <article-title>&#x0201C;Breast cancer image classification method based on deep transfer learning,&#x0201D;</article-title> in <source>Proceedings of the International Conference on Image Processing, Machine Learning and Pattern Recognition</source>, 190&#x02013;197. <pub-id pub-id-type="doi">10.1145/3700906.3700937</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Xu</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>D.</given-names></name> <name><surname>Tam</surname> <given-names>L.</given-names></name> <name><surname>Roth</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>D.</given-names></name></person-group> (<year>2024</year>). <article-title>&#x0201C;Learning quality labels for robust image classification,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision</source>, 1103&#x02013;1112. <pub-id pub-id-type="doi">10.1109/WACV57701.2024.00114</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xing</surname> <given-names>X.</given-names></name> <name><surname>Liang</surname> <given-names>G.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Jacobs</surname> <given-names>N.</given-names></name> <name><surname>Lin</surname> <given-names>A.-L.</given-names></name></person-group> (<year>2023</year>). <article-title>Self-supervised learning application on covid-19 chest x-ray image classification using masked autoencoder</article-title>. <source>Bioengineering</source> <volume>10</volume>:<fpage>901</fpage>. <pub-id pub-id-type="doi">10.3390/bioengineering10080901</pub-id><pub-id pub-id-type="pmid">37627786</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>X.</given-names></name> <name><surname>Wong</surname> <given-names>S. T.</given-names></name></person-group> (<year>2025</year>). <article-title>Contrastive learning in brain imaging</article-title>. <source>Computer. Med. Imag. Graph</source>. <volume>121</volume>:<fpage>102500</fpage>. <pub-id pub-id-type="doi">10.1016/j.compmedimag.2025.102500</pub-id><pub-id pub-id-type="pmid">39889467</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yan</surname> <given-names>Y.</given-names></name> <name><surname>Su</surname> <given-names>J.</given-names></name> <name><surname>He</surname> <given-names>J.</given-names></name> <name><surname>Fu</surname> <given-names>F.</given-names></name> <name><surname>Zheng</surname> <given-names>X.</given-names></name> <name><surname>Lyu</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>A survey of mathematical reasoning in the era of multimodal large language model: benchmark, method and challenges</article-title>. <source>arXiv preprint arXiv:2412.11936</source>.</citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Shi</surname> <given-names>R.</given-names></name> <name><surname>Ni</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). &#x0201C;Medmnist classification decathlon: a lightweight automl benchmark for medical image analysis,&#x0201D; analysis,&#x0201D; in <italic>2021 IEEE 18th International Symposium on Biomedical Imaging</italic> 191&#x02013;195. <pub-id pub-id-type="doi">10.1109/ISBI48211.2021.9434062</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Shi</surname> <given-names>R.</given-names></name> <name><surname>Wei</surname> <given-names>D.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Zhao</surname> <given-names>L.</given-names></name> <name><surname>Ke</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification</article-title>. <source>Sci. Data</source> <volume>10</volume>:<fpage>41</fpage>. <pub-id pub-id-type="doi">10.1038/s41597-022-01721-8</pub-id><pub-id pub-id-type="pmid">36658144</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Luo</surname> <given-names>X.</given-names></name> <name><surname>Lu</surname> <given-names>Z.</given-names></name> <name><surname>Shen</surname> <given-names>L.</given-names></name></person-group> (<year>2025</year>). <article-title>Medkan: an advanced kolmogorov-arnold network for medical image classification</article-title>. <source>arXiv preprint arXiv:2502.18416</source>.</citation>
</ref>
<ref id="B50">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>D.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Confidence-aware multi-teacher knowledge distillation,&#x0201D;</article-title> in <source>ICASSP 2022&#x02013;2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>4498</fpage>&#x02013;<lpage>4502</lpage>. <pub-id pub-id-type="doi">10.1109/ICASSP43922.2022.9747534</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>W.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Zheng</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name></person-group> (<year>2024</year>). <article-title>Goodsam: bridging domain and capacity gaps via segment anything model for distortion -aware panoramic semantic segmentation</article-title>. <source>arXiv preprint arXiv:2403.16370.</source></citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Xiao</surname> <given-names>Z.</given-names></name> <name><surname>Wu</surname> <given-names>X.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Zhao</surname> <given-names>J.</given-names></name> <name><surname>Hu</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Pyramid pixel context adaption network for medical image classification with supervised contrastive learning</article-title>. <source>IEEE Trans. Neural Netw. Learn. Syst</source>. <volume>36</volume>, <fpage>6802</fpage>&#x02013;<lpage>6815</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2024.3399164</pub-id><pub-id pub-id-type="pmid">38829749</pub-id></citation></ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>J.</given-names></name> <name><surname>Teng</surname> <given-names>F.</given-names></name> <name><surname>Luo</surname> <given-names>K.</given-names></name> <name><surname>Zhao</surname> <given-names>G.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Zheng</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2025</year>). <article-title>Unveiling the potential of segment anything model 2 for RGB-thermal semantic segmentation with language guidance</article-title>. <source>arXiv preprint arXiv:2503.02581</source>.</citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>X.</given-names></name> <name><surname>Fu</surname> <given-names>C.</given-names></name> <name><surname>Xie</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Sham</surname> <given-names>C.-W.</given-names></name></person-group> (<year>2022</year>). <article-title>Uncertainty-aware deep co-training for semi-supervised medical image segmentation</article-title>. <source>Comput. Biol. Med</source>. <volume>149</volume>:<fpage>106051</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2022.106051</pub-id><pub-id pub-id-type="pmid">36055155</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>X.</given-names></name> <name><surname>Weng</surname> <given-names>Z.</given-names></name> <name><surname>Lyu</surname> <given-names>Y.</given-names></name> <name><surname>Jiang</surname> <given-names>L.</given-names></name> <name><surname>Xue</surname> <given-names>H.</given-names></name> <name><surname>Ren</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2025</year>). <article-title>Retrieval augmented generation and understanding in vision: a survey and new outlook</article-title>. <source>arXiv preprint arXiv:2503.18016</source>.</citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhong</surname> <given-names>D.</given-names></name> <name><surname>Zheng</surname> <given-names>X.</given-names></name> <name><surname>Liao</surname> <given-names>C.</given-names></name> <name><surname>Lyu</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Wu</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2025</year>). <article-title>Omnisam: omnidirectional segment anything model for UDA in panoramic semantic segmentation</article-title>. <source>arXiv preprint arXiv:2503.07098</source>.</citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>C.</given-names></name> <name><surname>Xiao</surname> <given-names>B.</given-names></name> <name><surname>Shi</surname> <given-names>L.</given-names></name> <name><surname>Xu</surname> <given-names>S.</given-names></name> <name><surname>Zheng</surname> <given-names>X.</given-names></name></person-group> (<year>2024</year>). <article-title>Customize segment anything model for multi-modal semantic segmentation with mixture of lora experts</article-title>. <source>arXiv preprint arXiv:2412.04220</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2412.04220</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Yip</surname> <given-names>R.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Cai</surname> <given-names>Q.</given-names></name> <name><surname>Sun</surname> <given-names>Q.</given-names></name> <name><surname>Li</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Radiologic features of nodules attached to the mediastinal or diaphragmatic pleura at low-dose CT for lung cancer screening</article-title>. <source>Radiology</source> <volume>310</volume>:<fpage>e231219</fpage>.</citation>
</ref>
</ref-list>
</back>
</article>