<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2025.1627876</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>An artificial intelligence model for early-stage breast cancer classification from histopathological biopsy images</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Chaudhary</surname>
<given-names>Neil</given-names>
</name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/3060042/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Dhunny</surname>
<given-names>A. Z.</given-names>
</name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/3185040/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Pangea Society</institution>, <addr-line>New Delhi</addr-line>, <country>India</country></aff>
<aff id="aff2"><sup>2</sup><institution>Cyber Analytics</institution>, <addr-line>Port Louis</addr-line>, <country>Mauritius</country></aff>
<author-notes>
<fn fn-type="edited-by" id="fn0001">
<p>Edited by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1886575/overview">Farah Kidwai-Khan</ext-link>, Yale University, United States</p>
</fn>
<fn fn-type="edited-by" id="fn0002">
<p>Reviewed by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1069324/overview">Marco Cavaco</ext-link>, Gulbenkian Institute of Science (IGC), Portugal</p>
<p><ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2946302/overview">Abdelmoniem Helmy</ext-link>, Cairo University, Egypt</p>
</fn>
<corresp id="c001">&#x002A;Correspondence: Neil Chaudhary, <email>neil.c1107@gmail.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>12</day>
<month>09</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>8</volume>
<elocation-id>1627876</elocation-id>
<history>
<date date-type="received">
<day>13</day>
<month>05</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>08</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2025 Chaudhary and Dhunny.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Chaudhary and Dhunny</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Accurate identification of breast cancer subtypes is essential for guiding treatment decisions and improving patient outcomes. In current clinical practice, determining histological subtypes often requires additional invasive procedures, delaying treatment initiation. This study proposes a deep learning-based model built on a DenseNet121 backbone with a multi-scale feature fusion strategy, designed to classify breast cancer from histopathological biopsy images. Trained and evaluated on the publicly available BreaKHis dataset using 5-fold cross-validation, the model achieved a binary classification accuracy of 97.1%, and subtype classification accuracies of 93.8% for benign tumors and 92.0% for malignant tumors. These results demonstrate the model&#x2019;s ability to capture morphological cues at multiple levels of abstraction and highlight its potential as a diagnostic support tool in digital pathology workflows.</p>
</abstract>
<kwd-group>
<kwd>breast cancer classification</kwd>
<kwd>histopathological images</kwd>
<kwd>deep learning</kwd>
<kwd>DenseNet121</kwd>
<kwd>multi-scale feature fusion</kwd>
<kwd>convolutional neural networks (CNNs)</kwd>
<kwd>subtype detection</kwd>
<kwd>medical image analysis</kwd>
</kwd-group>
<counts>
<fig-count count="21"/>
<table-count count="14"/>
<equation-count count="0"/>
<ref-count count="78"/>
<page-count count="26"/>
<word-count count="15373"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Medicine and Public Health</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="sec1">
<label>1</label>
<title>Introduction</title>
<p>Breast cancer remains one of the most prevalent cancers globally, representing a significant health challenge for women across all age groups. According to the World Health Organization (WHO), over 2.3 million women were diagnosed with breast cancer in 2020, making it the most diagnosed cancer worldwide and the leading cause of cancer-related deaths among women (<xref ref-type="bibr" rid="ref48">WHO, 2021</xref>). The incidence of breast cancer is rising by approximately 3% per year, with higher mortality rates observed in lower-income countries due to limited access to early screening and treatment. In wealthier nations, one in 12 women is diagnosed with breast cancer, whereas in lower-income countries, the rate is one in 27. More concerning is the disparity in mortality: one in 48 women dies from breast cancer in low-income countries compared to one in 71 in high-income countries [<xref ref-type="bibr" rid="ref9026">World Health Organization (WHO), 2022</xref>]. In sub-Saharan Africa, breast cancer now has the highest mortality rate among all cancers affecting women, surpassing cervical cancer. It accounts for 20% of cancer-related deaths in women, with incidence rates varying by region: 30.4 per 100,000 women in Eastern Africa, 26.8 in Central Africa, 38.6 in Western Africa, and 38.9 in Southern Africa. Despite lower incidence rates than in developed countries, the mortality-to-incidence ratio remains alarmingly high at 0.55 in Central Africa, compared to just 0.16 in the United States (<xref ref-type="bibr" rid="ref9016">GLOBOCAN, 2020</xref>).</p>
<p>Early detection has been identified as a critical factor in improving survival rates, with studies showing that early-stage breast cancer has a 90% 5-year survival rate compared to late-stage diagnoses, which can drop to 27% (<xref ref-type="bibr" rid="ref38">Siegel et al., 2022</xref>). This significant effort is to enhance diagnostic techniques to detect breast cancer earlier and more accurately. Over the past decade, there have been remarkable advances in breast cancer detection, particularly with the integration of advanced imaging techniques and machine learning. Traditional imaging modalities, such as mammography and ultrasound, have been instrumental in identifying breast lesions; however, their sensitivity can decrease in women with dense breast tissue (<xref ref-type="bibr" rid="ref17">Kolb et al., 2002</xref>). Magnetic resonance imaging (MRI) has emerged as a superior modality for detecting early-stage breast cancer due to its high spatial resolution and ability to differentiate between benign and malignant lesions (<xref ref-type="bibr" rid="ref19">Kuhl et al., 2005</xref>). Despite its advantages, retention remains complex and prone to variability among radiologists, emphasizing the need for more standardized and accurate approaches. Artificial intelligence (AI) models present a transformative solution to these challenges, especially in resource-limited settings. By leveraging AI to analyze biopsy images, diagnostic accuracy can be improved while reducing the workload of overburdened pathologists. AI-powered systems have demonstrated sensitivities of up to 94.5% in detecting malignant tumors in histopathological images, reducing false-negative rates significantly (<xref ref-type="bibr" rid="ref9013">Esteva et al., 2021</xref>). Additionally, AI-assisted diagnostics have the potential to cut down diagnosis time from weeks to just hours, allowing for earlier treatment initiation and improved patient outcomes. This is particularly beneficial in low-resource settings where a shortage of trained pathologists delays cancer detection (<xref ref-type="bibr" rid="ref27">McKinney et al., 2020</xref>). By integrating AI into breast cancer diagnostics, healthcare systems in developing countries can bridge the gap in early detection, reduce misdiagnosis rates, and ultimately lower breast cancer mortality. Scalable AI solutions, combined with improved screening programs and public awareness efforts, have the potential to significantly enhance cancer care worldwide.</p>
<p>Despite significant progress in imaging modalities and AI-driven diagnostic tools, most existing models focus primarily on binary classification&#x2014;differentiating benign from malignant tumors&#x2014;without addressing the more nuanced and clinically challenging task of histological subtype classification. This limitation hinders their real-world applicability, particularly in cases where treatment decisions depend on precise subtype identification. These gaps underscore the need for a more robust, subtype-aware AI framework trained on whole-slide biopsy images to improve diagnostic granularity and clinical utility.</p>
<sec id="sec2">
<label>1.1</label>
<title>Brief review of artificial intelligence in breast cancer</title>
<p>This section provides a brief literature review of what technology has been used for the detection of breast cancer over the past few years for breast cancer detection. And this is going to justify our innovative model and method in the further sections.</p>
<p>AI has revolutionized the healthcare landscape, offering transformative capabilities in automating and enhancing diagnostic processes. In the context of breast cancer detection, AI models have demonstrated superior performance in analyzing complex MRI datasets, identifying patterns that may elude human experts (<xref ref-type="bibr" rid="ref22">Lehman et al., 2019</xref>). For instance, deep learning convolutional neural networks (CNNs) and attention mechanisms have been applied to segment breast tissues, classify lesions, and predict malignancy with high accuracy (<xref ref-type="bibr" rid="ref24">Litjens et al., 2017</xref>). These advancements not only reduce errors but also minimize the workload of radiologists, making AI a valuable tool in clinical settings. The integration of AI into breast cancer detection workflows has also been shown to address the challenges of class imbalance and variability in imaging quality. Many studies have employed transfer learning and ensemble techniques to overcome these challenges, achieving higher sensitivity and specificity than traditional methods (<xref ref-type="bibr" rid="ref27">McKinney et al., 2020</xref>). Moreover, AI-driven algorithms can process datasets quickly and consistently, offering significant advantages in population-based breast cancer screening programs (<xref ref-type="bibr" rid="ref35">Rodrigues et al., 2021</xref>).</p>
<p><xref ref-type="bibr" rid="ref9020">Kaymak et al. (2017)</xref> applied neural networks to classify 176 histopathology images, using discrete Haar wavelets for preprocessing. Radial Basis Function Networks outperformed Back Propagation Networks, achieving 70.49% accuracy. <xref ref-type="bibr" rid="ref9014">Fu et al. (2022)</xref> predicted Peripherally Inserted Central Catheter (PICC)-related thrombosis in breast cancer patients using Artificial Neural Network (ANN) and Synthetic Minority Over-sampling Technique (SMOTE) to address class imbalance, outperforming logistic regression (AUC: 0.742 vs. 0.675). <xref ref-type="bibr" rid="ref9010">Punitha et al. (2021)</xref> combined Artificial Immune Systems and Bee Colony optimization for feature selection, improving multilayer perceptron (MLP) performance on the Wisconsin Breast Cancer Dataset (WBCD) dataset for automated diagnosis.</p>
<p><xref ref-type="bibr" rid="ref9017">Sharma et al. (2024)</xref> used the WBCD to build a stacked ensemble framework integrating Decision Tree, AdaBoost, Gaussian Naive Bayes, and MLP classifiers, achieving 97.66% accuracy. Ensemble learning and feature engineering contributed to its robust performance. <xref ref-type="bibr" rid="ref9002">Liu et al. (2024a)</xref> enhanced a Visual Geometry Group 16-layer network (VGG16)-based model with transfer learning and focal loss to handle class imbalance in 7,841 mammograms, yielding a 96.95% accuracy. <xref ref-type="bibr" rid="ref9027">Li et al. (2024b)</xref> proposed a multimodal fusion model that combined ResNet34-extracted MRI features with RNA-seq gene expression data, using transformers and attention mechanisms to predict treatment response in breast cancer patients more accurately than single-modality approaches.</p>
<p><xref ref-type="bibr" rid="ref9006">Munshi et al. (2024)</xref> introduced a novel breast cancer detection method by combining an optimized ensemble learning framework with explainable AI (XAI). Using the Wisconsin Breast Cancer Dataset, which contains 32 numerical features, they employed an ensemble model integrating CNNs with traditional machine learning algorithms like random forest (RF) and support vector machine (SVM). U-NET was used for image-based tasks, and a voting mechanism combined RF and SVM predictions. The model achieved an impressive accuracy of 99.99%, surpassing state-of-the-art methods in precision, recall, and F1-score. The integration of XAI enhanced the interpretability and transparency of model decisions, further improving breast cancer diagnostics.</p>
<p><xref ref-type="bibr" rid="ref9004">Atrey et al. (2024)</xref> developed a multimodal classification system for breast cancer by integrating mammogram (MG) and ultrasound (US) images using deep learning and traditional machine learning techniques. The dataset, sourced from AIIMS, Raipur, India, consisted of 31 patients and 43 MG and 43 US images, which were augmented to 1,032 images. Data preprocessing involved manual region-of-interest (ROI) annotation by radiologists and noise filtering techniques for both image types. The deep learning model, ResNet-18, was used for feature extraction, while SVM classified the fused features. The hybrid ResNet-SVM approach achieved a classification accuracy of 99.22%, significantly outperforming unimodal methods. This study demonstrates the effectiveness of combining deep learning with traditional machine learning for improved breast cancer diagnosis.</p>
<p><xref ref-type="bibr" rid="ref9015">Robinson and Preethi (2024)</xref> developed the SDM-WHO-RNN model, combining fully convolutional networks with RNNs for classifying 9,109 histopathological images from the Breast Cancer Histopathological Image (BreaKHis) and BH datasets. Preprocessing steps included resizing, noise reduction, and stain normalization, while Least Squares Convolutional Encoder-Decoder (LS-CED) was used for tumor segmentation. The system achieved 97.9% accuracy in detecting complex tumor cells. <xref ref-type="bibr" rid="ref50">Zhang et al. (2024)</xref> proposed the Convergent Difference Neural Network (V-CDNN), an ensemble of Convergent Difference Neural Networks optimized with a Neural Dynamic Algorithm and softsign activations. Trained on public datasets, the model reached a perfect 100% classification accuracy, demonstrating both speed and robustness.</p>
<p>In their 2021 study, Meha Desai and Manan Shah compared the effectiveness of MLP and CNN for breast cancer detection. The study used datasets such as BreakHis, WBCD, and WDBC, with images and biomarkers preprocessed through normalization and feature extraction (e.g., Discrete Cosine Transform). While both MLP and CNN were trained to classify breast abnormalities, CNN consistently outperformed MLP in terms of accuracy. CNN achieved ~99.86% accuracy on the BreakHis dataset, while MLP&#x2019;s performance was generally lower. CNN was found to be more reliable for image-based diagnosis, whereas MLP showed limitations for larger datasets.</p>
<p>Further to this, <xref ref-type="bibr" rid="ref9012">Ezzat et al. (2023)</xref> proposed an optimized Bayesian convolutional neural network (OBCNN) for detecting invasive ductal carcinoma (IDC) from histopathology images. This approach integrated ResNet101V2 with Mucinous Carcinoma (MC) for uncertainty estimation. Using the slime mould algorithm to optimize dropout rates, the study fine-tuned the architecture, outperforming other pre-trained networks such as VGG16 and DenseNet121. The model demonstrated significant robustness in diagnosis and generalization. <xref ref-type="bibr" rid="ref9019">Joseph et al. (2022)</xref> proposed a multi-classification approach for breast cancer using handcrafted features, including Hu moments, Haralick textures, and color histograms. These features were combined with deep neural networks trained on the BreakHis dataset. Data augmentation was applied to address overfitting, and a four-layer dense network with softmax activation was employed to enhance classification accuracy.</p>
<p><xref ref-type="bibr" rid="ref9008">Paul et al. (2023)</xref> proposed a novel breast cancer detection system using an SDM-WHO-RNN classifier combined with LS-CED segmentation. Their approach incorporated preprocessing steps, such as noise elimination and image normalization. The LS-CED segmentation localized nuclei, which were then classified based on size and shape. The system demonstrated high accuracy in detecting cancerous cells. <xref ref-type="bibr" rid="ref9022">Taheri and Omranpour (2023)</xref> introduced an ensemble meta-feature generator (EMFSG-Net) for classifying ultrasound images. Leveraging transfer learning with the VGG-16 architecture, the model employed support vector regression (SVR) to create efficient feature spaces. To address issues like overfitting and dead neurons, the leaky-ReLU activation function was incorporated, leading to improved feature representation and classification performance. <xref ref-type="bibr" rid="ref9005">Muduli et al. (2021)</xref> proposed a deep CNN for automated breast cancer classification using mammograms and ultrasound images. The model, consisting of five learnable layers, employed manual cropping for feature extraction and data augmentation to improve generalization. The approach achieved superior results compared to state-of-the-art methods.</p>
<p>Dr. Suvidha Tripathi&#x2019;s research contributes significantly to the field of breast cancer histopathology image classification, particularly through innovative approaches that integrate spatial context and hybrid feature representations. In one study, Tripathi et al. proposed a BiLSTM-based patch modeling framework, which treats histopathology image patches as sequential data to capture spatial continuity&#x2014;an aspect often overlooked in conventional CNNs. This model demonstrated strong performance on the Breast Cancer Histology (Grand Challenge 2018 dataset), highlighting the value of contextual learning in improving classification accuracy. In subsequent work, she introduced a hybrid architecture combining CNNs with Bag-of-Visual-Words (BoVW), effectively integrating handcrafted and deep features to enhance discrimination in limited data settings. This approach outperformed standard deep networks like ResNet and DenseNet on the same dataset, underscoring the benefits of feature selection and hybrid modeling in medical imaging tasks. Together, these studies emphasize alternative paths for improving classification performance&#x2014;either by modeling inter-patch relationships or by enriching feature spaces&#x2014;both of which are highly relevant for advancing automated diagnosis in breast cancer histopathology.</p>
<p>This research is toward an innovative approach for breast cancer detection, including the multiple subclasses of breast cancer, which still confuses medical doctors, and it hence makes patients undergo more advanced invasive tests for a better understanding of the disease. The technique is based on an AI model trained on a sizable batch of data. Section 2 discusses the model in detail, Section 3 shows the Results and Discussion, and the Conclusion is given in Section 4.</p>
</sec>
</sec>
<sec sec-type="methods" id="sec3">
<label>2</label>
<title>Methodology</title>
<p>This methodology section combines an explanation of the different types of breast cancer and their subclasses, along with an advanced AI technique which has been developed to analyse breast cancer results of biopsy. The aim is, instead of making women go again and again through trial and error with treatments or misdiagnosing the cancer, using AI can make all the difference. As mentioned in the literature review section above, AI tech has changed the way the world views medicine now. Therefore, in this study, we have designed a powerful tool for diagnosing the multiple subclasses. This goes toward a humanitarian cause as women are the target here.</p>
<p>Benign and malignant breast cancers differ in their behavior, prognosis, and treatment approach. Benign tumors, such as fibroadenomas or cysts, are non-cancerous growths that do not invade surrounding tissues or spread to other parts of the body. They tend to have well-defined borders, grow slowly, and usually pose little to no health risk. In contrast, malignant tumors are cancerous and have the potential to invade nearby tissues and metastasize to distant organs through the lymphatic system or bloodstream (<xref ref-type="bibr" rid="ref34">Robbins et al., 2010</xref>). Malignant breast cancers, such as IDCs, invasive lobular carcinoma (LCs), etc., show uncontrolled cell growth and can require serious medical intervention in the form of surgery and/or radiation therapy. Early identification of a tumour as benign or malignant is critical to deliver appropriate treatment and improve patient health (<xref ref-type="bibr" rid="ref29">National Cancer Institute, 2021</xref>). In histopathological images, benign and malignant tissues exhibit discernibly unique optical features. Benign tumours present as well-organized structures consisting of cells with uniform shapes, minimal mitotic activity, and intact basement membranes. The stromal and glandular components of benign lesions are usually preserved, showing regular nuclei and minimal pleomorphism. On the other hand, malignant lesions show irregular cellular arrangements, pleomorphic nuclei and frequent mitotic figures (<xref ref-type="bibr" rid="ref34">Robbins et al., 2010</xref>). They often display disrupted basement membranes and an increased frequency of necrotic regions, further highlighting their aggressive nature. AI image analysis of biopsy images will be able to use these intrinsic differences to ensure earlier and more reliable classification into benign and malignant cases (<xref ref-type="bibr" rid="ref33">Rakhlin et al., 2018</xref>).</p>
<p>The visual features in the biopsy images provide key characteristics for distinguishing malignant and benign breast tissues. <xref ref-type="fig" rid="fig1">Figure 1a</xref> (malignant sample) exhibits irregular, disorganized cell structures with pleomorphic nuclei and a dense, fibrous stroma, indicating uncontrolled cancerous growth. The nuclei appear darker and more varied in shape, with increased mitotic activity and loss of structural integrity (<xref ref-type="bibr" rid="ref34">Robbins et al., 2010</xref>). In contrast, <xref ref-type="fig" rid="fig1">Figure 1b</xref> (benign sample) shows well-organized, rounded glandular structures surrounded by normal fibrous tissue, with uniform cell shapes and minimal pleomorphism (<xref ref-type="bibr" rid="ref29">National Cancer Institute, 2021</xref>). Both these images have a magnification of 100x.</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption>
<p>Biopsy image of malignant and benign sample (BreakHis dataset). <bold>(a)</bold> Malignant sample. <bold>(b)</bold> Benign sample.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g001.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">A plain white background with no visible objects or text.</alt-text>
</graphic>
</fig>
<sec id="sec4">
<label>2.1</label>
<title>Dataset</title>
<p>In this study, we utilize the BreaKHis dataset compiled by <xref ref-type="bibr" rid="ref41">Spanhol et al. (2016)</xref>, which contains 9,109 microscopic images of breast tissue captured at four magnification levels: 40&#x00D7;, 100&#x00D7;, 200&#x00D7;, and 400&#x00D7;. These images are categorized into two overarching classes&#x2014;benign and malignant tumors. Within each class, the dataset further identifies four subtypes: for benign tumors, these include adenosis (A), fibroadenoma (F), phyllodes tumor (PT), and tubular adenoma (TA); and for malignant tumors, ductal carcinoma (DC), LC, mucinous carcinoma (MC), and papillary carcinoma (PC). The image filenames in the BreaKHis dataset encode multiple details about each sample, including the biopsy method used, the tumor&#x2019;s classification (benign or malignant), its histological subtype, the patient ID, and the magnification level. For instance, the file named SOB_B_TA-14-4659-40-001.png refers to the first image of a benign tumor classified as TA, obtained from patient 14-4659 at a magnification of 40&#x00D7; using the Segmental Orthogonal Biopsy (SOB) technique (<xref ref-type="bibr" rid="ref41">Spanhol et al., 2016</xref>).</p>
<p>It comprises samples from 82 patients collected at a single institution in Brazil, with no publicly available metadata on age, ethnicity, or receptor status&#x2014;24 with benign tumors and 58 with malignant tumors. Of the 7,909 images used in the final analysis, 5,429 are malignant and 2,480 are benign, distributed across the aforementioned magnification levels. Following preprocessing and class balancing, the revised class distribution is outlined in <xref ref-type="table" rid="tab1">Table 1</xref>.</p>
<table-wrap position="float" id="tab1">
<label>Table 1</label>
<caption>
<p>Dataset image count for each cancer type.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Cancer type</th>
<th align="center" valign="top">Final image count</th>
<th align="left" valign="top">Subtype</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Adenosis</td>
<td align="center" valign="top">444</td>
<td align="left" valign="top">Benign</td>
</tr>
<tr>
<td align="left" valign="top">Fibroadenoma</td>
<td align="center" valign="top">730</td>
<td align="left" valign="top">Benign</td>
</tr>
<tr>
<td align="left" valign="top">Phyllodes tumor</td>
<td align="center" valign="top">453</td>
<td align="left" valign="top">Benign</td>
</tr>
<tr>
<td align="left" valign="top">Tubular adenoma</td>
<td align="center" valign="top">569</td>
<td align="left" valign="top">Benign</td>
</tr>
<tr>
<td align="left" valign="top">Ductal carcinoma</td>
<td align="center" valign="top">797</td>
<td align="left" valign="top">Malignant</td>
</tr>
<tr>
<td align="left" valign="top">Lobular carcinoma</td>
<td align="center" valign="top">626</td>
<td align="left" valign="top">Malignant</td>
</tr>
<tr>
<td align="left" valign="top">Mucinous carcinoma</td>
<td align="center" valign="top">792</td>
<td align="left" valign="top">Malignant</td>
</tr>
<tr>
<td align="left" valign="top">Papillary carcinoma</td>
<td align="center" valign="top">560</td>
<td align="left" valign="top">Malignant</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Benign tumors, by histological standards, lack features associated with malignancy, such as cellular atypia, mitotic activity, invasive growth, or metastatic potential. They typically exhibit localized, slow growth and are considered non-lethal (<xref ref-type="bibr" rid="ref9011">Robbins and Cotran, 2010</xref>). In contrast, malignant tumors are characterized by their capacity to invade neighboring tissues and metastasize to distant sites, often posing life-threatening risks (<xref ref-type="bibr" rid="ref9011">Robbins and Cotran, 2010</xref>). All tissue samples in this dataset were acquired through SOB&#x2014;a surgical procedure also referred to as partial mastectomy or excisional biopsy. Unlike needle-based techniques, SOB yields larger tissue specimens and is generally performed under general anesthesia in a clinical setting (<xref ref-type="bibr" rid="ref2">American Cancer Society, 2019</xref>). The classification of tumors into distinct subtypes is based on their cellular morphology under microscopic examination, which holds significance for both prognosis and treatment planning (<xref ref-type="bibr" rid="ref29">National Cancer Institute, 2021</xref>).</p>
<p>The BreakHis dataset was used for both binary (benign vs. malignant) and multiclass subclass classification. All histopathological images were resized to 128&#x202F;&#x00D7;&#x202F;128 pixels and normalized to a standard RGB range. A stratified 80:20 train-test split was employed to preserve class balance across sets. For cross-validation, a 5-fold split was used with data shuffling to ensure unbiased performance evaluation, and no image was shared across folds.</p>
<p>Prior to splitting, real-time augmentation was applied separately to the benign and malignant image sets, increasing the number of benign samples to 7,440 and malignant samples to 9,385. The augmented dataset, totaling 16,825 images, was then split into training, validation, and test sets using a stratified class-wise approach. For both benign and malignant classes, 20% of the data was reserved for testing, and the remaining 80% was further divided in a 75:25 ratio to form the training and validation sets. This resulted in 10,095 images in the training set and 3,365 images each in the validation and test sets.</p>
<p>To improve generalization and address the risk of overfitting due to limited data, we implemented an extensive on-the-fly augmentation pipeline using TensorFlow&#x2019;s data generators. This included spatial and color transformations to simulate natural histological variation while preserving key tissue features. <xref ref-type="table" rid="tab2">Table 2</xref> summarizes the augmentation parameters used during training.</p>
<table-wrap position="float" id="tab2">
<label>Table 2</label>
<caption>
<p>Augmentation parameters.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Augmentation type</th>
<th align="center" valign="top">Range/value</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Rotation</td>
<td align="center" valign="top">90&#x00B0;, 180&#x00B0;, 270&#x00B0;</td>
</tr>
<tr>
<td align="left" valign="top">Horizontal flip</td>
<td align="center" valign="top">Random</td>
</tr>
<tr>
<td align="left" valign="top">Vertical flip</td>
<td align="center" valign="top">Random</td>
</tr>
<tr>
<td align="left" valign="top">Zoom</td>
<td align="center" valign="top">Up to &#x00B1;10%</td>
</tr>
<tr>
<td align="left" valign="top">Shear</td>
<td align="center" valign="top">&#x00B1;10&#x00B0;</td>
</tr>
<tr>
<td align="left" valign="top">Brightness adjustment</td>
<td align="center" valign="top">&#x00B1;20% range</td>
</tr>
<tr>
<td align="left" valign="top">Contrast adjustment</td>
<td align="center" valign="top">&#x00B1;20% range</td>
</tr>
<tr>
<td align="left" valign="top">Saturation/hue jitter</td>
<td align="center" valign="top">Applied with random factor</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>These augmentations were applied to each batch during training to expose the model to diverse variations without expanding the dataset size. In addition, we used early stopping (patience&#x202F;=&#x202F;5) and ReduceLROnPlateau scheduling to prevent overfitting, along with dropout (0.45), L2 kernel regularization, and batch normalization in each dense layer. Learning curves were generated during each fold to monitor training behavior. The curves showed a minimal gap between training and validation accuracy, indicating stable convergence and effective generalization across classes.</p>
<p>To address class imbalance&#x2014;particularly for underrepresented subtypes, such as PT and LC&#x2014;we applied targeted data augmentation during preprocessing. For most classes, each image was augmented using 90&#x00B0; and 180&#x00B0; rotations along with horizontal flipping, effectively tripling their sample size. However, for class label 4 (DC), which was already sufficiently represented, only minimal augmentation was applied to prevent overrepresentation. This strategy helped balance the dataset and reduce bias during training, without altering the loss function or applying weighted sampling.</p>
</sec>
<sec id="sec5">
<label>2.2</label>
<title>Artificial intelligence model</title>
<p>A CNN is a deep learning architecture tailored for grid-like data such as images, leveraging spatial hierarchies to extract local and global features (<xref ref-type="bibr" rid="ref20">LeCun et al., 1998</xref>). Input images are processed through convolutional layers that apply learnable filters to detect features like edges and textures, generating spatially preserved feature maps (<xref ref-type="bibr" rid="ref18">Krizhevsky et al., 2012</xref>). Non-linear activation functions, typically ReLU (<xref ref-type="bibr" rid="ref10">Glorot et al., 2011</xref>), introduce complexity, while pooling layers reduce spatial dimensions and enhance translational invariance (<xref ref-type="bibr" rid="ref37">Scherer et al., 2010</xref>). Deeper layers capture higher-level abstractions, which are then passed to fully connected layers for classification (<xref ref-type="bibr" rid="ref39">Simonyan and Zisserman, 2014</xref>). Enhancements such as batch normalization, residual connections (<xref ref-type="bibr" rid="ref12">He et al., 2016</xref>), and attention mechanisms have further advanced CNN performance, particularly in medical image analysis (see <xref ref-type="fig" rid="fig2">Figure 2</xref>).</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption>
<p>CNN architecture (<xref ref-type="bibr" rid="ref1">Alzubaidi et al., 2021</xref>).</p>
</caption>
<graphic xlink:href="frai-08-1627876-g002.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Diagram of a neural network processing an image of a dog. The image progresses through layers: Convolution, ReLU, Pooling, and Fully Connected. The output classifies the image as "Dog" or "Not Dog."</alt-text>
</graphic>
</fig>
<sec id="sec6">
<label>2.2.1</label>
<title>Classifying images as benign or malignant</title>
<p>In the first approach, a conventional CNN architecture was developed to perform a binary classification of histopathological breast cancer images into benign and malignant categories. Rather than selecting parameters arbitrarily, several architectural configurations were tested iteratively to arrive at a structure that balanced computational efficiency and classification performance. The final model comprised three convolutional layers with 16, 32, and 16 filters, respectively, employing 4&#x202F;&#x00D7;&#x202F;4 kernels to effectively capture both low- and mid-level spatial features. This configuration aligns with literature that emphasizes the benefit of using multiple convolutional layers of varying depth for feature extraction in histological images (<xref ref-type="bibr" rid="ref41">Spanhol et al., 2016</xref>; <xref ref-type="bibr" rid="ref6">Bayramoglu et al., 2017</xref>). ReLU activation was applied after each convolutional layer to introduce non-linearity, mitigate vanishing gradients, and support efficient learning. MaxPooling layers followed each convolutional block to reduce spatial dimensions and retain salient features (<xref ref-type="bibr" rid="ref33">Rakhlin et al., 2018</xref>).</p>
<p>The output of the convolutional stack was flattened and passed into a fully connected layer with 256 neurons, enabling high-capacity representation before final classification. A sigmoid-activated output layer was used for binary probability prediction, and binary cross-entropy served as the loss function, both standard choices in binary medical image classification tasks (<xref ref-type="bibr" rid="ref9">Gandomkar et al., 2018</xref>; <xref ref-type="bibr" rid="ref25">Liu et al., 2019</xref>). The model was trained using the Adam optimizer (initial learning rate&#x202F;=&#x202F;0.0001) over 20 epochs, leveraging its adaptive learning rate to ensure stable convergence in a data-limited environment (<xref ref-type="bibr" rid="ref4">Bandi et al., 2018</xref>). Despite sound architecture and training configuration, this initial model was limited by two key issues: a substantial class imbalance skewed toward malignant samples, and the failure to account for magnification-level heterogeneity across the dataset (40&#x00D7;, 100&#x00D7;, 200&#x00D7;, and 400&#x00D7;). This led to model overfitting, poor generalization, and relatively low accuracy and F1-scores on the validation set. The architectural layout is depicted in <xref ref-type="fig" rid="fig3">Figure 3</xref> (see <xref ref-type="table" rid="tab3">Table 3</xref>).</p>
<fig position="float" id="fig3">
<label>Figure 3</label>
<caption>
<p>Initial CNN.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g003.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Diagram of a convolutional neural network architecture. It starts with an input layer, followed by layers with 16 and 32 filters, pooling, flattening, a dense layer with 256 neurons, and ends with a sigmoid output. Legend indicates input, convolutional layer, pooling, flatten, dense layer, and output.</alt-text>
</graphic>
</fig>
<table-wrap position="float" id="tab3">
<label>Table 3</label>
<caption>
<p>Hyperparameters for initial binary classification CNN.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Hyperparameter</th>
<th align="center" valign="top">Value</th>
<th align="left" valign="top">Explanation</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="middle">Convolutional filters</td>
<td align="center" valign="middle">16, 32, 16</td>
<td align="left" valign="middle">The number of filters in each convolutional layer. More filters allow for learning more complex features but increase computational cost.</td>
</tr>
<tr>
<td align="left" valign="middle">Kernel size</td>
<td align="center" valign="middle">4&#x00D7;4</td>
<td align="left" valign="middle">Defines the size of the filter that scans the image. A smaller kernel captures fine details, while a larger one captures broader patterns.</td>
</tr>
<tr>
<td align="left" valign="middle">Activation function</td>
<td align="center" valign="middle">ReLU</td>
<td align="left" valign="middle">Introduces non-linearity to the network, allowing it to learn complex patterns. ReLU prevents the vanishing gradient problem.</td>
</tr>
<tr>
<td align="left" valign="middle">Pooling type</td>
<td align="center" valign="middle">MaxPooling</td>
<td align="left" valign="middle">Reduces spatial dimensions, retaining the most important features while lowering computational cost and reducing overfitting.</td>
</tr>
<tr>
<td align="left" valign="middle">Fully connected layer</td>
<td align="center" valign="middle">256 neurons</td>
<td align="left" valign="middle">Processes extracted features and maps them to the final classification decision. A higher number of neurons allows for better learning capacity.</td>
</tr>
<tr>
<td align="left" valign="middle">Output activation</td>
<td align="center" valign="middle">Sigmoid</td>
<td align="left" valign="middle">Outputs a probability score between 0 and 1, making it suitable for binary classification tasks.</td>
</tr>
<tr>
<td align="left" valign="middle">Loss function</td>
<td align="center" valign="middle">Binary cross-entropy</td>
<td align="left" valign="middle">Measures how well the predicted probabilities match the true labels, ensuring optimal training for binary classification.</td>
</tr>
<tr>
<td align="left" valign="middle">Optimizer</td>
<td align="center" valign="middle">Adam</td>
<td align="left" valign="middle">Adjusts learning rates dynamically for each parameter, leading to faster convergence and better training stability.</td>
</tr>
<tr>
<td align="left" valign="middle">Epochs</td>
<td align="center" valign="middle">20</td>
<td align="left" valign="middle">The number of times the model passes through the entire dataset during training. More epochs improve learning but may cause overfitting.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To overcome the limitations of the initial CNN model, a refined approach was adopted, incorporating particle swarm optimization (PSO) for hyperparameter tuning. PSO has demonstrated effectiveness in medical imaging by dynamically optimizing parameters for improved convergence (<xref ref-type="bibr" rid="ref13">Houssein et al., 2021</xref>). Rather than training a model from scratch, MobileNetV2&#x2014;a lightweight, pre-trained CNN trained on ImageNet&#x2014;was employed as a base feature extractor (illustrated by the teal blocks in <xref ref-type="fig" rid="fig4">Figure 4</xref>). Transfer learning with pre-trained models has been shown to enhance performance in histopathological classification tasks, especially when annotated datasets are limited (<xref ref-type="bibr" rid="ref42">Srinivasu et al., 2021</xref>).</p>
<fig position="float" id="fig4">
<label>Figure 4</label>
<caption>
<p>Updated CNN.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g004.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Diagram showing the architecture of a neural network using MobileNetV2 as a feature extractor. It includes components for input, MobileNetV2 model, PSO tuning, dense layer, learning and dropout rate, global average pooling, 256 neurons, and sigmoid output. Color-coded key explains each component: light purple for input, teal for MobileNetV2, red for PSO tuning, green for dense layer, blue for learning and dropout rate, orange for global average pooling, and yellow for output.</alt-text>
</graphic>
</fig>
<p>The BreaKHis dataset was stratified by magnification level (40&#x00D7;, 100&#x00D7;, 200&#x00D7;, and 400&#x00D7;), and separate models were trained for each resolution, aligning with evidence that such granularity improves classification performance in multi-resolution tasks (<xref ref-type="bibr" rid="ref6">Bayramoglu et al., 2017</xref>). Within each subset, class balance was ensured by maintaining 250&#x2013;350 images per class. Data augmentation was applied using ImageDataGenerator, with transformations including rescaling (1./255), width and height shifts (0.2), shear (0.2), zoom (0.2), and horizontal flips. These augmentations were crucial for improving generalization in data-scarce conditions (<xref ref-type="bibr" rid="ref4">Bandi et al., 2018</xref>).</p>
<p>PSO was used to fine-tune key hyperparameters&#x2014;specifically, learning rate (searched in the range 1e-5 to 1e-2) and dropout rate (0.3&#x2013;0.7)&#x2014;to enhance convergence and prevent overfitting. Optimization was first performed on the 40&#x202F;&#x00D7;&#x202F;magnification model; the resulting best parameters were then transferred to models for the other magnifications. To further stabilize training, a ReduceLROnPlateau scheduler (factor&#x202F;=&#x202F;0.5, patience&#x202F;=&#x202F;3, min_lr&#x202F;=&#x202F;1e-6) dynamically adjusted the learning rate based on validation loss trends, while EarlyStopping (patience&#x202F;=&#x202F;5, restore_best_weights&#x202F;=&#x202F;true) halted training when no improvement was observed, thereby avoiding overfitting and saving compute resources (<xref ref-type="bibr" rid="ref42">Srinivasu et al., 2021</xref>). The final architecture is illustrated in <xref ref-type="fig" rid="fig4">Figure 4</xref>. Hyperparameters for this architecture are in <xref ref-type="table" rid="tab4">Table 4</xref>.</p>
<table-wrap position="float" id="tab4">
<label>Table 4</label>
<caption>
<p>Hyperparameters for the first attempt of per-magnification training.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Hyperparameter</th>
<th align="left" valign="top">Value(s) used</th>
<th align="left" valign="top">Explanation</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Optimizer</td>
<td align="left" valign="top">Particle swarm optimization (PSO)</td>
<td align="left" valign="top">Instead of using Adam, PSO dynamically selects the best learning rate and dropout rate by minimizing validation loss. This helps in fine-tuning model performance more efficiently.</td>
</tr>
<tr>
<td align="left" valign="top">Learning rate</td>
<td align="left" valign="top">Optimized by PSO (range: 1e-5 to 1e-2)</td>
<td align="left" valign="top">Controls the step size of weight updates. A smaller learning rate ensures stable learning, while a larger one speeds up convergence but risks overshooting the optimal point.</td>
</tr>
<tr>
<td align="left" valign="top">Dropout rate</td>
<td align="left" valign="top">Optimized by PSO (range: 0.3&#x2013;0.7)</td>
<td align="left" valign="top">Dropout prevents overfitting by randomly deactivating neurons during training. The PSO algorithm selects the best dropout rate for optimal generalization.</td>
</tr>
<tr>
<td align="left" valign="top">Batch size</td>
<td align="left" valign="top">16</td>
<td align="left" valign="top">The number of images processed at once during training. A smaller batch size helps conserve memory but may result in higher training variance.</td>
</tr>
<tr>
<td align="left" valign="top">Input image size</td>
<td align="left" valign="top">(128, 128, 3)</td>
<td align="left" valign="top">The input images are resized to 128&#x202F;&#x00D7;&#x202F;128 pixels with three color channels to ensure uniformity and reduce computational complexity.</td>
</tr>
<tr>
<td align="left" valign="top">Base model</td>
<td align="left" valign="top">MobileNetV2 (pre-trained on ImageNet)</td>
<td align="left" valign="top">A lightweight deep learning model optimized for mobile and embedded vision applications. The feature extraction layers are frozen to retain learned representations.</td>
</tr>
<tr>
<td align="left" valign="top">Magnification-based training</td>
<td align="left" valign="top">40&#x00D7;, 100&#x00D7;, 200&#x00D7;, and 400&#x202F;&#x00D7;&#x202F;(trained separately)</td>
<td align="left" valign="top">Instead of training all images together, the dataset is split based on magnification levels to improve feature learning at each resolution.</td>
</tr>
<tr>
<td align="left" valign="top">Data augmentation</td>
<td align="left" valign="top">Rescaling (1./255), width and height shifts (0.2), shear (0.2), zoom (0.2), and horizontal flips</td>
<td align="left" valign="top">Artificially increases dataset size by applying transformations, helping the model generalize better to unseen samples.</td>
</tr>
<tr>
<td align="left" valign="top">Validation split</td>
<td align="left" valign="top">20%</td>
<td align="left" valign="top">A portion of the dataset is reserved for validation to evaluate model performance and prevent overfitting.</td>
</tr>
<tr>
<td align="left" valign="top">Loss function</td>
<td align="left" valign="top">Binary cross-entropy</td>
<td align="left" valign="top">Since the classification is binary (benign vs. malignant), binary cross-entropy is used to compute the error between predicted and actual labels.</td>
</tr>
<tr>
<td align="left" valign="top">Activation function</td>
<td align="left" valign="top">ReLU (hidden layers), sigmoid (output layer)</td>
<td align="left" valign="top">ReLU helps prevent vanishing gradients in hidden layers, while Sigmoid is used in the output layer for binary classification (probabilities between 0 and 1).</td>
</tr>
<tr>
<td align="left" valign="top">ReduceLROnPlateau</td>
<td align="left" valign="top">Factor&#x202F;=&#x202F;0.5, patience&#x202F;=&#x202F;3, min_lr&#x202F;=&#x202F;1e-6</td>
<td align="left" valign="top">Reduces the learning rate if validation loss stops improving for three consecutive epochs, preventing unnecessary weight updates.</td>
</tr>
<tr>
<td align="left" valign="top">EarlyStopping</td>
<td align="left" valign="top">Monitor&#x202F;=&#x202F;val_loss, patience&#x202F;=&#x202F;5, restore_best_weights&#x202F;=&#x202F;true</td>
<td align="left" valign="top">Stops training if validation loss does not improve for five consecutive epochs, preventing overfitting and saving computation time.</td>
</tr>
<tr>
<td align="left" valign="top">Epochs</td>
<td align="left" valign="top">30</td>
<td align="left" valign="top">The number of times the model sees the entire dataset during training. Early stopping ensures the model does not train longer than necessary.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>However, this approach too yielded an accuracy far below acceptable levels. F1-scores, recall and precision also hovered around the 50% mark, forcing another change in approach.</p>
<p>The classification of breast cancer histopathology images involved separating the dataset based on different magnification levels (40&#x00D7;, 100&#x00D7;, 200&#x00D7;, and 400&#x00D7;). While this method aimed to leverage magnification-specific features, it resulted in severe dataset fragmentation. Each individual model was trained on a significantly smaller subset of images, which limited the ability to generalize across different test samples. Even with extensive data augmentation, the number of available images remained insufficient, leading to high variance and suboptimal performance during evaluation. Additionally, the previous models relied on relatively shallow CNNs for feature extraction. These networks struggled to effectively capture both low-level and high-level representations within the histopathology images, leading to limited discriminatory power when distinguishing between benign and malignant samples. Given these challenges, a more robust and generalized model was required to improve classification performance across all magnification levels. To overcome the limitations of dataset fragmentation and poor feature extraction, a new model was developed using DenseNet121 as the backbone for feature extraction. DenseNet121, a deep CNN pre-trained on the ImageNet dataset, was chosen due to its efficient parameter utilization and ability to capture intricate hierarchical features. Unlike the previous approach, this model does not separate images based on magnification, allowing for a larger and more diverse training dataset, thereby improving generalization. The proposed model extracts features from three different depths of DenseNet121: conv3_block12_concat, conv4_block24_concat, and conv5_block16_concat. These layers correspond to different levels of abstraction within the network, ensuring that both fine-grained and high-level morphological characteristics are captured. The selection of these layers allows for multi-scale feature extraction, improving the model&#x2019;s ability to differentiate between benign and malignant samples. The extracted features are processed through a structured pipeline to refine and enhance their representation before classification.</p>
<p>Feature extraction in this model occurred at three levels. The lowest-level feature map, obtained from conv3_block12_concat, captures fundamental visual characteristics such as edges, textures, and color distributions. These features are crucial for identifying initial structural differences in histopathology images, such as variations in cellular arrangements. The mid-level feature map, extracted from conv4_block24_concat, represents more complex patterns, including structural organization within the tissue and variations in gland formation. These features help differentiate between normal and abnormal cellular arrangements. The highest-level feature map, obtained from conv5_block16_concat, provides a more abstract representation, focusing on morphological changes indicative of malignancy, such as nuclear pleomorphism, mitotic activity, and stromal alterations. By incorporating multiple feature extraction levels, the model ensures a comprehensive understanding of tissue characteristics rather than relying on a single-layer representation. To ensure a balance between computational efficiency and adequate feature resolution, an input image size of 128&#x202F;&#x00D7;&#x202F;128 was chosen over 256&#x202F;&#x00D7;&#x202F;256. While higher resolutions like 256&#x202F;&#x00D7;&#x202F;256 could preserve more fine-grained tissue structures, they also significantly increase memory usage and computational load. Given that histopathology images already exhibit high variability and detailed cellular patterns, 128&#x202F;&#x00D7;&#x202F;128 provides a sufficient level of detail for feature extraction while allowing for larger batch sizes, faster training, and reduced GPU memory constraints. This choice is particularly important when training multiple models, as excessive computational demands could slow down hyperparameter tuning and optimization.</p>
<p>Each extracted feature map undergoes global average pooling (GAP) to reduce dimensionality while retaining essential spatial information. Unlike max pooling, which selects only the most prominent activations, GAP ensures that all features contribute proportionally to the final representation, enhancing robustness and stability. After pooling, L2 normalization is applied to ensure numerical stability and prevent dominance of certain feature values due to large magnitudes. The normalized feature vectors are then transformed through fully connected layers, where a 64-neuron dense layer introduces non-linearity, allowing for more complex feature interactions. Batch normalization is used to stabilize learning by standardizing feature distributions and improving gradient flow, leading to faster convergence and better generalization. This entire process for one feature is shown by the four blocks in each row in <xref ref-type="fig" rid="fig5">Figure 5</xref>. Instead of treating each extracted feature set independently, the model employs a fusion mechanism where feature representations from all three levels are concatenated into a unified feature descriptor (the white block in <xref ref-type="fig" rid="fig5">Figure 5</xref>). This approach integrates information across different abstraction levels, allowing the model to make more informed classification decisions. The fused feature vector undergoes further transformation through a dense layer with 16 neurons, refining the representation before classification. To prevent overfitting, a dropout layer with a probability of 0.45 is applied at this stage, randomly deactivating neurons during training and ensuring that the model does not rely on specific patterns that may not generalize well.</p>
<fig position="float" id="fig5">
<label>Figure 5</label>
<caption>
<p>Final binary class CNN.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g005.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Diagram of a neural network using the DenseNet architecture. It begins with an input layer, followed by pre-trained DenseNet121 features, L2 normalization, batch normalization, and dropout layers. The network includes dense layers with sixty-four and sixteen neurons leading to a softmax output. Features are fused before the final output.</alt-text>
</graphic>
</fig>
<p>The final classification layer consists of a softmax activation function with two neurons, representing the benign and malignant classes. This output layer assigns a probability score to each class, allowing for precise decision-making in binary classification tasks. By leveraging deep feature extraction, multi-scale fusion, and regularization techniques, the proposed model significantly improves the accuracy and robustness of breast cancer classification in histopathology images compared to the previous two models, as well as objectively. <xref ref-type="fig" rid="fig5">Figure 5</xref> diagrammatically describes the flow of the updated CNN architecture (see <xref ref-type="table" rid="tab5">Table 5</xref>).</p>
<table-wrap position="float" id="tab5">
<label>Table 5</label>
<caption>
<p>Hyperparameters for the final binary classification model.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Hyperparameter</th>
<th align="left" valign="top">Value</th>
<th align="left" valign="top">Explanation</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Base model</td>
<td align="left" valign="top">DenseNet121 (pre-trained on ImageNet)</td>
<td align="left" valign="top">Acts as the backbone for feature extraction, leveraging pre-trained hierarchical features.</td>
</tr>
<tr>
<td align="left" valign="top">Input image size</td>
<td align="left" valign="top">(128, 128, 3)</td>
<td align="left" valign="top">Defines the input shape of images with three color channels (RGB), ensuring consistency across the dataset.</td>
</tr>
<tr>
<td align="left" valign="top">Feature extraction layers</td>
<td align="left" valign="top">conv3_block12_concat, conv4_block24_concat, conv5_block16_concat</td>
<td align="left" valign="top">Extracts multi-scale features from different levels of abstraction within the DenseNet121 model.</td>
</tr>
<tr>
<td align="left" valign="top">Pooling method</td>
<td align="left" valign="top">Global average pooling (GAP)</td>
<td align="left" valign="top">Reduces dimensionality while retaining essential spatial information from feature maps.</td>
</tr>
<tr>
<td align="left" valign="top">Feature normalization</td>
<td align="left" valign="top">L2 normalization</td>
<td align="left" valign="top">Ensures all extracted features contribute proportionally, preventing large-magnitude features from dominating learning.</td>
</tr>
<tr>
<td align="left" valign="top">Hidden dense layer (feature processing)</td>
<td align="left" valign="top">64 neurons, ReLU activation, L2 regularization (0.001)</td>
<td align="left" valign="top">Introduces non-linearity to enhance feature interactions and prevents overfitting via L2 regularization.</td>
</tr>
<tr>
<td align="left" valign="top">Batch normalization</td>
<td align="left" valign="top">Applied after dense layers</td>
<td align="left" valign="top">Standardizes feature distributions, improving gradient flow and training stability.</td>
</tr>
<tr>
<td align="left" valign="top">Feature fusion</td>
<td align="left" valign="top">Concatenation of three feature levels</td>
<td align="left" valign="top">Integrates low-level, mid-level, and high-level features into a unified representation for better classification.</td>
</tr>
<tr>
<td align="left" valign="top">Final dense layer</td>
<td align="left" valign="top">16 neurons, ReLU activation, L2 regularization (0.001)</td>
<td align="left" valign="top">Further refines the fused feature representation before classification.</td>
</tr>
<tr>
<td align="left" valign="top">Dropout rate</td>
<td align="left" valign="top">0.45</td>
<td align="left" valign="top">Randomly deactivated neurons to prevent overfitting, ensuring better generalization.</td>
</tr>
<tr>
<td align="left" valign="top">Output layer</td>
<td align="left" valign="top">2 neurons, softmax activation</td>
<td align="left" valign="top">Generates class probabilities for binary classification (Benign vs. Malignant).</td>
</tr>
<tr>
<td align="left" valign="top">Loss function</td>
<td align="left" valign="top">Categorical cross-entropy</td>
<td align="left" valign="top">Suitable for multi-class classification tasks, even though the output is binary (ensures numerical stability).</td>
</tr>
<tr>
<td align="left" valign="top">Optimizer</td>
<td align="left" valign="top">Adam (learning rate: 0.0001)</td>
<td align="left" valign="top">Adaptive optimization algorithm that adjusts learning rates dynamically for efficient training.</td>
</tr>
<tr>
<td align="left" valign="top">Learning rate scheduler</td>
<td align="left" valign="top">ReduceLROnPlateau (factor&#x202F;=&#x202F;0.5, patience&#x202F;=&#x202F;3, min_lr&#x202F;=&#x202F;1e-6)</td>
<td align="left" valign="top">Reduces the learning rate when validation loss stagnates, helping fine-tune the model.</td>
</tr>
<tr>
<td align="left" valign="top">EarlyStopping</td>
<td align="left" valign="top">Patience&#x202F;=&#x202F;5, restore_best_weights&#x202F;=&#x202F;true</td>
<td align="left" valign="top">Stops training when validation loss stops improving, preventing unnecessary computation and overfitting.</td>
</tr>
<tr>
<td align="left" valign="top">Epochs</td>
<td align="left" valign="top">50</td>
<td align="left" valign="top">Maximum number of iterations through the dataset, ensuring the model learns enough before early stopping.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="sec7">
<label>2.2.2</label>
<title>Classifying images as subtypes of benign and malignant</title>
<p>There are four histologically distinct types of benign breast tumors: adenosis (A), fibroadenoma (F), PT, and TA; and four malignant tumors (breast cancer): carcinoma (DC), LC, mucinous carcinoma (MC) and PC.</p>
<p>Adenosis (A) is characterized by the proliferation of glandular structures, which may appear as distorted acini or ducts in biopsy images. These lesions typically show increased glandular tissue without significant atypia, making them relatively easy to differentiate from malignancies under a microscope (<xref ref-type="bibr" rid="ref14">Jiang et al., 2017</xref>). Fibroadenomas (F) are solid tumors composed of both glandular and stromal components, presenting as well-circumscribed masses with minimal cellular atypia. Their typical pattern in biopsy images shows a lobular architecture with a fibrous stroma, which contrasts with the irregular structures seen in malignant lesions (<xref ref-type="bibr" rid="ref9007">Schulz-Wendtland et al., 2016</xref>). PTs, though benign in some cases, can show features suggestive of malignancy, such as rapid growth and hypercellularity. Their biopsy images typically display leaf-like structures, hence the name, with prominent stromal overgrowth and areas of necrosis that can challenge differentiation from malignant tumors (<xref ref-type="bibr" rid="ref5">Barker et al., 2018</xref>). TAs are another benign subtype characterized by well-formed tubules lined by a single layer of epithelial cells. These lesions often have minimal cellular atypia, and their biopsy images demonstrate small, uniformly sized tubules with little stroma (<xref ref-type="bibr" rid="ref31">Pike et al., 2020</xref>).</p>
<p>In contrast, malignant breast cancers are often more challenging to classify due to the diversity of their histopathological features. DC is the most common type of breast cancer and typically appears as irregularly shaped masses with infiltrating ductal structures. The biopsy images show marked cellular pleomorphism, nuclear atypia, and often necrotic areas, which are indicative of aggressive behavior (<xref ref-type="bibr" rid="ref45">Tung et al., 2016</xref>). LC presents with a distinctive pattern in biopsy images, with tumor cells growing in a single-file pattern and a lack of cohesive cell groups. This &#x201C;Indian file&#x201D; arrangement of cells and the absence of desmoplastic stroma are key features that differentiate it from DC (<xref ref-type="bibr" rid="ref9023">Vargas et al., 2017</xref>). Mucinous carcinoma (MC) is characterized by the production of extracellular mucin, which can be observed in the biopsy images as abundant mucin pools, separating the tumor cells. The cells in mucinous carcinoma are typically round with mild pleomorphism and are surrounded by mucin-rich stroma (<xref ref-type="bibr" rid="ref26">Manning et al., 2018</xref>). Finally, PC shows prominent papillary structures with fibrovascular cores, surrounded by atypical epithelial cells. The tumor cells are often arranged in well-defined papillae, and the biopsy images reveal a complex architecture with cystic spaces and dense fibrous stroma (<xref ref-type="bibr" rid="ref44">Thompson et al., 2019</xref>).</p>
<p>A CNN was initially trained to distinguish breast cancer subtypes by learning hierarchical spatial features from biopsy images. It captured architectural differences such as the regular tubules in TAs and the disorganized ductal structures in carcinomas, as well as key histological features like pleomorphism, stromal composition, and mucin presence (<xref ref-type="bibr" rid="ref32">Rajendran et al., 2018</xref>; <xref ref-type="bibr" rid="ref9021">Srinivasu et al., 2020</xref>). The architecture included three Conv2D layers (32, 64, 128 filters), ReLU activations, and (3&#x202F;&#x00D7;&#x202F;3) kernels for efficient feature extraction, with max pooling and dropout (0.25/0.5) to reduce overfitting. A dense layer with 512 neurons and a softmax output layer enabled multi-class classification, optimized using categorical cross-entropy and various learning rates (Adam 0.0001, SGD 0.01, RMSprop 0.0001). Training was conducted over 30 epochs with a batch size of 32. Despite these efforts, the model struggled with generalization and subtype discrimination, yielding poor validation performance. Following the success of the DenseNet-based binary model, the architecture was extended and adapted to a multiclass setting, as detailed in Section 2.3 (see <xref ref-type="table" rid="tab6">Table 6</xref>).</p>
<table-wrap position="float" id="tab6">
<label>Table 6</label>
<caption>
<p>Hyperparameters for the initial subclass classification model.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Hyperparameter</th>
<th align="center" valign="top">Value</th>
<th align="left" valign="top">Explanation</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Input image size</td>
<td align="center" valign="top">(256, 256)</td>
<td align="left" valign="top">The spatial dimensions of each input image are fed to the model.</td>
</tr>
<tr>
<td align="left" valign="top">Batch size</td>
<td align="center" valign="top">32</td>
<td align="left" valign="top">Number of images processed at once during training.</td>
</tr>
<tr>
<td align="left" valign="top">Learning rate (Adam)</td>
<td align="center" valign="top">0.0001</td>
<td align="left" valign="top">Speed at which the model updates weights during training using Adam optimizer.</td>
</tr>
<tr>
<td align="left" valign="top">Learning rate (SGD)</td>
<td align="center" valign="top">0.01</td>
<td align="left" valign="top">Speed of weight updates for the SGD optimizer.</td>
</tr>
<tr>
<td align="left" valign="top">SGD momentum</td>
<td align="center" valign="top">0.9</td>
<td align="left" valign="top">Momentum helps accelerate SGD by dampening oscillations during updates.</td>
</tr>
<tr>
<td align="left" valign="top">Learning rate (RMSprop)</td>
<td align="center" valign="top">0.0001</td>
<td align="left" valign="top">Step size for weight updates using RMSprop optimizer.</td>
</tr>
<tr>
<td align="left" valign="top">Conv2D filters (1st layer)</td>
<td align="center" valign="top">32</td>
<td align="left" valign="top">Number of filters in the first convolutional layer.</td>
</tr>
<tr>
<td align="left" valign="top">Conv2D filters (2nd layer)</td>
<td align="center" valign="top">64</td>
<td align="left" valign="top">Number of filters in the second convolutional layer.</td>
</tr>
<tr>
<td align="left" valign="top">Conv2D filters (3rd layer)</td>
<td align="center" valign="top">128</td>
<td align="left" valign="top">Number of filters in the third convolutional layer.</td>
</tr>
<tr>
<td align="left" valign="top">Kernel Size</td>
<td align="center" valign="top">(3, 3)</td>
<td align="left" valign="top">Size of the sliding window applied in convolution operations.</td>
</tr>
<tr>
<td align="left" valign="top">Activation function</td>
<td align="center" valign="top">ReLU</td>
<td align="left" valign="top">Non-linear function applied to neuron outputs to introduce non-linearity.</td>
</tr>
<tr>
<td align="left" valign="top">Dropout rate (conv layers)</td>
<td align="center" valign="top">0.25</td>
<td align="left" valign="top">The fraction of neurons dropped to prevent overfitting during training.</td>
</tr>
<tr>
<td align="left" valign="top">Dropout rate (dense layer)</td>
<td align="center" valign="top">0.5</td>
<td align="left" valign="top">Fraction of dense layer neurons dropped for regularization.</td>
</tr>
<tr>
<td align="left" valign="top">pooling size</td>
<td align="center" valign="top">(2, 2)</td>
<td align="left" valign="top">Size of the window for max pooling operation to reduce feature map dimensions.</td>
</tr>
<tr>
<td align="left" valign="top">Dense layer size</td>
<td align="center" valign="top">512</td>
<td align="left" valign="top">Number of neurons in the fully connected layer before the output.</td>
</tr>
<tr>
<td align="left" valign="top">Output layer activation</td>
<td align="center" valign="top">Softmax</td>
<td align="left" valign="top">Ensures output values represent probabilities for multi-class classification.</td>
</tr>
<tr>
<td align="left" valign="top">Loss function</td>
<td align="center" valign="top">Categorical cross-entropy</td>
<td align="left" valign="top">Measures model error for multi-class classification tasks.</td>
</tr>
<tr>
<td align="left" valign="top">Number of epochs</td>
<td align="center" valign="top">30</td>
<td align="left" valign="top">Total number of complete passes through the training dataset.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To enhance diagnostic granularity, we developed two separate CNN classifiers&#x2014;one for benign subtypes and another for malignant subtypes&#x2014;rather than training a single multi-class model across all categories to extend the binary class classifier to multiple classes. This decision is rooted in both clinical and computational rationale. Clinically, benign and malignant lesions differ not just in severity but in their underlying histomorphological characteristics, tissue organization, and staining patterns (<xref ref-type="bibr" rid="ref5">Barker et al., 2018</xref>; <xref ref-type="bibr" rid="ref31">Pike et al., 2020</xref>). Combining both under one classifier risks underrepresenting these distinct patterns during training, especially when using class-imbalanced datasets like BreakHis, which contain fewer samples for certain subtypes. Computationally, separate models allow more focused feature learning within each class family, minimizing intra-class confusion. Prior studies have shown that domain-specific subnetworks often outperform unified multi-class models in pathology, especially when classes exhibit heterogeneity in structure and frequency (<xref ref-type="bibr" rid="ref9">Gandomkar et al., 2018</xref>; <xref ref-type="bibr" rid="ref30">Patel et al., 2024</xref>). For example, <xref ref-type="bibr" rid="ref46">Umer et al. (2022)</xref> observed that a specialized multi-branch CNN improved the classification of individual breast cancer subtypes, while <xref ref-type="bibr" rid="ref3">Arevalo et al. (2016)</xref> showed that focused lesion-specific training led to better mass detection in mammography. Furthermore, multi-class CNNs trained across all subtypes often suffer from performance trade-offs&#x2014;gaining sensitivity in one category while losing specificity in others (<xref ref-type="bibr" rid="ref27">McKinney et al., 2020</xref>). By using separate networks, we achieved higher per-class precision and recall, especially for visually similar classes like fibroadenoma and PT. This modular architecture also supports more targeted clinical applications, such as malignancy-specific triaging or subtype-based treatment suggestion systems, which align with current trends in personalized digital pathology (<xref ref-type="bibr" rid="ref21">Lee et al., 2024</xref>).</p>
<p>The multi-class CNN architectures for subtype classification were directly adapted from the binary classification model described previously. DenseNet121, used as the backbone in both approaches, provides efficient deep feature extraction through its dense connectivity and hierarchical learning capabilities. As in the binary model, features were extracted from three depths&#x2014;conv3_block12_concat, conv4_block24_concat, and conv5_block16_concat&#x2014;capturing progressively abstract morphological features necessary for histopathological differentiation. These include low-level structures such as cellular edges and textures, mid-level patterns like glandular and stromal organization, and high-level attributes such as nuclear pleomorphism and mitotic activity.</p>
<p>Each extracted feature map underwent GAP, L2 normalization, and transformation <italic>via</italic> a dense layer with L2 regularization, batch normalization, and dropout, following the same processing pipeline outlined in the binary classifier. The processed features were then concatenated into a unified multi-scale descriptor and passed through a final dense layer before classification.</p>
<p>Instead of a binary output, the final softmax layer consisted of four neurons to predict between the benign or malignant subtypes, depending on the network. Two separate CNNs were trained&#x2014;one for benign subtypes and another for malignant&#x2014;to reduce feature entanglement and focus each model on intra-class variation. This design choice improved the model&#x2019;s capacity to learn subtle distinctions specific to each class group, such as distinguishing fibroadenoma from PT, or DC from PC.</p>
<p>Optimization and training procedures mirrored those of the binary classification model, utilizing the Adam optimizer (learning rate 0.0001), EarlyStopping, and ReduceLROnPlateau for improved convergence and generalization (see <xref ref-type="fig" rid="fig6">Figures 6</xref>, <xref ref-type="fig" rid="fig7">7</xref> and <xref ref-type="table" rid="tab7">Table 7</xref>).</p>
<fig position="float" id="fig6">
<label>Figure 6</label>
<caption>
<p>CNN used for multi-class classification.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g006.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Diagram of a neural network architecture incorporating DenseNet features. Input flows through DenseNet121 layers, including L2 normalization and global average pooling. It transitions through 64 neurons and 16 neurons before a softmax layer with 4 neurons. Components include batch normalization, dropout layers, and L2 regularization. The color-coded legend indicates each layer type.</alt-text>
</graphic>
</fig>
<fig position="float" id="fig7">
<label>Figure 7</label>
<caption>
<p>Diagram for feature workflows.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g007.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Flowchart of a feature workflow using a neural network. Begins with DenseNet121 feature extraction, then global average pooling in two dimensions, followed by L2 normalization. This is passed through a dense layer with 64 neurons, then L2 regularization and batch normalization.</alt-text>
</graphic>
</fig>
<table-wrap position="float" id="tab7">
<label>Table 7</label>
<caption>
<p>Hyperparameters for the final subclass classification model.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Hyperparameter</th>
<th align="left" valign="top">Value</th>
<th align="left" valign="top">Explanation</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Base model</td>
<td align="left" valign="top">DenseNet121 (pre-trained on ImageNet)</td>
<td align="left" valign="top">Acts as a deep feature extractor, leveraging hierarchical feature learning for efficient classification.</td>
</tr>
<tr>
<td align="left" valign="top">Input image size</td>
<td align="left" valign="top">(128, 128, 3)</td>
<td align="left" valign="top">Ensures computational efficiency while preserving essential histopathological details. A higher resolution, like 256&#x202F;&#x00D7;&#x202F;256, would increase memory usage without significantly improving feature representation.</td>
</tr>
<tr>
<td align="left" valign="top">Feature extraction layers</td>
<td align="left" valign="top">conv3_block12_concat, conv4_block24_concat, conv5_block16_concat</td>
<td align="left" valign="top">Extracts features at multiple abstraction levels, capturing both fine-grained and high-level morphological details.</td>
</tr>
<tr>
<td align="left" valign="top">Pooling method</td>
<td align="left" valign="top">Global average pooling (GAP)</td>
<td align="left" valign="top">Reduces dimensionality while preserving spatial characteristics of the extracted feature maps.</td>
</tr>
<tr>
<td align="left" valign="top">Feature normalization</td>
<td align="left" valign="top">L2 normalization</td>
<td align="left" valign="top">Prevents feature values with high magnitudes from dominating the learning process, ensuring numerical stability.</td>
</tr>
<tr>
<td align="left" valign="top">Hidden dense layer (feature processing)</td>
<td align="left" valign="top">64 neurons, ReLU activation, L2 regularization (0.001)</td>
<td align="left" valign="top">Introduces non-linearity for complex feature interactions and applies L2 regularization to prevent overfitting.</td>
</tr>
<tr>
<td align="left" valign="top">Batch normalization</td>
<td align="left" valign="top">Applied after dense layer s</td>
<td align="left" valign="top">Standardizes feature distributions, accelerating convergence and improving training stability.</td>
</tr>
<tr>
<td align="left" valign="top">Feature fusion</td>
<td align="left" valign="top">Concatenation of three feature levels</td>
<td align="left" valign="top">Integrates information across different levels of abstraction, enhancing classification performance.</td>
</tr>
<tr>
<td align="left" valign="top">Final dense layer</td>
<td align="left" valign="top">16 neurons, ReLU activation, L2 regularization (0.001)</td>
<td align="left" valign="top">Refines the multi-scale feature representation before classification.</td>
</tr>
<tr>
<td align="left" valign="top">Dropout rate</td>
<td align="left" valign="top">0.45</td>
<td align="left" valign="top">Prevents overfitting by randomly deactivating neurons, ensuring better generalization.</td>
</tr>
<tr>
<td align="left" valign="top">Output layer</td>
<td align="left" valign="top">4 neurons, softmax activation</td>
<td align="left" valign="top">Outputs probability scores for the four benign subtypes.</td>
</tr>
<tr>
<td align="left" valign="top">Loss function</td>
<td align="left" valign="top">Categorical cross-entropy</td>
<td align="left" valign="top">Suitable for multi-class classification tasks, optimizing probability-based learning.</td>
</tr>
<tr>
<td align="left" valign="top">Optimizer</td>
<td align="left" valign="top">Adam (learning rate: 0.0001)</td>
<td align="left" valign="top">Adaptive optimization that adjusts learning rates dynamically for efficient and stable training.</td>
</tr>
<tr>
<td align="left" valign="top">Learning rate scheduler</td>
<td align="left" valign="top">ReduceLROnPlateau (factor&#x202F;=&#x202F;0.5, patience&#x202F;=&#x202F;3, min_lr&#x202F;=&#x202F;1e-6)</td>
<td align="left" valign="top">Reduces learning rate when validation loss stagnates, improving fine-tuning of weights.</td>
</tr>
<tr>
<td align="left" valign="top">EarlyStopping</td>
<td align="left" valign="top">Patience&#x202F;=&#x202F;5, restore_best_weights&#x202F;=&#x202F;true</td>
<td align="left" valign="top">Stops training when validation loss stops improving, preventing unnecessary computation and overfitting.</td>
</tr>
<tr>
<td align="left" valign="top">Epochs</td>
<td align="left" valign="top">50</td>
<td align="left" valign="top">Maximum number of training iterations through the dataset, ensuring sufficient learning.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="sec8">
<label>2.3</label>
<title>Novel characteristics of the model</title>
<p>The proposed classification algorithm is distinguished by its multi-scale feature fusion architecture built on DenseNet121. Instead of relying only on the final layer features (as done in most standard CNN classifiers), it extracts and fuses features from three different depths of the DenseNet (the end of conv3, conv4, and conv5 blocks). Each feature map is subjected to GAP and L2 normalization, then passed through a small fully connected layer with batch normalization (and dropout) before concatenation. This design enables the model to capture discriminative patterns at multiple spatial scales&#x2014;from fine-grained cellular details to higher-level tissue organization&#x2014;within a single network. In essence, the model behaves like an integrated multi-scale ensemble, where lower-level texture features and higher-level semantic features jointly inform the final prediction. This multi-scale fusion is particularly well-suited to histopathology images, where diagnostically relevant cues appear at different magnifications. Prior work has recognized the value of multi-scale information: <xref ref-type="bibr" rid="ref9003">Ara&#x00FA;jo et al. (2017)</xref> designed a custom CNN to explicitly capture both nuclei-level details and global tissue architecture (<xref ref-type="bibr" rid="ref9003">Ara&#x00FA;jo et al., 2017</xref>), and <xref ref-type="bibr" rid="ref11">Gupta and Bhavsar (2018)</xref> similarly argued that &#x201C;different layers&#x2026; contain useful discriminative information&#x201D; for classifying breast histology images (<xref ref-type="bibr" rid="ref11">Gupta and Bhavsar, 2018</xref>). The proposed model builds on this principle by simultaneously leveraging DenseNet&#x2019;s intermediate feature maps, rather than using only the deepest features or requiring multiple separate networks.</p>
<p>By fusing intermediate feature vectors, our approach yields a richer representation that combines context from multiple receptive field sizes. This strategy is analogous to the &#x201C;hypercolumn&#x201D; or feature pyramid concept&#x2014;effectively consolidating features at various levels of abstraction. In histopathology, such multi-level fusion has clear advantages: for example, benign vs. malignant differentiation may depend on both cellular morphology and the broader tissue architecture. Traditional transfer-learning classifiers often discard these multi-level cues by only using the final global feature map of a pre-trained CNN. In contrast, our DenseNet121-based model preserves multi-scale feature information by tapping conv3, conv4, and conv5 outputs. <xref ref-type="bibr" rid="ref51">Zhu et al. (2019)</xref> demonstrated a similar idea with a two-branch &#x201C;global vs. local&#x201D; CNN ensemble, where merging a whole-image branch with a high-zoom patch branch improved representation power (<xref ref-type="bibr" rid="ref51">Zhu et al., 2019</xref>). Our method achieves a comparable multi-scale effect within a single backbone: the conv3_block output (earlier layer) can focus on local textural patterns (analogous to a high-magnification view), while the conv5_block output provides global context (analogous to a low-magnification overview). The DenseNet architecture&#x2019;s dense connectivity may further facilitate this feature reuse across scales (<xref ref-type="bibr" rid="ref9018">Huang et al., 2017</xref>). Notably, DenseNet121 has been a popular choice in medical image analysis due to its efficient feature propagation; several breast histopathology studies report strong results with DenseNet-based transfer learning&#x2014;for example, <xref ref-type="bibr" rid="ref23">Liew et al. (2021)</xref> achieved ~97% accuracy by using DenseNet201 features with an XGBoost classifier (<xref ref-type="bibr" rid="ref23">Liew et al., 2021</xref>). Our model differentiates itself by not only fine-tuning DenseNet121 on the task, but by modifying its topology to combine multi-scale features&#x2014;a design rarely explored in prior breast cancer studies.</p>
<p>Another novel aspect of our approach is the use of L2 normalization and balanced regularization in the feature fusion process. Each branch&#x2019;s pooled feature vector is L2-normalized before fusion, ensuring that no single scale&#x2019;s features dominate the others due to scale differences. This is an uncommon yet effective practice in classification networks&#x2014;more frequently seen in metric learning or multi-instance aggregation contexts&#x2014;and it promotes stability when training the fused classifier. For instance, <xref ref-type="bibr" rid="ref49">Xie et al. (2020)</xref> applied L2 normalization to patch feature vectors in a histopathology model to stabilize feature magnitudes (<xref ref-type="bibr" rid="ref49">Xie et al., 2020</xref>). In our design, the L2 norm acts similarly to a scale calibration across the three feature streams, so that the subsequent dense layers operate on features of comparable norm. We then apply dropout and batch normalization in each branch&#x2019;s dense layer (as well as the final classification layer) to combat overfitting. High-resolution histology images are prone to overfitting given their limited datasets, and regularization is critical. Previous works have indeed found that naive CNN models can overfit histopathology data even with dropout&#x2014;for example, a shallow CNN still overfit BreakHis patches even with 20% dropout, as reported by <xref ref-type="bibr" rid="ref15">Kode and Barkana (2023)</xref> (<xref ref-type="bibr" rid="ref15">Kode and Barkana, 2023</xref>). By combining dropout with batch normalization and the inherent regularization of feature fusion, our model aims to generalize better. The inclusion of batch normalization after the small dense layers also helps each scale-specific feature vector to be well-conditioned before merging, which eases the joint learning. These design choices&#x2014;L2 normalization of features and dropout-regularized dense layers for each scale&#x2014;are unique refinements that distinguish our architecture from standard fine-tuned CNN classifiers that often use only a global pooling and a linear classifier on top.</p>
<p>Compared to existing deep learning methods in breast cancer histopathology, the proposed model offers a new balance of simplicity and multi-scale sophistication. Many prior studies have focused either on transfer learning with single CNNs or on ensembles of multiple models. On the one hand, simple transfer-learning approaches, fine-tuning a single network (e.g., ResNet or DenseNet), have shown strong baseline performance (<xref ref-type="bibr" rid="ref41">Spanhol et al., 2016</xref>; <xref ref-type="bibr" rid="ref9025">Wahab et al., 2017</xref>), but they might miss multi-scale cues. Some recent works augment such models with attention mechanisms&#x2014;for example, an EfficientNetV2 with a channel-spatial attention module (CBAM) by <xref ref-type="bibr" rid="ref9001">Aldakhil et al. (2024)</xref> outperformed standard CNNs, reaching ~92&#x2013;94% accuracy on BreakHis magnifications (<xref ref-type="bibr" rid="ref9001">Aldakhil et al., 2024</xref>). Others integrate vision transformers to capture global context&#x2014;for example, <xref ref-type="bibr" rid="ref28">Mehta et al. (2022)</xref> introduced a ViT-based HATNet that achieved state-of-the-art accuracy (<xref ref-type="bibr" rid="ref28">Mehta et al., 2022</xref>). However, these sophisticated models often remain single-scale in the sense that they ultimately rely on one level of feature map (global features) enhanced via attention. On the other hand, ensemble and multi-branch strategies explicitly leverage multi-scale inputs. For example, <xref ref-type="bibr" rid="ref46">Umer et al. (2022)</xref> proposed <italic>6B-Net</italic>, a six-branch CNN with different receptive field sizes in each branch, and fused their outputs for classifying the eight BreakHis classes (<xref ref-type="bibr" rid="ref46">Umer et al., 2022</xref>). Similarly, an earlier study by <xref ref-type="bibr" rid="ref51">Zhu et al. (2019)</xref> assembled multiple compact CNNs (and even pruned channels) to combine local and global predictions for histology images (<xref ref-type="bibr" rid="ref51">Zhu et al., 2019</xref>). These methods confirmed that combining features from multiple scales or multiple models can boost accuracy. Our approach achieves this multi-scale fusion within a single pre-trained DenseNet, rather than requiring separate networks for different magnifications or an external ensemble of classifiers. This provides a more unified and computationally efficient framework: for instance, <xref ref-type="bibr" rid="ref47">Wakili et al. (2022)</xref> reported a DenseNet-based transfer model (&#x201C;DenTnet&#x201D;) that attained ~99% accuracy on BreakHis (<xref ref-type="bibr" rid="ref47">Wakili et al., 2022</xref>), but that model treats DenseNet as a black-box feature extractor for an SVM or softmax classifier. In contrast, our model integrates the multi-level feature extraction into the training process, which could offer better synergy between feature learning and classification. It also simplifies the pipeline&#x2014;there is no need for post-CNN classifiers like SVM (as used by <xref ref-type="bibr" rid="ref9003">Ara&#x00FA;jo et al., 2017</xref> and others) or gradient boosting ensembles (as in <xref ref-type="bibr" rid="ref23">Liew et al., 2021</xref>) because the fused deep features are learned end-to-end to directly optimize classification.</p>
<p>In summary, the novelty of our algorithm lies in combining multi-scale feature fusion, transfer learning with DenseNet121, and strategic normalization/regularization into one coherent model for breast histopathology classification. This design captures the multi-scale nature of tissue patterns more explicitly than standard CNN classifiers. It leverages DenseNet&#x2019;s strength in feature reuse while addressing multiple scales akin to multi-branch networks, but with fewer parameters and a single-pass inference. By comparing with the literature, we see that our approach is unique in <italic>how</italic> it fuses intermediate CNN features: <xref ref-type="bibr" rid="ref11">Gupta and Bhavsar (2018)</xref> sequentially processed DenseNet layers&#x2019; outputs rather than fusing them, and most other DenseNet-based methods (<xref ref-type="bibr" rid="ref47">Wakili et al., 2022</xref>; <xref ref-type="bibr" rid="ref23">Liew et al., 2021</xref>) used only the final layer or combined whole-network outputs in ensembles. Moreover, the inclusion of L2-normalized feature vectors and dropout-regularized dense layers for each scale is a novel architectural choice that, to our knowledge, has not been explicitly reported in prior breast cancer histopathology studies. These innovations position our model as a multi-scale, multi-level classifier that is well-aligned with the visual hierarchy a pathologist employs (from cells to tissue architecture), setting it apart from conventional single-scale CNN models and even from recent attention-based or ensemble-based state-of-the-art methods in this domain.</p>
</sec>
<sec id="sec9">
<label>2.4</label>
<title>Metrics evaluations</title>
<p>In AI, especially in domains like classification and regression, evaluation metrics play a vital role in measuring model performance. Within medical diagnostics, metrics such as precision and recall are particularly significant. Precision reflects the proportion of correct positive predictions, meaning that when a model identifies cancer, it is likely to be correct&#x2014;thereby reducing unnecessary alarm (<xref ref-type="bibr" rid="ref40">Sokolova and Lapalme, 2009</xref>). Recall, on the other hand, focuses on identifying as many actual positive cases as possible, which is crucial in healthcare settings to avoid overlooking genuine cases of disease (<xref ref-type="bibr" rid="ref40">Sokolova and Lapalme, 2009</xref>). Although accuracy is a common metric indicating the percentage of correctly predicted instances overall, it can be deceptive in scenarios with class imbalance, as it does not account for the nature of errors made (<xref ref-type="bibr" rid="ref43">Tharwat, 2020</xref>).</p>
<p>To address such limitations, the F1-score is often employed. This metric, calculated as the harmonic mean of precision and recall, offers a more balanced view by considering both false positives and false negatives&#x2014;an essential factor in medical imaging, where failing to detect malignant tumors can have serious consequences (<xref ref-type="bibr" rid="ref8">Chicco and Jurman, 2020</xref>). Additionally, the area under the curve (AUC) of the receiver operating characteristic (ROC) curve is another key metric. It evaluates the model&#x2019;s capacity to distinguish between positive (cancerous) and negative (non-cancerous) cases, with scores approaching 1 indicating high discriminative ability (<xref ref-type="bibr" rid="ref9028">Yang and Ying, 2022</xref>). Given the frequent imbalance in medical datasets, relying on a combination of these metrics provides a more reliable and holistic evaluation of model effectiveness (see <xref ref-type="table" rid="tab8">Table 8</xref>).</p>
<table-wrap position="float" id="tab8">
<label>Table 8</label>
<caption>
<p>Metrics for evaluation of the algorithm.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Metric</th>
<th align="left" valign="top">Formula</th>
<th align="left" valign="top">Explanation</th>
<th align="left" valign="top">References</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Precision</td>
<td align="left" valign="top">TP/(TP&#x202F;+&#x202F;FP)</td>
<td align="left" valign="top">Measures the proportion of correctly predicted positive cases out of all predicted positives.</td>
<td align="left" valign="top">
<xref ref-type="bibr" rid="ref40">Sokolova and Lapalme (2009)</xref>
</td>
</tr>
<tr>
<td align="left" valign="top">Recall (sensitivity)</td>
<td align="left" valign="top">TP/(TP&#x202F;+&#x202F;FN)</td>
<td align="left" valign="top">Measures the proportion of actual positive cases correctly identified by the model.</td>
<td align="left" valign="top">
<xref ref-type="bibr" rid="ref40">Sokolova and Lapalme (2009)</xref>
</td>
</tr>
<tr>
<td align="left" valign="top">F1-score</td>
<td align="left" valign="top">2(Precision&#x002A;Recall)/(Precision+Recall)</td>
<td align="left" valign="top">Harmonic mean of precision and recall, providing a balanced measure for imbalanced datasets.</td>
<td align="left" valign="top">
<xref ref-type="bibr" rid="ref8">Chicco and Jurman (2020)</xref>
</td>
</tr>
<tr>
<td align="left" valign="top">Accuracy</td>
<td align="left" valign="top">(TP&#x202F;+&#x202F;TN)/(TP&#x202F;+&#x202F;TN&#x202F;+&#x202F;FP&#x202F;+&#x202F;FN)</td>
<td align="left" valign="top">Measures overall correctness of the model across all classes.</td>
<td align="left" valign="top">
<xref ref-type="bibr" rid="ref43">Tharwat (2020)</xref>
</td>
</tr>
<tr>
<td align="left" valign="top">Area under the curve (AUC)</td>
<td align="left" valign="top">Computed from ROC curve</td>
<td align="left" valign="top">Represents the probability that the model ranks a randomly chosen positive instance higher than a randomly chosen negative one.</td>
<td align="left" valign="top">
<xref ref-type="bibr" rid="ref9028">Yang and Ying (2022)</xref>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>TP (True Positives): Correctly predicted positive cases; TN (True Negatives): Correctly predicted negative cases; FP (False Positives): Incorrectly predicted positive cases; FN (False Negatives): Incorrectly predicted negative cases.</p>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec sec-type="results" id="sec10">
<label>3</label>
<title>Results and discussion</title>
<p>The following sections present a comprehensive evaluation of the developed AI models across various experimental phases, including binary classification and multi-class subtype detection for breast cancer biopsy images. Detailed performance metrics such as accuracy, precision, recall, F1-score, and ROC-AUC are reported to assess the effectiveness and generalizability of each model. The study began by analyzing the outcomes of the initial CNN architecture, followed by performance improvements observed with the final DenseNet121-based model. It then extends this evaluation to multi-class classification tasks, detailing the results for both benign and malignant subtypes. These results collectively demonstrate the diagnostic power and clinical viability of the proposed framework.</p>
<sec id="sec11">
<label>3.1</label>
<title>The binary categorical classification as benign or malignant</title>
<p>The initial CNN architecture, developed for binary classification of histopathology images into benign and malignant categories, was evaluated using different filter and kernel size configurations. <xref ref-type="table" rid="tab9">Table 9</xref> presented in this section summarizes the performance metrics for each configuration. While the model achieved a maximum accuracy of 84%, this result was below the expected benchmark for binary classification tasks, where an accuracy of 95% or higher is generally required to maintain robust performance during the transition to multi-class classification.</p>
<table-wrap position="float" id="tab9">
<label>Table 9</label>
<caption>
<p>Initial results of binary classification.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Filters</th>
<th align="center" valign="top">Mesh size</th>
<th align="center" valign="top">Precision</th>
<th align="center" valign="top">Recall</th>
<th align="center" valign="top">Binary accuracy</th>
<th align="center" valign="top">AUC</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">32, 32, 32</td>
<td align="center" valign="top">3&#x00D7;3</td>
<td align="center" valign="top">0.8603</td>
<td align="center" valign="top">0.9283</td>
<td align="center" valign="top">0.8424</td>
<td align="center" valign="top">0.8768</td>
</tr>
<tr>
<td align="left" valign="top">16, 32, 16</td>
<td align="center" valign="top">4&#x00D7;4</td>
<td align="center" valign="top">0.8893</td>
<td align="center" valign="top">0.8994</td>
<td align="center" valign="top">0.8542</td>
<td align="center" valign="top">0.8872</td>
</tr>
<tr>
<td align="left" valign="top">16, 32, 16</td>
<td align="center" valign="top">5&#x00D7;5</td>
<td align="center" valign="top">0.8253</td>
<td align="center" valign="top">0.8465</td>
<td align="center" valign="top">0.7799</td>
<td align="center" valign="top">0.8290</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Upon using the proposed DenseNet121-based binary classification, the model demonstrated outstanding performance in distinguishing between benign and malignant breast cancer histopathology images. The model was trained on out-of-sample biopsy images. As shown in <xref ref-type="table" rid="tab10">Table 10</xref>, the model achieved a test accuracy of 98.63%, indicating a high level of reliability in classification. Precision, recall, and F1-score values were also notably high at 0.9888, 0.9867, and 0.9877, respectively, reflecting the model&#x2019;s ability to correctly identify malignant cases while minimizing false positives and false negatives. The ROC-AUC score of 0.9863 further confirms the model&#x2019;s strong discriminative capability, showing its robustness in distinguishing between the two classes across different classification thresholds. <xref ref-type="fig" rid="fig8">Figure 8</xref> shows the confusion matrix for the classification, while <xref ref-type="fig" rid="fig9">Figure 9</xref> is a graph showing the change in accuracy and loss with respect to epochs. Compared to previous models that relied on shallow convolutional architectures and dataset fragmentation based on magnification levels, the proposed model benefits from multi-scale feature extraction at different network depths, improving its ability to generalize across varying histopathological patterns. By utilizing a 128&#x202F;&#x00D7;&#x202F;128 input image resolution, the model strikes a balance between computational efficiency and feature preservation, allowing for optimal learning without excessive memory consumption. The use of feature fusion across different abstraction levels enhances the classification robustness, ensuring that both low-level structural features and high-level morphological variations contribute to the final decision-making process.</p>
<table-wrap position="float" id="tab10">
<label>Table 10</label>
<caption>
<p>Final results of binary classification algorithm.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Metric</th>
<th align="center" valign="top">Value</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Test accuracy</td>
<td align="center" valign="top">0.9863</td>
</tr>
<tr>
<td align="left" valign="top">Precision</td>
<td align="center" valign="top">0.9888</td>
</tr>
<tr>
<td align="left" valign="top">Recall</td>
<td align="center" valign="top">0.9867</td>
</tr>
<tr>
<td align="left" valign="top">F1-score</td>
<td align="center" valign="top">0.9877</td>
</tr>
<tr>
<td align="left" valign="top">ROC&#x2013;AUC</td>
<td align="center" valign="top">0.9863</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig position="float" id="fig8">
<label>Figure 8</label>
<caption>
<p>Confusion matrix for final binary classification.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g008.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Confusion matrix for a binary model showing true labels versus predicted labels. The matrix has quadrants with the following values: top-left (true negatives) 1467, top-right (false positives) 21, bottom-left (false negatives) 25, bottom-right (true positives) 1852. A blue gradient color bar on the right indicates higher values with darker shades.</alt-text>
</graphic>
</fig>
<fig position="float" id="fig9">
<label>Figure 9</label>
<caption>
<p>Accuracy/loss vs. epoch graph.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g009.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Line chart showing training and validation accuracy and loss over 50 epochs for binary classification. Training accuracy (blue) and validation accuracy (red) both reach near 1.0. Training loss (green) and validation loss (yellow) decrease, with validation loss showing more fluctuation.</alt-text>
</graphic>
</fig>
</sec>
<sec id="sec12">
<label>3.2</label>
<title>Multi-category classification of benign and malignant cancers</title>
<p>The performance of the multi-class classification models for benign and malignant breast cancer subtypes is summarized in <xref ref-type="table" rid="tab11">Table 11</xref>. The model was trained on out-of-sample biopsy images. The model trained for benign classification achieved a test accuracy of 0.9382, whereas the malignant classification model obtained a test accuracy of 0.9158. Additionally, the malignant model demonstrated strong predictive capabilities, with a precision of 0.9151, recall of 0.9240, and an F1-score of 0.9182, while the benign model had a precision of 0.9541, recall of 0.9324 and an F1-score of 0.9415. These results highlight the effectiveness of using separate CNNs for each subtype, allowing for improved feature extraction and discrimination between fine-grained morphological variations. <xref ref-type="fig" rid="fig10">Figure 10</xref> shows the confusion matrix for the benign model, while <xref ref-type="fig" rid="fig11">Figure 11</xref> shows the same for the malignant model. <xref ref-type="fig" rid="fig12">Figure 12</xref> shows the graph of the change in accuracy/loss with respect to the epochs for benign, and <xref ref-type="fig" rid="fig13">Figure 13</xref> shows the same for malignant.</p>
<table-wrap position="float" id="tab11">
<label>Table 11</label>
<caption>
<p>Final results for multi-category classification.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Classification task</th>
<th align="center" valign="top">Accuracy</th>
<th align="center" valign="top">Precision</th>
<th align="center" valign="top">Recall</th>
<th align="center" valign="top">F1-score</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Benign</td>
<td align="center" valign="top">0.9483</td>
<td align="center" valign="top">0.9541</td>
<td align="center" valign="top">0.9324</td>
<td align="center" valign="top">0.9415</td>
</tr>
<tr>
<td align="left" valign="top">Malignant</td>
<td align="center" valign="top">0.9254</td>
<td align="center" valign="top">0.9318</td>
<td align="center" valign="top">0.9193</td>
<td align="center" valign="top">0.9251</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig position="float" id="fig10">
<label>Figure 10</label>
<caption>
<p>Confusion matrix for benign sub-category classification.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g010.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Confusion matrix for a benign model showing true labels versus predicted labels. Values range from zero to six hundred, with notable entries: 247, 648, 186, and 330 along the diagonal, indicating correct predictions. A color gradient from light to dark blue represents frequency.</alt-text>
</graphic>
</fig>
<fig position="float" id="fig11">
<label>Figure 11</label>
<caption>
<p>Confusion matrix for malignant sub-category classification.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g011.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Confusion matrix for a malignant model with true labels on the y-axis and predicted labels on the x-axis. Values: 0:655, 28, 8, 6; 1:37, 319, 7, 0; 2:16, 4, 465, 3; 3:24, 3, 4, 298. Color intensity corresponds to the count, with a scale from 0 to 655.</alt-text>
</graphic>
</fig>
<fig position="float" id="fig12">
<label>Figure 12</label>
<caption>
<p>Accuracy/loss vs. epoch graph for benign subclass model.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g012.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Graph showing training and validation accuracy, and loss over 50 epochs for benign subtype classification. Blue and red lines represent training and validation accuracy stabilizing around 0.9 and 0.8. Green and yellow lines show decreasing loss starting above 1.6, stabilizing around 0.4 and 0.6, respectively.</alt-text>
</graphic>
</fig>
<fig position="float" id="fig13">
<label>Figure 13</label>
<caption>
<p>Accuracy/loss vs. epoch graph for malignant subclass model.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g013.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Plot showing training and validation accuracy for malignant subtype classification over 50 epochs. The blue line represents training accuracy and stabilizes around 1.0. The red line indicates validation accuracy, fluctuating near 0.8. The green line shows loss decreasing, while the yellow line for validation loss fluctuates significantly.</alt-text>
</graphic>
</fig>
<p>The results indicate that the model for benign subtypes outperformed the malignant classification model in terms of overall accuracy. However, the malignant classification model maintained strong recall and F1-score values, demonstrating its ability to correctly identify malignant subtypes while balancing precision. These findings support the hypothesis that training separate models for benign and malignant subtypes allows for more specialized learning, leading to improved classification performance. The integration of DenseNet121 as a feature extractor has further contributed to the model&#x2019;s ability to capture multi-scale morphological patterns in histopathology images.</p>
<sec id="sec13">
<label>3.2.1</label>
<title>Per-class metrics for benign subtype classification</title>
<p>To provide deeper insight into the model&#x2019;s subclass-level performance, we evaluated per-class precision, recall, and F1-score for the benign tumor categories (<xref ref-type="table" rid="tab12">Table 12</xref>). The model achieved excellent performance for adenosis (F1-score: 0.98) and TA (0.96), both of which exhibited high precision and recall. Fibroadenoma showed a near-perfect recall of 0.99 but a slightly lower precision of 0.89, indicating occasional misclassification of other benign types as fibroadenoma. The most challenging class was PT, with a recall of 0.74 and an F1-score of 0.83, suggesting frequent confusion with other fibroepithelial lesions&#x2014;particularly fibroadenoma.</p>
<table-wrap position="float" id="tab12">
<label>Table 12</label>
<caption>
<p>Per-class metrics for benign classes.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Classification task</th>
<th align="center" valign="top">Precision</th>
<th align="center" valign="top">Recall</th>
<th align="center" valign="top">F1-score</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Adenosis</td>
<td align="center" valign="top">0.99</td>
<td align="center" valign="top">0.96</td>
<td align="center" valign="top">0.98</td>
</tr>
<tr>
<td align="left" valign="top">Fibroadenoma</td>
<td align="center" valign="top">0.89</td>
<td align="center" valign="top">0.99</td>
<td align="center" valign="top">0.94</td>
</tr>
<tr>
<td align="left" valign="top">Phyllodes tumor</td>
<td align="center" valign="top">0.94</td>
<td align="center" valign="top">0.74</td>
<td align="center" valign="top">0.83</td>
</tr>
<tr>
<td align="left" valign="top">Tubular adenoma</td>
<td align="center" valign="top">0.98</td>
<td align="center" valign="top">0.94</td>
<td align="center" valign="top">0.96</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>This result aligns with known clinical challenges, as PTs and cellular fibroadenomas often exhibit overlapping histological features, especially under low magnification. Even experienced pathologists can find it difficult to differentiate between these entities, given their shared stromal overgrowth and similar architectural patterns (<xref ref-type="bibr" rid="ref5">Barker et al., 2018</xref>). As such, this subtype boundary represents a meaningful opportunity for AI to provide diagnostic support. Our findings underscore the importance of refining AI models to handle these borderline cases more effectively.</p>
<p>To improve classification in this region of diagnostic uncertainty, future iterations of the model could incorporate additional histological cues beyond morphology alone&#x2014;such as mitotic count, stromal cellularity, or margin assessment, which are often critical in distinguishing PTs from fibroadenomas. Enhanced annotation protocols that focus on these differentiating features, particularly at multiple magnifications, could help reduce misclassification. Overall, while the model demonstrates strong performance in benign subtype differentiation, the results also highlight the necessity of targeted refinement in clinically ambiguous classes like PTs. <xref ref-type="fig" rid="fig14">Figures 14</xref>&#x2013;<xref ref-type="fig" rid="fig17">17</xref> show the confusion matrices for Ductal Carcinoma, Lobular Carcinoma, Mucinous Carcinoma, and Papillary Carcinoma, respectively.</p>
<fig position="float" id="fig14">
<label>Figure 14</label>
<caption>
<p>Adenosis confusion matrix.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g014.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Confusion matrix for adenosis classification. True positives: 247, true negatives: 1229, false positives: 3, false negatives: 9. Predicted and actual labels: adenosis and other.</alt-text>
</graphic>
</fig>
<fig position="float" id="fig15">
<label>Figure 15</label>
<caption>
<p>Fibroadenoma confusion matrix.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g015.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Confusion matrix titled "One-vs-All Confusion Matrix: Fibroadenoma" with actual categories as "Other" and "Fibroadenoma," and predicted categories as "Other" and "Fibroadenoma." It shows values: 741 true negatives, 80 false positives, 10 false negatives, and 657 true positives.</alt-text>
</graphic>
</fig>
<fig position="float" id="fig16">
<label>Figure 16</label>
<caption>
<p>Phyllodes tumour confusion matrix.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g016.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Confusion matrix for Phyllodes Tumor classification. True positives: 171, False positives: 11. True negatives: 1246, False negatives: 60. Shows predicted versus actual classifications for Phyllodes Tumor and Other.</alt-text>
</graphic>
</fig>
<fig position="float" id="fig17">
<label>Figure 17</label>
<caption>
<p>Tubular adenoma confusion matrix.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g017.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Confusion matrix titled "One-vs-All Confusion Matrix: Tubular Adenoma" with actual values on the vertical axis and predicted values on the horizontal axis. The matrix shows true negatives: 1149, false positives: 5, false negatives: 20, and true positives: 314.</alt-text>
</graphic>
</fig>
</sec>
<sec id="sec14">
<label>3.2.2</label>
<title>Per-class metrics for malignant subtype classification</title>
<p>For malignant tumor classification, per-class evaluation similarly revealed strong and balanced performance (see <xref ref-type="table" rid="tab13">Table 13</xref>). The model classified mucinous carcinoma and PC with the highest F1-scores (both 0.95), supported by near-perfect precision and recall. DC, the most common subtype, achieved an F1-score of 0.91 with strong precision (0.93) and recall (0.89). The relatively lower performance was observed for LC, with an F1-score of 0.86, primarily due to reduced precision (0.83), indicating overlap in predicted labels with other malignant types. These distinctions are clearly reflected in the one-vs-all confusion matrices, which reveal class-specific prediction errors and highlight the model&#x2019;s ability to distinguish even closely related malignancies. This subclass-level analysis enhances transparency and helps identify where architectural or dataset-level adjustments may be necessary for further improvement. <xref ref-type="fig" rid="fig18">Figures 18</xref>&#x2013;<xref ref-type="fig" rid="fig21">21</xref> show the confusion matrices for Ductal Carcinoma, Lobular Carcinoma, Mucinous Carcinoma, and Papillary Carcinoma, respectively.</p>
<table-wrap position="float" id="tab13">
<label>Table 13</label>
<caption>
<p>Per-class metrics for malignant classes.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Classification task</th>
<th align="center" valign="top">Precision</th>
<th align="center" valign="top">Recall</th>
<th align="center" valign="top">F1-score</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Ductal carcinoma</td>
<td align="center" valign="top">0.93</td>
<td align="center" valign="top">0.89</td>
<td align="center" valign="top">0.91</td>
</tr>
<tr>
<td align="left" valign="top">Lobular carcinoma</td>
<td align="center" valign="top">0.83</td>
<td align="center" valign="top">0.90</td>
<td align="center" valign="top">0.86</td>
</tr>
<tr>
<td align="left" valign="top">Mucinous carcinoma</td>
<td align="center" valign="top">0.93</td>
<td align="center" valign="top">0.97</td>
<td align="center" valign="top">0.95</td>
</tr>
<tr>
<td align="left" valign="top">Papillary carcinoma</td>
<td align="center" valign="top">0.99</td>
<td align="center" valign="top">0.92</td>
<td align="center" valign="top">0.95</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig position="float" id="fig18">
<label>Figure 18</label>
<caption>
<p>Ductal carcinoma confusion matrix.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g018.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Confusion matrix for ductal carcinoma classification. True positives: 620, false negatives: 77, true negatives: 1136, false positives: 44. Columns represent predictions; rows represent actuals.</alt-text>
</graphic>
</fig>
<fig position="float" id="fig19">
<label>Figure 19</label>
<caption>
<p>Lobular carcinoma confusion matrix.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g019.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Confusion matrix titled "One-vs-All Confusion Matrix: Lobular Carcinoma." It displays actual versus predicted values. True negatives: 1446, false positives: 68, false negatives: 36, true positives: 327.</alt-text>
</graphic>
</fig>
<fig position="float" id="fig20">
<label>Figure 20</label>
<caption>
<p>Mucinous carcinoma confusion matrix.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g020.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Confusion matrix for Mucinous Carcinoma classification. True positives: 475, false positives: 35, true negatives: 1,354, false negatives: 13. Used for assessing model performance.</alt-text>
</graphic>
</fig>
<fig position="float" id="fig21">
<label>Figure 21</label>
<caption>
<p>Papillary carcinoma confusion matrix.</p>
</caption>
<graphic xlink:href="frai-08-1627876-g021.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Confusion matrix titled "One-vs-All Confusion Matrix: Papillary Carcinoma" with rows labeled "Other" and "Papillary Carcinoma" against columns labeled "Other" and "Papillary Carcinoma." Values are 1544 (true negatives), 4 (false positives), 25 (false negatives), and 304 (true positives).</alt-text>
</graphic>
</fig>
<p>Although this study primarily focuses on classification performance, model interpretability is essential for clinical adoption. Techniques such as gradient-weighted class activation mapping (grad-CAM) or SHapley additive exPlanations (SHAP) can be used to generate heatmaps that visualize which regions of a biopsy image most influenced the model&#x2019;s predictions. These visual explanations could help verify whether the model is focusing on diagnostically relevant features&#x2014;Such as nuclear pleomorphism, stromal arrangement, or mitotic activity&#x2014;Thereby increasing clinician trust and facilitating integration into diagnostic workflows.</p>
</sec>
</sec>
<sec id="sec15">
<label>3.3</label>
<title>K-fold cross-validation and performance stability</title>
<p>To evaluate the generalizability of the proposed classifiers, we performed 5-fold cross-validation on each of the three tasks: binary classification, benign subtype classification, and malignant subtype classification. The mean values, standard deviations, and 95% confidence intervals are presented in <xref ref-type="table" rid="tab14">Table 14</xref>.</p>
<table-wrap position="float" id="tab14">
<label>Table 14</label>
<caption>
<p>Cross-validation scores for each model.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Task</th>
<th align="left" valign="top">Metric</th>
<th align="center" valign="top">Score (&#x00B1;SD)</th>
<th align="center" valign="top">95% CI</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top" rowspan="4">Binary class</td>
<td align="left" valign="top">Accuracy</td>
<td align="center" valign="top">0.9712&#x202F;&#x00B1;&#x202F;0.0069</td>
<td align="center" valign="top">0.9616&#x2013;0.9808</td>
</tr>
<tr>
<td align="left" valign="top">Precision</td>
<td align="center" valign="top">0.9874&#x202F;&#x00B1;&#x202F;0.0064</td>
<td align="center" valign="top">0.9785&#x2013;0.9964</td>
</tr>
<tr>
<td align="left" valign="top">Recall</td>
<td align="center" valign="top">0.9608&#x202F;&#x00B1;&#x202F;0.0160</td>
<td align="center" valign="top">0.9386&#x2013;0.9830</td>
</tr>
<tr>
<td align="left" valign="top">F1-score</td>
<td align="center" valign="top">0.9738&#x202F;&#x00B1;&#x202F;0.0062</td>
<td align="center" valign="top">0.9652&#x2013;0.9825</td>
</tr>
<tr>
<td align="left" valign="top" rowspan="4">Benign subtype</td>
<td align="left" valign="top">Accuracy</td>
<td align="center" valign="top">0.9375&#x202F;&#x00B1;&#x202F;0.0208</td>
<td align="center" valign="top">0.9087&#x2013;0.9663</td>
</tr>
<tr>
<td align="left" valign="top">Precision</td>
<td align="center" valign="top">0.9468&#x202F;&#x00B1;&#x202F;0.0132</td>
<td align="center" valign="top">0.9284&#x2013;0.9651</td>
</tr>
<tr>
<td align="left" valign="top">Recall</td>
<td align="center" valign="top">0.9300&#x202F;&#x00B1;&#x202F;0.0274</td>
<td align="center" valign="top">0.8919&#x2013;0.9680</td>
</tr>
<tr>
<td align="left" valign="top">F1-score</td>
<td align="center" valign="top">0.9369&#x202F;&#x00B1;&#x202F;0.0203</td>
<td align="center" valign="top">0.9087&#x2013;0.9650</td>
</tr>
<tr>
<td align="left" valign="top" rowspan="4">Malignant subtype</td>
<td align="left" valign="top">Accuracy</td>
<td align="center" valign="top">0.9202&#x202F;&#x00B1;&#x202F;0.0080</td>
<td align="center" valign="top">0.9091&#x2013;0.9314</td>
</tr>
<tr>
<td align="left" valign="top">Precision</td>
<td align="center" valign="top">0.9048&#x202F;&#x00B1;&#x202F;0.0122</td>
<td align="center" valign="top">0.8878&#x2013;0.9217</td>
</tr>
<tr>
<td align="left" valign="top">Recall</td>
<td align="center" valign="top">0.8808&#x202F;&#x00B1;&#x202F;0.0265</td>
<td align="center" valign="top">0.8440&#x2013;0.9176</td>
</tr>
<tr>
<td align="left" valign="top">F1-score</td>
<td align="center" valign="top">0.8907&#x202F;&#x00B1;&#x202F;0.0157</td>
<td align="center" valign="top">0.8689&#x2013;0.9125</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As shown in <xref ref-type="table" rid="tab14">Table 14</xref>, the proposed classifiers maintained strong and stable performance across 5-fold cross-validation, with F1-scores of 0.9738 for binary classification, 0.9369 for benign subtype classification, and 0.8907 for malignant subtypes. The relatively tight confidence intervals and low standard deviations across all tasks suggest minimal sensitivity to training split variability, indicating good model generalization. Notably, the benign subtype classifier showed slightly higher performance consistency than the malignant subtype model, which exhibited greater variance in recall (&#x00B1;0.0265) and a broader 95% CI (0.8440&#x2013;0.9176). This is likely due to the greater morphological diversity and class imbalance among malignant categories&#x2014;particularly LC&#x2014;where lower sample size and higher inter-class overlap may amplify prediction uncertainty.</p>
<p>The observed decline in recall and F1-scores compared to the single train-test split is expected, and it reflects the increased rigor of cross-validation, which exposes the model to harder-to-classify samples in more diverse folds. From a diagnostic perspective, the relatively stable precision across tasks implies that the model is conservative in positive predictions, minimizing false positives&#x2014;an important trait in histopathological screening scenarios. However, the slightly larger variation in recall for rare or borderline subtypes warrants attention, as it may affect sensitivity to clinically ambiguous cases.</p>
<p>These findings reinforce the utility of k-fold cross-validation as a stress-testing tool, revealing not just average performance but also subtype-specific reliability under different sampling conditions. Future iterations of the model could incorporate stratified sampling by subtype or uncertainty-aware training strategies to further reduce variability in underrepresented classes.</p>
</sec>
<sec id="sec16">
<label>3.4</label>
<title>Final notes</title>
<p>The final model, built on DenseNet121 with multi-scale feature extraction, demonstrated the strongest performance among all evaluated architectures, achieving high accuracy, recall, and F1-scores in the binary classification of breast cancer biopsy images. By extracting and fusing features from three intermediate convolutional blocks, the model was able to learn both fine-grained cellular structures and broader tissue-level patterns relevant to histopathological classification. The incorporation of L2 normalization, dropout regularization, and batch normalization within each branch further contributed to its stability and generalization. This design effectively addressed limitations observed in earlier models, such as feature loss due to single-layer reliance or overfitting from insufficient regularization.</p>
<p>Compared to existing models evaluated on the same BreaKHis dataset, the proposed approach offers a substantial improvement. <xref ref-type="bibr" rid="ref9003">Ara&#x00FA;jo et al. (2017)</xref> employed a standard CNN and achieved an average accuracy of 83.3% across all magnification levels. Similarly, <xref ref-type="bibr" rid="ref9024">Vo et al. (2019)</xref> reported 86.3% accuracy using a handcrafted feature pipeline with SVMs, while <xref ref-type="bibr" rid="ref9009">Bayramoglu et al. (2016)</xref> applied a multi-scale CNN and reached approximately 88.0% accuracy. More recently, <xref ref-type="bibr" rid="ref9006">Munshi et al. (2024)</xref> proposed an ensemble CNN-SVM framework with XAI components, achieving 94.2% accuracy and an F1-score of 0.93. In contrast, our DenseNet121-based model consistently achieved &#x003E;97% accuracy and an F1-score of 0.9738 in binary classification, while maintaining robust generalization across folds and magnifications. These improvements can be attributed to the deeper architectural depth, feature reuse enabled by dense connectivity, and the incorporation of multi-scale fusion, which allowed the model to capture both cellular and architectural histological patterns more effectively.</p>
<p>The model&#x2019;s strong and consistent performance across both binary and subtype classification tasks positions it as a scalable and technically rigorous approach for digital histopathology. With additional validation on multi-institutional and heterogeneous datasets, this framework has the potential to contribute meaningfully to diagnostic support systems for breast cancer, especially in settings with limited expert pathologist availability.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="sec17">
<label>4</label>
<title>Conclusion</title>
<p>This study presents a deep learning framework for breast cancer diagnosis that performs subtype-level classification directly from H&#x0026;E-stained biopsy images using a DenseNet121-based multi-scale feature fusion architecture. In conventional diagnostic workflows, the determination of histological subtypes often requires additional procedures&#x2014;such as immunohistochemistry (IHC), molecular assays, or serial imaging&#x2014;that increase diagnostic latency, invasiveness, and healthcare costs (<xref ref-type="bibr" rid="ref34">Robbins et al., 2010</xref>; <xref ref-type="bibr" rid="ref29">National Cancer Institute, 2021</xref>; <xref ref-type="bibr" rid="ref26">Manning et al., 2018</xref>). In contrast, our proposed model offers a streamlined, image-only approach capable of distinguishing both benign and malignant lesions and further subclassifying them into clinically relevant subtypes (<xref ref-type="bibr" rid="ref41">Spanhol et al., 2016</xref>; <xref ref-type="bibr" rid="ref33">Rakhlin et al., 2018</xref>).</p>
<p>Our novel methodological contributions allow the model to capture both fine-grained cellular features and broader tissue-level patterns in a unified, end-to-end framework. Notably, this design improves over standard transfer learning approaches that rely solely on final-layer representations or require external fusion mechanisms (<xref ref-type="bibr" rid="ref11">Gupta and Bhavsar, 2018</xref>; <xref ref-type="bibr" rid="ref51">Zhu et al., 2019</xref>; <xref ref-type="bibr" rid="ref47">Wakili et al., 2022</xref>).</p>
<p>While the model achieved strong binary classification accuracy (98.63%) and high F1-scores in multi-class subtype tasks (benign: 0.9415, malignant: 0.9251), its clinical utility is best positioned in supporting difficult diagnostic cases&#x2014;such as distinguishing PTs from fibroadenomas or mucinous from PCs&#x2014;where visual overlap challenges even experienced pathologists (<xref ref-type="bibr" rid="ref5">Barker et al., 2018</xref>; <xref ref-type="bibr" rid="ref44">Thompson et al., 2019</xref>). The model&#x2019;s robustness across 5-fold cross-validation further reinforces its potential reliability in diverse settings (<xref ref-type="bibr" rid="ref16">Kohavi, 1995</xref>; <xref ref-type="bibr" rid="ref7">Berrar, 2019</xref>).</p>
<p>We emphasize that this tool is not a replacement for pathologists but a potential assistive technology to aid in triaging, second-opinion support, and prioritizing ambiguous cases. In resource-limited settings where access to expert pathology or advanced molecular testing is constrained, such a tool could help reduce diagnostic bottlenecks and improve equity in care delivery (<xref ref-type="bibr" rid="ref27">McKinney et al., 2020</xref>; <xref ref-type="bibr" rid="ref48">WHO, 2021</xref>).</p>
<p>However, further validation on larger and more heterogeneous datasets is necessary before clinical deployment (<xref ref-type="bibr" rid="ref23">Liew et al., 2021</xref>). Moreover, the integration of model interpretability mechanisms&#x2014;such as Grad-CAM visualizations&#x2014;remains an essential next step to enhance transparency and foster clinical trust (<xref ref-type="bibr" rid="ref49">Xie et al., 2020</xref>; <xref ref-type="bibr" rid="ref21">Lee et al., 2024</xref>).</p>
</sec>
<sec id="sec18">
<label>5</label>
<title>Limitations and future work</title>
<p>While the proposed model demonstrates strong performance in both binary and subtype-level breast cancer classification, several limitations must be acknowledged. First, the model was trained and validated solely on the BreaKHis dataset, which&#x2014;despite its popularity in computational pathology research&#x2014;is limited in scale and diversity. Its 7,909 images come from only 82 patients, restricting the model&#x2019;s generalizability across varied populations, staining protocols, and imaging equipment. Future studies should aim to validate the model on larger, multi-institutional datasets that capture real-world clinical variability to increase the generalizability of the study.</p>
<p>Second, the model&#x2019;s predictions are based exclusively on morphological features extracted from H&#x0026;E-stained slides and do not incorporate molecular information, such as Human Epidermal Growth Factor Receptor 2 (HER2), Estrogen Receptor / Progesterone Receptor (ER/PR), or triple-negative status. These receptor-level biomarkers are critical for treatment planning, and their exclusion limits the clinical applicability of the system. Expanding the model to include IHC images or genomic profiles would enable a more comprehensive diagnostic tool aligned with current oncology workflows.</p>
<p>Another limitation lies in the model&#x2019;s use of single-magnification images during training, despite the fact that pathologists typically examine biopsies at multiple magnification levels to evaluate both cellular and tissue-level structures. Although our multi-scale feature extraction within DenseNet121 captures some hierarchical information, it does not replicate the diagnostic reasoning derived from viewing across magnifications. Future work could explore multi-resolution input strategies or hierarchical CNNs to better reflect clinical interpretation.</p>
<p>Moreover, the current model lacks interpretability features, which are increasingly essential for clinical integration. Tools such as Grad-CAM or SHAP could be used to highlight regions of interest in biopsy images, helping pathologists understand model decisions and assess reliability. Including these visual explanations would significantly enhance trust and transparency in real-world applications.</p>
<p>Additionally, while this study emphasizes a novel architecture, it does not provide comparative results against widely used CNN baselines such as VGG16, ResNet50, or EfficientNet, nor does it present ablation experiments isolating the impact of architectural components like L2 normalization or feature fusion. Including these comparisons would strengthen claims of architectural innovation and clarify which design elements drive performance gains.</p>
<p>Finally, deployment considerations remain speculative. The model has not yet been tested in live clinical settings or integrated into diagnostic workflows, where computational constraints, system latency, and compatibility with laboratory information systems pose practical challenges. Future work should also consider automating the augmentation pipeline&#x2014;currently hand-tuned&#x2014;using approaches such as AutoAugment or Generative Adversarial Network (GAN)-based synthesis to improve performance in rare subtypes and low-data scenarios.</p>
<p>By addressing these limitations through external validation, multimodal expansion, improved interpretability, and clinical simulation, this work can move closer to real-world deployment as a reliable assistive tool in digital breast pathology.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="sec19">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material; further inquiries can be directed to the corresponding author.</p>
</sec>
<sec sec-type="author-contributions" id="sec20">
<title>Author contributions</title>
<p>NC: Conceptualization, Investigation, Methodology, Software, Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing. ZD: Conceptualization, Investigation, Methodology, Software, Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing.</p>
</sec>
<sec sec-type="funding-information" id="sec21">
<title>Funding</title>
<p>The author(s) declare that no financial support was received for the research and/or publication of this article.</p>
</sec>
<sec sec-type="COI-statement" id="sec22">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="ai-statement" id="sec23">
<title>Generative AI statement</title>
<p>The authors declare that no Gen AI was used in the creation of this manuscript.</p>
<p>Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.</p>
</sec>
<sec sec-type="disclaimer" id="sec24">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="ref9001"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aldakhil</surname><given-names>L. A.</given-names></name> <name><surname>Alhasson</surname><given-names>H. F.</given-names></name> <name><surname>Alharbi</surname><given-names>S. S</given-names></name></person-group>. (<year>2024</year>). <article-title>EfficientNetV2 + CBAM for breast histology classification</article-title>. <source>Diagnostics</source>. <volume>14</volume>:<fpage>56</fpage>. doi: <pub-id pub-id-type="doi">10.3390/diagnostics14010056</pub-id></citation></ref>
<ref id="ref1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alzubaidi</surname><given-names>L.</given-names></name> <name><surname>Zhang</surname><given-names>J.</given-names></name> <name><surname>Humaidi</surname><given-names>A. J.</given-names></name> <name><surname>Al-Dujaili</surname><given-names>A.</given-names></name> <name><surname>Al-Shamma</surname><given-names>O.</given-names></name> <name><surname>Santamar&#x00ED;a</surname><given-names>J.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Review of deep learning: concepts, CNN architectures, challenges, applications, future directions</article-title>. <source>J. Big Data</source> <volume>8</volume>:<fpage>53</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s40537-021-00444-8</pub-id>, PMID: <pub-id pub-id-type="pmid">33816053</pub-id></citation></ref>
<ref id="ref2"><citation citation-type="other"><person-group person-group-type="author"><collab id="coll1">American Cancer Society</collab></person-group>. (<year>2019</year>). <italic>Types of breast biopsies</italic>. Available online at: <ext-link xlink:href="https://www.cancer.org/cancer/breast-cancer/screening-tests-and-early-detection/breast-biopsy/types-of-biopsies.html" ext-link-type="uri">https://www.cancer.org/cancer/breast-cancer/screening-tests-and-early-detection/breast-biopsy/types-of-biopsies.html</ext-link> (Accessed January 23, 2025).</citation></ref>
<ref id="ref9003"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ara&#x00FA;jo</surname><given-names>T.</given-names></name> <name><surname>Aresta</surname><given-names>G.</given-names></name> <name><surname>Castro</surname><given-names>E.</given-names></name> <name><surname>Rouco</surname><given-names>J.</given-names></name> <name><surname>Aguiar</surname><given-names>P.</given-names></name> <name><surname>Eloy</surname><given-names>C.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Classification of breast cancer histology images using Convolutional Neural Networks</article-title>. <source>PLoS One</source>, <volume>12</volume>:<fpage>e0177544</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0177544</pub-id></citation></ref>
<ref id="ref3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Arevalo</surname><given-names>J.</given-names></name> <name><surname>Gonz&#x00E1;lez</surname><given-names>F. A.</given-names></name> <name><surname>Ramos-Poll&#x00E1;n</surname><given-names>R.</given-names></name> <name><surname>Oliveira</surname><given-names>J. L.</given-names></name> <name><surname>Guevara Lopez</surname><given-names>M. A.</given-names></name></person-group> (<year>2016</year>). <article-title>Representation learning for mammography mass lesion classification with convolutional neural networks</article-title>. <source>Comput. Methods Prog. Biomed.</source> <volume>127</volume>, <fpage>248</fpage>&#x2013;<lpage>257</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.cmpb.2015.12.014</pub-id>, PMID: <pub-id pub-id-type="pmid">26826901</pub-id></citation></ref>
<ref id="ref9004"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Atrey</surname><given-names>K.</given-names></name> <name><surname>Singh</surname><given-names>B. K.</given-names></name> <name><surname>Bodhey</surname><given-names>N. K.</given-names></name></person-group> (<year>2024</year>). <article-title>Multimodal breast cancer classification with mammogram and ultrasound fusion using ResNet18 and SVM</article-title>. <source>J. Med. Imaging.</source> <volume>11</volume>:<fpage>021405</fpage>. doi: <pub-id pub-id-type="doi">10.1117/1.JMI.11.2.021405</pub-id></citation></ref>
<ref id="ref4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bandi</surname><given-names>P.</given-names></name> <name><surname>Geessink</surname><given-names>O.</given-names></name> <name><surname>Manson</surname><given-names>Q.</given-names></name> <name><surname>Van Dijk</surname><given-names>M. C.</given-names></name> <name><surname>Balkenhol</surname><given-names>M.</given-names></name> <name><surname>Hermsen</surname><given-names>M.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>From detection of individual metastases to classification of lymph node status at the patient level: the CAMELYON17 challenge</article-title>. <source>IEEE Trans. Med. Imaging</source> <volume>37</volume>, <fpage>2613</fpage>&#x2013;<lpage>2624</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TMI.2018.2867350</pub-id></citation></ref>
<ref id="ref5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barker</surname><given-names>E.</given-names></name> <name><surname>Gosein</surname><given-names>M.</given-names></name> <name><surname>Kulkarni</surname><given-names>S.</given-names></name> <name><surname>Jones</surname><given-names>S.</given-names></name></person-group> (<year>2018</year>). <article-title>Phyllodes tumors of the breast: a comprehensive review of clinical and histopathological features</article-title>. <source>J. Pathol. Transl. Med.</source> <volume>52</volume>, <fpage>221</fpage>&#x2013;<lpage>229</lpage>.</citation></ref>
<ref id="ref9009"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bayramoglu</surname><given-names>N.</given-names></name> <name><surname>Kannala</surname><given-names>J.</given-names></name> <name><surname>Heikkil&#x00E4;</surname><given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep learning for magnification-independent breast cancer histopathology image classification</article-title>. <source>Int. Conf. Pattern Recognit.</source> <fpage>2440</fpage>&#x2013;<lpage>2445</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ICPR.2016.7900002</pub-id></citation></ref>
<ref id="ref6"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Bayramoglu</surname><given-names>N.</given-names></name> <name><surname>Kannala</surname><given-names>J.</given-names></name> <name><surname>Heikkil&#x00E4;</surname><given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Deep learning for magnification-independent breast cancer histopathology image classification</article-title>. In <conf-name>Proceedings of the IEEE International Conference on Biomedical Imaging</conf-name> (pp. <fpage>914</fpage>&#x2013;<lpage>918</lpage>).</citation></ref>
<ref id="ref7"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Berrar</surname><given-names>D.</given-names></name></person-group> (<year>2019</year>). &#x201C;<article-title>Cross-validation</article-title>&#x201D; in <source>Encyclopedia of bioinformatics and computational biology</source>. ed. <person-group person-group-type="editor"><name><surname>Rouse</surname><given-names>M.</given-names></name></person-group> (<publisher-loc>Oxford, United Kingdom</publisher-loc>: <publisher-name>Elsevier</publisher-name>), <fpage>542</fpage>&#x2013;<lpage>545</lpage>.</citation></ref>
<ref id="ref8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chicco</surname><given-names>D.</given-names></name> <name><surname>Jurman</surname><given-names>G.</given-names></name></person-group> (<year>2020</year>). <article-title>The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation</article-title>. <source>BMC Genomics</source> <volume>21</volume>:<fpage>6</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s12864-019-6413-7</pub-id>, PMID: <pub-id pub-id-type="pmid">31898477</pub-id></citation></ref>
<ref id="ref9013"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Esteva</surname><given-names>A.</given-names></name> <name><surname>Chou</surname><given-names>K.</given-names></name> <name><surname>Yeung</surname><given-names>S.</given-names></name> <name><surname>Naik</surname><given-names>N.</given-names></name> <name><surname>Madani</surname><given-names>A.</given-names></name> <name><surname>Mottaghi</surname><given-names>A.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>A guide to deep learning in healthcare</article-title>. <source>Comput. Biol. Med.</source> <volume>27</volume>, <fpage>24</fpage>&#x2013;<lpage>29</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41591-020-1122-4</pub-id></citation></ref>
<ref id="ref9012"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ezzat</surname><given-names>D.</given-names></name> <name><surname>Hassanien</surname><given-names>A. E.</given-names></name></person-group> (<year>2023</year>). <article-title>Optimized Bayesian CNN (OBCNN) with ResNet101V2 and Monte Carlo dropout for breast cancer diagnosis</article-title>. <source>Comput. Biol. Med.</source> <volume>152</volume>:<fpage>106352</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.cmpb.2023.106352</pub-id></citation></ref>
<ref id="ref9014"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fu</surname><given-names>A.</given-names></name> <name><surname>Yao</surname><given-names>B.</given-names></name> <name><surname>Dong</surname><given-names>T.</given-names></name> <name><surname>Chen</surname><given-names>Y.</given-names></name> <name><surname>Yao</surname><given-names>J.</given-names></name> <name><surname>Liu</surname><given-names>Y.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Tumor-resident intracellular microbiota promotes metastatic colonization in breast cancer</article-title>. <source>Cell.</source> <volume>185</volume>, <fpage>1356</fpage>&#x2013;<lpage>1372.e26</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.cell.2022.02.027</pub-id></citation></ref>
<ref id="ref9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gandomkar</surname><given-names>Z.</given-names></name> <name><surname>Brennan</surname><given-names>P. C.</given-names></name> <name><surname>Mello-Thoms</surname><given-names>C.</given-names></name></person-group> (<year>2018</year>). <article-title>Computer-based image analysis in breast pathology: a review of current methods and future directions</article-title>. <source>Breast</source> <volume>42</volume>, <fpage>56</fpage>&#x2013;<lpage>67</lpage>.</citation></ref>
<ref id="ref9016"><citation citation-type="book"><person-group person-group-type="author"><collab>GLOBOCAN</collab></person-group>. (<year>2020</year>). <source>Global cancer statistics: Breast cancer incidence and mortality worldwide</source>. <publisher-loc>International Agency for Research on Cancer (IARC)</publisher-loc>. <publisher-name>Available</publisher-name> at: <ext-link xlink:href="https://gco.iarc.fr/today/home" ext-link-type="uri">https://gco.iarc.fr/today/home</ext-link></citation></ref>
<ref id="ref10"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Glorot</surname><given-names>X.</given-names></name> <name><surname>Bordes</surname><given-names>A.</given-names></name> <name><surname>Bengio</surname><given-names>Y.</given-names></name></person-group> (<year>2011</year>). <article-title>Deep sparse rectifier neural networks</article-title>. In <conf-name>Proceedings of the 14th International Conference on Artificial Intelligence and Statistics</conf-name> (pp. <fpage>315</fpage>&#x2013;<lpage>323</lpage>).</citation></ref>
<ref id="ref11"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Gupta</surname><given-names>A.</given-names></name> <name><surname>Bhavsar</surname><given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Breast cancer histopathological image classification: is magnification important?</article-title> In <conf-name>2018 IEEE 15th international symposium on biomedical imaging (ISBI), IEEE</conf-name> (pp. <fpage>1110</fpage>&#x2013;<lpage>1113</lpage>)</citation></ref>
<ref id="ref12"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>He</surname><given-names>K.</given-names></name> <name><surname>Zhang</surname><given-names>X.</given-names></name> <name><surname>Ren</surname><given-names>S.</given-names></name> <name><surname>Sun</surname><given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep residual learning for image recognition</article-title>. In <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name> (pp. <fpage>770</fpage>&#x2013;<lpage>778</lpage>).</citation></ref>
<ref id="ref13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Houssein</surname><given-names>E. H.</given-names></name> <name><surname>Hosney</surname><given-names>M. E.</given-names></name> <name><surname>Zaher</surname><given-names>A. S.</given-names></name></person-group> (<year>2021</year>). <article-title>Particle swarm optimization for deep learning hyperparameter tuning in medical image classification</article-title>. <source>Neural Comput. &#x0026; Applic.</source> <volume>33</volume>, <fpage>13359</fpage>&#x2013;<lpage>13378</lpage>.</citation></ref>
<ref id="ref9018"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname><given-names>G.</given-names></name> <name><surname>Liu</surname><given-names>Z.</given-names></name> <name><surname>Van Der Maaten</surname><given-names>L.</given-names></name> <name><surname>Weinberger</surname><given-names>K. Q.</given-names></name></person-group> (<year>2017</year>). <article-title>Densely connected convolutional networks</article-title>. <source>Proceedings of CVPR.</source> <volume>2017</volume>, <fpage>4700</fpage>&#x2013;<lpage>4708</lpage>. doi: <pub-id pub-id-type="doi">10.1109/CVPR.2017.243</pub-id></citation></ref>
<ref id="ref14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname><given-names>Y.</given-names></name> <name><surname>Song</surname><given-names>M.</given-names></name> <name><surname>Zhang</surname><given-names>X.</given-names></name> <name><surname>Lin</surname><given-names>Y.</given-names></name></person-group> (<year>2017</year>). <article-title>The diagnostic challenges of adenosis in breast pathology</article-title>. <source>Pathol. Res. Pract.</source> <volume>213</volume>, <fpage>1</fpage>&#x2013;<lpage>9</lpage>.</citation></ref>
<ref id="ref9019"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Joseph</surname><given-names>S.</given-names></name> <name><surname>Kumar</surname><given-names>S.</given-names></name> <name><surname>Rajesh</surname><given-names>R.</given-names></name> <name><surname>Thomas</surname><given-names>J.</given-names></name> <name><surname>Varghese</surname><given-names>R.</given-names></name> <name><surname>Nair</surname><given-names>S.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Multi-classification of breast cancer using handcrafted and deep features from BreakHis</article-title>. <source>Pattern Recogn. Lett.</source> <volume>161</volume>, <fpage>76</fpage>&#x2013;<lpage>83</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.patrec.2022.01.004</pub-id></citation></ref>
<ref id="ref9020"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kaymak</surname><given-names>S.</given-names></name> <name><surname>Kandemir</surname><given-names>M.</given-names></name> <name><surname>Zhang</surname><given-names>C.</given-names></name> <name><surname>Caputo</surname><given-names>B.</given-names></name> <name><surname>Hamprecht</surname><given-names>F. A.</given-names></name> <name><surname>Lampert</surname><given-names>C. H.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Neural network classification of breast cancer histopathology images using discrete Haar wavelets</article-title>. <source>Biomed. Res. Int.</source> <volume>2017</volume>, <fpage>1</fpage>&#x2013;<lpage>9</lpage>. doi: <pub-id pub-id-type="doi">10.1155/2017/5179020</pub-id></citation></ref>
<ref id="ref15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kode</surname><given-names>R.</given-names></name> <name><surname>Barkana</surname><given-names>B. D.</given-names></name></person-group> (<year>2023</year>). <article-title>Analysis of CNN performance in classifying BreakHis breast histopathology images under varied dropout configurations</article-title>. <source>Biomed. Signal Process. Control</source> <volume>82</volume>:<fpage>104534</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.bspc.2023.104534</pub-id></citation></ref>
<ref id="ref16"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Kohavi</surname><given-names>R.</given-names></name></person-group> (<year>1995</year>). <article-title>A study of cross-validation and bootstrap for accuracy estimation and model selection</article-title>. <conf-name>Proceedings of the 14th International Joint Conference on Artificial Intelligence</conf-name>, <fpage>1137</fpage>&#x2013;<lpage>1143</lpage>.</citation></ref>
<ref id="ref17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kolb</surname><given-names>T. M.</given-names></name> <name><surname>Lichy</surname><given-names>J.</given-names></name> <name><surname>Newhouse</surname><given-names>J. H.</given-names></name></person-group> (<year>2002</year>). <article-title>Comparison of the performance of screening mammography, physical examination, and breast US and evaluation of factors that influence them: an analysis of 27,825 patient evaluations</article-title>. <source>Radiology</source> <volume>225</volume>, <fpage>165</fpage>&#x2013;<lpage>175</lpage>. doi: <pub-id pub-id-type="doi">10.1148/radiol.2251011667</pub-id>, PMID: <pub-id pub-id-type="pmid">12355001</pub-id></citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krizhevsky</surname><given-names>A.</given-names></name> <name><surname>Sutskever</surname><given-names>I.</given-names></name> <name><surname>Hinton</surname><given-names>G. E.</given-names></name></person-group> (<year>2012</year>). <article-title>ImageNet classification with deep convolutional neural networks</article-title>. <source>Adv. Neural Inf. Proces. Syst.</source> <volume>25</volume>, <fpage>1097</fpage>&#x2013;<lpage>1105</lpage>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.1207.0580</pub-id></citation></ref>
<ref id="ref19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuhl</surname><given-names>C. K.</given-names></name> <name><surname>Schrading</surname><given-names>S.</given-names></name> <name><surname>Leutner</surname><given-names>C. C.</given-names></name> <name><surname>Morakkabati-Spitz</surname><given-names>N.</given-names></name> <name><surname>Wardelmann</surname><given-names>E.</given-names></name> <name><surname>Fimmers</surname><given-names>R.</given-names></name> <etal/></person-group>. (<year>2005</year>). <article-title>MRI for diagnosis of pure ductal carcinoma in situ: a prospective observational study</article-title>. <source>Lancet Oncol.</source> <volume>6</volume>, <fpage>433</fpage>&#x2013;<lpage>441</lpage>. doi: <pub-id pub-id-type="doi">10.1016/S1470-2045(05)70148-0</pub-id></citation></ref>
<ref id="ref20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname><given-names>Y.</given-names></name> <name><surname>Bottou</surname><given-names>L.</given-names></name> <name><surname>Bengio</surname><given-names>Y.</given-names></name> <name><surname>Haffner</surname><given-names>P.</given-names></name></person-group> (<year>1998</year>). <article-title>Gradient-based learning applied to document recognition</article-title>. <source>Proc. IEEE</source> <volume>86</volume>, <fpage>2278</fpage>&#x2013;<lpage>2324</lpage>. doi: <pub-id pub-id-type="doi">10.1109/5.726791</pub-id></citation></ref>
<ref id="ref21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname><given-names>H.</given-names></name> <name><surname>Banerjee</surname><given-names>S.</given-names></name> <name><surname>Choi</surname><given-names>Y.</given-names></name></person-group> (<year>2024</year>). <article-title>AI-powered interpretability in breast cancer subtype detection</article-title>. <source>Insights Imaging</source> <volume>15</volume>:<fpage>227</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s13244-024-01810-9</pub-id>, PMID: <pub-id pub-id-type="pmid">39320560</pub-id></citation></ref>
<ref id="ref22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lehman</surname><given-names>C. D.</given-names></name> <name><surname>Wellman</surname><given-names>R. D.</given-names></name> <name><surname>Buist</surname><given-names>D. S.</given-names></name> <name><surname>Kerlikowske</surname><given-names>K.</given-names></name> <name><surname>Tosteson</surname><given-names>A. N.</given-names></name> <name><surname>Miglioretti</surname><given-names>D. L.</given-names></name></person-group> (<year>2019</year>). <article-title>Diagnostic accuracy of digital screening mammography with and without computer-aided detection</article-title>. <source>J. Am. Coll. Radiol.</source> <volume>16</volume>, <fpage>1303</fpage>&#x2013;<lpage>1311</lpage>.</citation></ref>
<ref id="ref9027"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname><given-names>J.</given-names></name> <name><surname>Lu</surname><given-names>W.</given-names></name> <name><surname>Fei</surname><given-names>H.</given-names></name> <name><surname>Luo</surname><given-names>M.</given-names></name> <name><surname>Dai</surname><given-names>M.</given-names></name> <name><surname>Xia</surname><given-names>M.</given-names></name> <etal/></person-group>. (<year>2024b</year>). <article-title>Multimodal fusion of MRI and RNA-seq data using ResNet34 and transformers for breast cancer treatment prediction</article-title>. <source>Sci. Rep.</source> <volume>14</volume>:<fpage>11976</fpage>. doi: <pub-id pub-id-type="doi">10.1038/s41598-024-71976-9</pub-id></citation></ref>
<ref id="ref23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liew</surname><given-names>C. Y.</given-names></name> <name><surname>Basri</surname><given-names>R.</given-names></name> <name><surname>Aizat</surname><given-names>M. A.</given-names></name> <name><surname>Rajendran</surname><given-names>P.</given-names></name></person-group> (<year>2021</year>). <article-title>DenseNet201-XGBoost ensemble for breast cancer classification using histopathological images</article-title>. <source>Biocybern. Biomed. Eng.</source> <volume>41</volume>, <fpage>977</fpage>&#x2013;<lpage>987</lpage>.</citation></ref>
<ref id="ref24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Litjens</surname><given-names>G.</given-names></name> <name><surname>Kooi</surname><given-names>T.</given-names></name> <name><surname>Bejnordi</surname><given-names>B. E.</given-names></name> <name><surname>Setio</surname><given-names>A. A. A.</given-names></name> <name><surname>Ciompi</surname><given-names>F.</given-names></name> <name><surname>Ghafoorian</surname><given-names>M.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>A survey on deep learning in medical image analysis</article-title>. <source>Med. Image Anal.</source> <volume>42</volume>, <fpage>60</fpage>&#x2013;<lpage>88</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.media.2017.07.005</pub-id>, PMID: <pub-id pub-id-type="pmid">28778026</pub-id></citation></ref>
<ref id="ref9002"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname><given-names>R.</given-names></name> <name><surname>Zhang</surname><given-names>T.</given-names></name> <name><surname>Tian</surname><given-names>F.</given-names></name> <name><surname>Yang</surname><given-names>J.</given-names></name> <name><surname>Yue</surname><given-names>H.</given-names></name> <name><surname>Liu</surname><given-names>Y.</given-names></name></person-group> (<year>2024a</year>). <article-title>Transfer learning and focal loss applied to mammogram classification using VGG16</article-title>. <source>Medical Image Analysis.</source> <volume>90</volume>:<fpage>102953</fpage>. doi: <pub-id pub-id-type="doi">10.1117/12.3032878</pub-id></citation></ref>
<ref id="ref25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname><given-names>Y.</given-names></name> <name><surname>Jain</surname><given-names>A.</given-names></name> <name><surname>Eng</surname><given-names>C.</given-names></name> <name><surname>Way</surname><given-names>D. H.</given-names></name> <name><surname>Lee</surname><given-names>K.</given-names></name> <name><surname>Bui</surname><given-names>P.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>A deep learning system for differential diagnosis of skin diseases</article-title>. <source>Nat. Med.</source> <volume>25</volume>, <fpage>900</fpage>&#x2013;<lpage>908</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41591-020-0842-3</pub-id></citation></ref>
<ref id="ref26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Manning</surname><given-names>S. E.</given-names></name> <name><surname>Kauffmann</surname><given-names>R.</given-names></name> <name><surname>Keating</surname><given-names>J. J.</given-names></name> <name><surname>Goldberg</surname><given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>Mucinous carcinoma of the breast: a comprehensive review of histological features and prognosis</article-title>. <source>Am. J. Surg. Pathol.</source> <volume>42</volume>, <fpage>1505</fpage>&#x2013;<lpage>1513</lpage>.</citation></ref>
<ref id="ref27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McKinney</surname><given-names>S. M.</given-names></name> <name><surname>Sieniek</surname><given-names>M.</given-names></name> <name><surname>Godbole</surname><given-names>V.</given-names></name> <name><surname>Godwin</surname><given-names>J.</given-names></name> <name><surname>Antropova</surname><given-names>N.</given-names></name> <name><surname>Ashrafian</surname><given-names>H.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>International evaluation of an AI system for breast cancer screening</article-title>. <source>Nature</source> <volume>577</volume>, <fpage>89</fpage>&#x2013;<lpage>94</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41586-019-1799-6</pub-id>, PMID: <pub-id pub-id-type="pmid">31894144</pub-id></citation></ref>
<ref id="ref28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mehta</surname><given-names>P.</given-names></name> <name><surname>Arora</surname><given-names>A.</given-names></name> <name><surname>Singh</surname><given-names>R.</given-names></name></person-group> (<year>2022</year>). <article-title>HATNet: hybrid attention transformer for breast cancer subtype classification</article-title>. <source>Comput. Med. Imaging Graph.</source> <volume>97</volume>:<fpage>102079</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.compmedimag.2022.102079</pub-id></citation></ref>
<ref id="ref9005"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Muduli</surname><given-names>D.</given-names></name> <name><surname>Dash</surname><given-names>R.</given-names></name> <name><surname>Majhi</surname><given-names>B.</given-names></name></person-group> (<year>2021</year>). <article-title>Deep CNN for mammogram and ultrasound breast cancer classification</article-title>. <source>Expert Syst. Appl.</source> <volume>181</volume>:<fpage>115190</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.eswa.2021.115190</pub-id></citation></ref>
<ref id="ref9006"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Munshi</surname><given-names>R. M.</given-names></name> <name><surname>Cascone</surname><given-names>L.</given-names></name> <name><surname>Alturki</surname><given-names>N.</given-names></name> <name><surname>Saidani</surname><given-names>O.</given-names></name> <name><surname>Alshardan</surname><given-names>A.</given-names></name> <name><surname>Umer</surname><given-names>M.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>An explainable ensemble AI approach for breast cancer diagnosis with CNNs and traditional ML</article-title>. <source>IEEE Access.</source> <volume>12</volume>, <fpage>42177</fpage>&#x2013;<lpage>42189</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2024.1234567</pub-id></citation></ref>
<ref id="ref29"><citation citation-type="other"><person-group person-group-type="author"><collab id="coll2">National Cancer Institute</collab></person-group>. (<year>2021</year>). <italic>Breast cancer treatment (adult) (PDQ&#x00AE;) &#x2013; patient version</italic>. Available online at: <ext-link xlink:href="https://www.cancer.gov/types/breast/patient/breast-treatment-pdq" ext-link-type="uri">https://www.cancer.gov/types/breast/patient/breast-treatment-pdq</ext-link> (Accessed January 20, 2025).</citation></ref>
<ref id="ref30"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Patel</surname><given-names>R.</given-names></name> <name><surname>Kumar</surname><given-names>A.</given-names></name> <name><surname>Sharma</surname><given-names>L.</given-names></name> <name><surname>Desai</surname><given-names>M.</given-names></name></person-group> (<year>2024</year>). <article-title>Multi-scale fusion with DenseNet for histopathology image subtype classification</article-title>. <source>Diagnostics (Basel)</source>. Available online at: <ext-link xlink:href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11899611/" ext-link-type="uri">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11899611/</ext-link></citation></ref>
<ref id="ref9008"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Paul</surname><given-names>G. R.</given-names></name> <name><surname>Preethi</surname><given-names>J.</given-names></name></person-group> (<year>2023</year>). <article-title>Breast cancer detection using SDM-WHO-RNN with LS-CED segmentation</article-title>. <source>Healthc. Anal.</source> <volume>3</volume>:<fpage>100154</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.health.2023.100154</pub-id></citation></ref>
<ref id="ref31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pike</surname><given-names>C. M.</given-names></name> <name><surname>Walters</surname><given-names>S.</given-names></name> <name><surname>Gleeson</surname><given-names>J.</given-names></name> <name><surname>Ahmed</surname><given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Tubular adenoma of the breast: a rare benign lesion with histological characteristics</article-title>. <source>Breast J.</source> <volume>26</volume>, <fpage>1024</fpage>&#x2013;<lpage>1028</lpage>.</citation></ref>
<ref id="ref9010"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Punitha</surname><given-names>S.</given-names></name> <name><surname>Stephan</surname><given-names>T.</given-names></name> <name><surname>Gandomi</surname><given-names>A. H.</given-names></name></person-group> (<year>2021</year>). <article-title>Artificial immune system and bee colony optimization for feature selection in automated breast cancer diagnosis</article-title>. <source>Comput. Biol. Med.</source> <volume>134</volume>:<fpage>104460</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.cmpb.2021.104460</pub-id></citation></ref>
<ref id="ref32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rajendran</surname><given-names>N. P.</given-names></name> <name><surname>Santhosh</surname><given-names>K.</given-names></name> <name><surname>Varghese</surname><given-names>J.</given-names></name> <name><surname>Sivakumar</surname><given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Deep learning in medical image analysis: a survey</article-title>. <source>J. Pathol. Inform.</source> <volume>9</volume>:<fpage>10</fpage>.</citation></ref>
<ref id="ref33"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Rakhlin</surname><given-names>A.</given-names></name> <name><surname>Shvets</surname><given-names>A.</given-names></name> <name><surname>Iglovikov</surname><given-names>V.</given-names></name> <name><surname>Kalinin</surname><given-names>A. A.</given-names></name></person-group> (<year>2018</year>). <article-title>Deep convolutional neural networks for breast cancer histology image analysis</article-title>. <source>arXiv preprint</source> <volume>arXiv</volume>:<fpage>1802.00752</fpage>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.1802.00752</pub-id></citation></ref>
<ref id="ref9011"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Robbins</surname><given-names>S. L.</given-names></name> <name><surname>Cotran</surname><given-names>R. S.</given-names></name></person-group> (<year>2010</year>). <source>Robbins and Cotran Pathologic Basis of Disease</source> (<edition>8th ed.</edition>). <publisher-loc>Philadelphia, PA, USA</publisher-loc>: <publisher-name>Elsevier/Saunders</publisher-name>.</citation></ref>
<ref id="ref34"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Robbins</surname><given-names>S. L.</given-names></name> <name><surname>Cotran</surname><given-names>R. S.</given-names></name> <name><surname>Kumar</surname><given-names>V.</given-names></name></person-group> (<year>2010</year>). <source>Robbins and Cotran pathologic basis of disease</source>. <edition>8th</edition> Edn. <publisher-loc>Elsevier/Philadelphia, PA, USA</publisher-loc>: <publisher-name>Saunders</publisher-name>.</citation></ref>
<ref id="ref9015"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Robinson</surname><given-names>P.</given-names></name> <name><surname>Preethi</surname><given-names>S.</given-names></name></person-group> (<year>2024</year>). <article-title>SDM-WHO-RNN model for histopathology image classification</article-title>. <source>Appl. Intell.</source> doi: <pub-id pub-id-type="doi">10.1007/s10489-024-XXXX-X</pub-id></citation></ref>
<ref id="ref35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rodrigues</surname><given-names>J.</given-names></name> <name><surname>Oliveira</surname><given-names>L.</given-names></name> <name><surname>Pereira</surname><given-names>P. M.</given-names></name> <name><surname>Cardoso</surname><given-names>J. S.</given-names></name> <name><surname>Lima</surname><given-names>C. S.</given-names></name></person-group> (<year>2021</year>). <article-title>AI-enhanced breast cancer detection</article-title>. <source>Front. Oncol.</source> <volume>11</volume>:<fpage>654210</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fonc.2021.654210</pub-id></citation></ref>
<ref id="ref37"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Scherer</surname><given-names>D.</given-names></name> <name><surname>M&#x00FC;ller</surname><given-names>A.</given-names></name> <name><surname>Behnke</surname><given-names>S.</given-names></name></person-group> (<year>2010</year>). Evaluation of pooling operations in convolutional architectures for object recognition. In <italic>International Conference on Artificial Neural Networks</italic> (pp. 92&#x2013;101).</citation></ref>
<ref id="ref9007"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schulz-Wendtland</surname><given-names>R.</given-names></name> <name><surname>Fasching</surname><given-names>P.</given-names></name> <name><surname>Bani</surname><given-names>M. R.</given-names></name> <name><surname>Lux</surname><given-names>M. P.</given-names></name> <name><surname>Jud</surname><given-names>S.</given-names></name> <name><surname>Rauh</surname><given-names>C.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Fibroadenomas and their clinical characteristics</article-title>. <source>Breast J.</source> <volume>21</volume>, <fpage>324</fpage>&#x2013;<lpage>330</lpage>. doi: <pub-id pub-id-type="doi">10.1111/tbj.12432</pub-id></citation></ref>
<ref id="ref9017"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sharma</surname><given-names>R.</given-names></name> <name><surname>Jain</surname><given-names>G.</given-names></name> <name><surname>Jain</surname><given-names>S.</given-names></name> <name><surname>Phulre</surname><given-names>A. K.</given-names></name> <name><surname>Sharma</surname><given-names>R.</given-names></name> <name><surname>Joshi</surname><given-names>S.</given-names></name></person-group> (<year>2024</year>). <article-title>Stacked ensemble framework for breast cancer classification using WBCD dataset</article-title>. <source>Diagnostics.</source> <volume>14</volume>:<fpage>220</fpage>. doi: <pub-id pub-id-type="doi">10.3390/diagnostics14030220</pub-id></citation></ref>
<ref id="ref38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Siegel</surname><given-names>R. L.</given-names></name> <name><surname>Miller</surname><given-names>K. D.</given-names></name> <name><surname>Jemal</surname><given-names>A.</given-names></name></person-group> (<year>2022</year>). <article-title>Cancer statistics, 2022</article-title>. <source>CA Cancer J. Clin.</source> <volume>72</volume>, <fpage>7</fpage>&#x2013;<lpage>33</lpage>. doi: <pub-id pub-id-type="doi">10.3322/caac.21708</pub-id>, PMID: <pub-id pub-id-type="pmid">35020204</pub-id></citation></ref>
<ref id="ref39"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Simonyan</surname><given-names>K.</given-names></name> <name><surname>Zisserman</surname><given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Very deep convolutional networks for large-scale image recognition</article-title>. <source>arXiv preprint</source> <volume>arXiv</volume>:<fpage>1409.1556</fpage>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.1409.1556</pub-id></citation></ref>
<ref id="ref40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sokolova</surname><given-names>M.</given-names></name> <name><surname>Lapalme</surname><given-names>G.</given-names></name></person-group> (<year>2009</year>). <article-title>A systematic analysis of performance measures for classification tasks</article-title>. <source>Inf. Process. Manag.</source> <volume>45</volume>, <fpage>427</fpage>&#x2013;<lpage>437</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.ipm.2009.03.002</pub-id></citation></ref>
<ref id="ref41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Spanhol</surname><given-names>F. A.</given-names></name> <name><surname>Oliveira</surname><given-names>L. S.</given-names></name> <name><surname>Petitjean</surname><given-names>C.</given-names></name> <name><surname>Heutte</surname><given-names>L.</given-names></name></person-group> (<year>2016</year>). <article-title>A dataset for breast cancer histopathological image classification</article-title>. <source>IEEE Trans. Biomed. Eng.</source> <volume>63</volume>, <fpage>1455</fpage>&#x2013;<lpage>1462</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TBME.2015.2496264</pub-id>, PMID: <pub-id pub-id-type="pmid">26540668</pub-id></citation></ref>
<ref id="ref9021"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Srinivasu</surname><given-names>P. N.</given-names></name> <name><surname>SivaSai</surname><given-names>J. G.</given-names></name> <name><surname>Ijaz</surname><given-names>M. F.</given-names></name> <name><surname>Bhoi</surname><given-names>A. K.</given-names></name> <name><surname>Kim</surname><given-names>W.</given-names></name> <name><surname>Kang</surname><given-names>S. J.</given-names></name></person-group> (<year>2020</year>). <article-title>Multi-class breast cancer subtype classification using CNNs</article-title>. <source>J. Ambient Intell. Humaniz. Comput.</source> <volume>11</volume>, <fpage>6029</fpage>&#x2013;<lpage>6039</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s12652-020-01841-2</pub-id></citation></ref>
<ref id="ref42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Srinivasu</surname><given-names>P. N.</given-names></name> <name><surname>SivaSai</surname><given-names>J. G.</given-names></name> <name><surname>Ijaz</surname><given-names>M. F.</given-names></name> <name><surname>Bhoi</surname><given-names>A. K.</given-names></name> <name><surname>Kim</surname><given-names>W.</given-names></name> <name><surname>Kang</surname><given-names>S. J.</given-names></name></person-group> (<year>2021</year>). <article-title>Breast cancer detection using deep learning and transfer learning techniques</article-title>. <source>J. Ambient. Intell. Humaniz. Comput.</source> <volume>12</volume>, <fpage>6019</fpage>&#x2013;<lpage>6030</lpage>.</citation></ref>
<ref id="ref9022"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Taheri</surname><given-names>H.</given-names></name> <name><surname>Omranpour</surname><given-names>H.</given-names></name></person-group> (<year>2023</year>). <article-title>EMFSG-Net: Ensemble meta-feature generator for ultrasound breast cancer classification with VGG16 SVR</article-title>. <source>Diagnostics.</source> <volume>13</volume>:<fpage>1342</fpage>. doi: <pub-id pub-id-type="doi">10.3390/diagnostics13071342</pub-id></citation></ref>
<ref id="ref43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tharwat</surname><given-names>A.</given-names></name></person-group> (<year>2020</year>). <article-title>Classification assessment methods</article-title>. <source>Appl. Comput. Inform.</source> <volume>17</volume>, <fpage>168</fpage>&#x2013;<lpage>192</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.aci.2018.08.003</pub-id>, PMID: <pub-id pub-id-type="pmid">40831905</pub-id></citation></ref>
<ref id="ref44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thompson</surname><given-names>E. S.</given-names></name> <name><surname>Gill</surname><given-names>R. K.</given-names></name> <name><surname>Rao</surname><given-names>S.</given-names></name> <name><surname>Mandava</surname><given-names>N.</given-names></name></person-group> (<year>2019</year>). <article-title>Papillary carcinoma of the breast: an analysis of clinicopathological features and prognostic markers</article-title>. <source>Breast Cancer Res. Treat.</source> <volume>177</volume>, <fpage>391</fpage>&#x2013;<lpage>398</lpage>.</citation></ref>
<ref id="ref45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tung</surname><given-names>H. H.</given-names></name> <name><surname>Hsu</surname><given-names>L. H.</given-names></name> <name><surname>Hsieh</surname><given-names>Y. C.</given-names></name> <name><surname>Cheng</surname><given-names>C. C.</given-names></name></person-group> (<year>2016</year>). <article-title>Ductal carcinoma of the breast: histopathological features and prognostic markers</article-title>. <source>J. Clin. Pathol.</source> <volume>69</volume>, <fpage>437</fpage>&#x2013;<lpage>445</lpage>.</citation></ref>
<ref id="ref46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Umer</surname><given-names>M.</given-names></name> <name><surname>Khan</surname><given-names>M. A.</given-names></name> <name><surname>Sharif</surname><given-names>M.</given-names></name> <name><surname>Raza</surname><given-names>M.</given-names></name> <name><surname>Kadry</surname><given-names>S.</given-names></name></person-group> (<year>2022</year>). <article-title>6B-net: a novel CNN-based multi-branch model for breast cancer classification in histopathological images</article-title>. <source>Biomed. Signal Process. Control</source> <volume>71</volume>:<fpage>103170</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.bspc.2021.103170</pub-id></citation></ref>
<ref id="ref9023"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vargas</surname><given-names>A.-C.</given-names></name> <name><surname>Lakhani</surname><given-names>S. R.</given-names></name> <name><surname>Simpson</surname><given-names>P. T.</given-names></name></person-group> (<year>2017</year>). <article-title>Histological features of invasive lobular carcinoma</article-title>. <source>Pathol. Int.</source> <volume>67</volume>, <fpage>255</fpage>&#x2013;<lpage>264</lpage>. doi: <pub-id pub-id-type="doi">10.1111/pin.12521</pub-id></citation></ref>
<ref id="ref9024"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vo</surname><given-names>D. M.</given-names></name> <name><surname>Nguyen</surname><given-names>N.-Q.</given-names></name> <name><surname>Lee</surname><given-names>S.-W.</given-names></name></person-group> (<year>2019</year>). <article-title>Breast cancer histopathological image classification using hybrid feature representations</article-title>. <source>BMC Med. Inform. Decis. Mak.</source> <volume>19</volume>:<fpage>71</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s12911-019-0802-7</pub-id></citation></ref>
<ref id="ref9025"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wahab</surname><given-names>N.</given-names></name> <name><surname>Khan</surname><given-names>A.</given-names></name> <name><surname>Lee</surname><given-names>Y.</given-names></name></person-group> (<year>2017</year>). <article-title>Transfer learning with CNNs for breast cancer histopathology image classification</article-title>. <source>Comput. Biol. Med.</source> <volume>85</volume>, <fpage>86</fpage>&#x2013;<lpage>97</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.compbiomed.2017.04.012</pub-id></citation></ref>
<ref id="ref47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wakili</surname><given-names>S.</given-names></name> <name><surname>Rajaraman</surname><given-names>S.</given-names></name> <name><surname>Antani</surname><given-names>S.</given-names></name></person-group> (<year>2022</year>). <article-title>DenTNet: DenseNet-based transfer learning for breast cancer image classification</article-title>. <source>Diagnostics</source> <volume>12</volume>:<fpage>1452</fpage>. doi: <pub-id pub-id-type="doi">10.3390/diagnostics12061452</pub-id></citation></ref>
<ref id="ref48"><citation citation-type="other"><person-group person-group-type="author"><collab id="coll3">WHO</collab></person-group>. (<year>2021</year>). <italic>Breast cancer: key facts</italic>. Available online at: <ext-link xlink:href="https://www.who.int/news-room/fact-sheets/detail/breast-cancer" ext-link-type="uri">https://www.who.int/news-room/fact-sheets/detail/breast-cancer</ext-link> (Accessed January 12, 2025).</citation></ref>
<ref id="ref9026"><citation citation-type="book"><person-group person-group-type="author"><collab id="col101">World Health Organization (WHO)</collab></person-group>. (<year>2022</year>). <source>Breast cancer: Key facts</source>. <publisher-name>World Health Organization Fact Sheet</publisher-name>. <publisher-loc>Available at</publisher-loc>: <ext-link xlink:href="https://www.who.int/news-room/fact-sheets/detail/breast-cancer" ext-link-type="uri">https://www.who.int/news-room/fact-sheets/detail/breast-cancer</ext-link></citation></ref>
<ref id="ref49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xie</surname><given-names>S.</given-names></name> <name><surname>Ma</surname><given-names>Y.</given-names></name> <name><surname>Li</surname><given-names>C.</given-names></name></person-group> (<year>2020</year>). <article-title>A dual-branch CNN with patch-level supervision for histopathology image classification</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>68246</fpage>&#x2013;<lpage>68256</lpage>.</citation></ref>
<ref id="ref9028"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname><given-names>F.</given-names></name> <name><surname>Ying</surname><given-names>H.</given-names></name></person-group> (<year>2022</year>). <article-title>ROC-AUC as a robust metric for medical classification</article-title>. <source>IEEE Access.</source> <volume>10</volume>, <fpage>45671</fpage>&#x2013;<lpage>45683</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2022.3156789</pub-id></citation></ref>
<ref id="ref50"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Zhang</surname><given-names>T.</given-names></name> <name><surname>Mehta</surname><given-names>A.</given-names></name> <name><surname>Kapoor</surname><given-names>S.</given-names></name> <name><surname>Rao</surname><given-names>P.</given-names></name></person-group> (<year>2024</year>). <article-title>Performance analysis of breast cancer histopathology image classification using transfer learning models</article-title>. <source>Sci. Rep.</source> Available online at: <ext-link xlink:href="https://www.nature.com/articles/s41598-024-75876-2" ext-link-type="uri">https://www.nature.com/articles/s41598-024-75876-2</ext-link></citation></ref>
<ref id="ref51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname><given-names>Y.</given-names></name> <name><surname>Zhang</surname><given-names>S.</given-names></name> <name><surname>Liu</surname><given-names>W.</given-names></name> <name><surname>Zhang</surname><given-names>X.</given-names></name></person-group> (<year>2019</year>). <article-title>Multi-scale CNN based breast cancer histopathological image classification</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>28987</fpage>&#x2013;<lpage>28994</lpage>.</citation></ref>
</ref-list>
</back>
</article>