Beyond Accuracy: Protocol-Dependent Robustness and Reliability in Patch-Based Breast Histopathology Classification

Authors

  • SULEYMAN AL-SHOWARAH Software Engineering Department, Faculty of Information Technology, Mutah University
  • Wael Alzyadat Department of Software Engineering, Al-Zaytoonah University of Jordan, Amman, Jordan
  • Aysh Alhroob Department of Software Engineering, Al-Zaytoonah University of Jordan, Amman, Jordan
  • Ameen Shaheen Department of Software Engineering, Al-Zaytoonah University of Jordan, Amman, Jordan

DOI:

https://doi.org/10.15849/ijasca.v18i2.69

Keywords:

Deep Learning, Breast Cancer, Histopathology, Convolutional Neural Networks, Evaluation protocol, Computational Pathology

Abstract

Breast cancer is one of the diseases that cause a dead worldwide among women. Early diagnosis can save many lives and this will lead to a proper treatment. Histopathological image analysis is an important diagnostic method for detecting breast cancer.  Computer-aided diagnosis of breast images helps radiologists do the task more efficiently and appropriately. The reliable evaluation of deep learning models in computational pathology depends strongly on how the data is split during validation. In patch-based histopathology classification, random image-level splitting can obscure inter-patient heterogeneity and produce overly stable performance estimates.This study investigates the effect of the evaluation protocol on Invasive Ductal Carcinoma (IDC) detection using the Breast Histopathology Images dataset. A lightweight Convolutional Neural Network (CNN) was trained using balanced sampling and evaluated under two 5-fold cross-validation schemes: strict patient-level separation and stratified random image-level splitting. Under patient-level evaluation, the model achieved a mean accuracy of 0.7826 ± 0.0295 and a mean AUC of 0.8691 ± 0.0257, with a Brier score of 0.1582 ± 0.0159. Random splitting produced similar mean accuracy (0.7867 ± 0.0096) and AUC (0.8652 ± 0.0053), with slightly improved Brier score (0.1509 ± 0.0047). While paired statistical testing showed no significant differences in mean performance (p > 0.05), patient-level evaluation exhibited substantially higher fold-wise variability. Calibration analysis indicated moderate reliability with mild overconfidence, and error inspection revealed consistent morphological failure patterns. Overall, the evaluation protocol mainly affects perceived robustness rather than average performance, highlighting the importance of patient-level validation and variability reporting.

Downloads

All Downloads: 6

Download data is not yet available.

Downloads

Published

2026-06-21

How to Cite

AL-SHOWARAH, S., Alzyadat, W., Alhroob, A., & Shaheen, A. (2026). Beyond Accuracy: Protocol-Dependent Robustness and Reliability in Patch-Based Breast Histopathology Classification. International Journal of Advances in Soft Computing and Its Applications, 18(2), 132–149. https://doi.org/10.15849/ijasca.v18i2.69

Google Scholar Link