TY - GEN
T1 - Explainable Deep Learning for Kidney CT Scan Classification
AU - Vasu, Artham
AU - Srivastava, Suzaan S.
AU - Olivia, Diana
AU - Divya, S.
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - This study presents a clear and accurate framework for automatically classifying kidney computed tomography (CT) scans into four important categories: normal, cyst, stone, and tumour. While current methods show high diagnostic accuracy, their lack of interpretability and confidence calibration limits their use in clinical settings. To address these issues, the proposed system uses an improved image analysis architecture along with three supportive interpretability methods: Gradient-weighted Class Activation Mapping (Grad-CAM), Local Interpretable Model-Agnostic Explanations (LIME), and Score-CAM. These methods help visualize decision patterns at different levels. A realtime reliability diagram is included to regularly evaluate and ensure probability calibration for each classification. This helps provide reliable confidence estimates. The framework was built and tested using a publicly available dataset of 1 2, 4 4 6 kidney CT images that cover different anatomical orientations and contrast phases. It achieved a classification accuracy of 94.2%, a macroaveraged F1-score of 0.925, and an Expected Calibration Error (ECE) of 2.0%. The system shows strong ability to differentiate and maintain robust confidence. These results emphasize the importance of combining explainability and calibration to develop transparent, trustworthy, and clinically valuable diagnostic systems.
AB - This study presents a clear and accurate framework for automatically classifying kidney computed tomography (CT) scans into four important categories: normal, cyst, stone, and tumour. While current methods show high diagnostic accuracy, their lack of interpretability and confidence calibration limits their use in clinical settings. To address these issues, the proposed system uses an improved image analysis architecture along with three supportive interpretability methods: Gradient-weighted Class Activation Mapping (Grad-CAM), Local Interpretable Model-Agnostic Explanations (LIME), and Score-CAM. These methods help visualize decision patterns at different levels. A realtime reliability diagram is included to regularly evaluate and ensure probability calibration for each classification. This helps provide reliable confidence estimates. The framework was built and tested using a publicly available dataset of 1 2, 4 4 6 kidney CT images that cover different anatomical orientations and contrast phases. It achieved a classification accuracy of 94.2%, a macroaveraged F1-score of 0.925, and an Expected Calibration Error (ECE) of 2.0%. The system shows strong ability to differentiate and maintain robust confidence. These results emphasize the importance of combining explainability and calibration to develop transparent, trustworthy, and clinically valuable diagnostic systems.
UR - https://www.scopus.com/pages/publications/105036275520
UR - https://www.scopus.com/pages/publications/105036275520#tab=citedBy
U2 - 10.1109/MoSICom67153.2025.11398278
DO - 10.1109/MoSICom67153.2025.11398278
M3 - Conference contribution
AN - SCOPUS:105036275520
T3 - Proceedings of IEEE International Conference on Modelling, Simulation and Intelligent Computing, MoSICom 2025
SP - 141
EP - 146
BT - Proceedings of IEEE International Conference on Modelling, Simulation and Intelligent Computing, MoSICom 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - IEEE International Conference on Modelling, Simulation and Intelligent Computing, MoSICom 2025
Y2 - 10 December 2025 through 12 December 2025
ER -