Artificial Intelligence for Medical Imaging Diagnosis: From Accuracy to Clinical Reliability through Multimodal Fusion, Validation, and Regulatory Perspectives
DOI:
https://doi.org/10.70882/josrar.2026.v3i4.230Keywords:
Artificial Intelligence, Medical Imaging, Multimodal Fusion, Clinical Validation, Explainable AI, Regulatory ComplianceAbstract
Artificial intelligence (AI)-based medical imaging diagnosis has demonstrated remarkable performance across multiple clinical domains, with deep learning models frequently reporting diagnostic accuracy, sensitivity, and specificity exceeding 90% under controlled experimental conditions. However, translating these results into clinically reliable, regulatory-compliant systems remains a critical challenge. As a narrative survey rather than an original benchmark study, this paper reports no new experimental results; instead, it introduces a modality-aware analytical framework organizing the existing literature across four dimensions: imaging modality, data provenance, validation maturity, and model architecture. Using this taxonomy, the survey synthesizes unimodal and multimodal fusion approaches spanning radiology (CT, MRI, X-ray), pathology (whole-slide images), ophthalmology (fundus photography, OCT), and multi-source fusion combining imaging with electronic health records (EHR) and genomic data. The synthesis indicates that high reported accuracy is strongly contingent on data characteristics and evaluation conditions, with many models relying on low-maturity validation lacking evidence of generalization in real-world settings. To address these limitations, an engineering-oriented deployment framework is proposed, integrating modality-driven model selection, structured preprocessing pipelines, multi-level clinical validation, computational feasibility assessment, and explainability, together with a clinical deployment readiness model spanning validation maturity, data diversity, interpretability, and regulatory alignment. Key challenges include the single-site generalization gap, algorithmic bias across demographic groups, limited clinical adoption of explainable AI, insufficient alignment with regulatory frameworks including FDA 510(k), De Novo, and EU MDR/IVDR pathways, and a continuing need for prospective multicenter validation. Future directions toward federated learning, foundation models, certification-aware design, and multimodal digital biomarker integration are outlined.
References
Abramoff, M. D., Lavin, P. T., Birch, M., Shah, N., & Folk, J. C. (2018). Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. NPJ Digital Medicine, 1(1), 39. https://doi.org/10.1038/s41746-018-0040-6
Abramoff, M. D., Tarver, M. E., Loyo-Berrios, N., Trujillo, S., Char, D., Obermeyer, Z., ... & Topol, E. J. (2023). Considerations for addressing bias in artificial intelligence for health equity. NPJ Digital Medicine, 6(1), 170. https://doi.org/10.1038/s41746-023-00913-9
Acosta, J. N., Falcone, G. J., Rajpurkar, P., & Topol, E. J. (2022). Multimodal biomedical AI. Nature Medicine, 28(9), 1773–1784. https://doi.org/10.1038/s41591-022-01981-2
Albadawy, E. A., Saha, A., & Mazurowski, M. A. (2018). Deep learning for segmentation of brain tumors: Impact of cross-institutional training and testing. Medical Physics, 45(3), 1150–1158. https://doi.org/10.1002/mp.12752
Ardila, D., Kiraly, A. P., Bharadwaj, S., Choi, B., Reicher, J. J., Peng, L., ... & Shetty, S. (2019). End-to-end lung cancer screening with three-dimensional deep learning on low-dose chest computed tomography. Nature Medicine, 25(6), 954–961. https://doi.org/10.1038/s41591-019-0447-x
Badgeley, M. A., Zech, J. R., Oakden-Rayner, L., Glicksberg, B. S., Liu, M., Gale, W., ... & Dudley, J. T. (2019). Deep learning predicts hip fracture using confounding patient and healthcare variables. NPJ Digital Medicine, 2(1), 31. https://doi.org/10.1038/s41746-019-0105-1
Bilal, A., Imran, A., Baig, T. I., Liu, X., Long, H., Alzahrani, A., & Shafiq, M. (2024). DeepSVDNet: A deep learning-based approach for detecting and classifying vision-threatening diabetic retinopathy in retinal fundus images. Computer Systems Science and Engineering, 48(2), 511–528.
Chen, R. J., Ding, T., Lu, M. Y., Williamson, D. F. K., Jaume, G., Song, A. H., ... & Mahmood, F. (2024). Towards a general-purpose foundation model for computational pathology. Nature Medicine, 30(3), 850–862. https://doi.org/10.1038/s41591-024-02857-3
Chen, R. J., Lu, M. Y., Williamson, D. F. K., Chen, T. Y., Lipkova, J., Noor, Z., ... & Mahmood, F. (2022). Pan-cancer integrative histology-genomic analysis via multimodal deep learning. Cancer Cell, 40(8), 865–878. https://doi.org/10.1016/j.ccell.2022.07.004
Cui, C., Yang, H., Wang, Y., Zhao, S., Asad, Z., Coburn, L. A., ... & Huo, Y. (2023). Deep multimodal fusion of image and non-image data in disease diagnosis and prognosis: a review. Progress in Biomedical Engineering, 5(2), 022001. https://doi.org/10.1088/2516-1091/acc2fe
Doi, K. (2007). Computer-aided diagnosis in medical imaging: Historical review, current status and future potential. Computerized Medical Imaging and Graphics, 31(4–5), 198–211. https://doi.org/10.1016/j.compmedimag.2007.02.002
Esteva, A., Kuprel, B., Novoa, R. A., Ko, J., Swetter, S. M., Blau, H. M., & Thrun, S. (2017). Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542(7639), 115–118. https://doi.org/10.1038/nature21056
Ghassemi, M., Oakden-Rayner, L., & Beam, A. L. (2021). The false hope of current approaches to explainable artificial intelligence in health care. The Lancet Digital Health, 3(11), e745–e750. https://doi.org/10.1016/S2589-7500(21)00208-9
Gulshan, V., Peng, L., Coram, M., Stumpe, M. C., Wu, D., Narayanaswamy, A., ... & Webster, D. R. (2016). Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA, 316(22), 2402–2410. https://doi.org/10.1001/jama.2016.17216
He, J., Wang, J., Han, Z., Ma, J., Wang, C., & Qi, M. (2023). An interpretable transformer network for the retinal disease classification using optical coherence tomography. Scientific Reports, 13(1), 3637. https://doi.org/10.1038/s41598-023-30853-z
Larson, D. B., Harvey, H., Rubin, D. L., Irani, N., Justin, R. T., & Langlotz, C. P. (2021). Regulatory frameworks for development and evaluation of artificial intelligence–based diagnostic imaging algorithms: summary and recommendations. Journal of the American College of Radiology, 18(3), 413–424. https://doi.org/10.1016/j.jacr.2020.09.060
Lipkova, J., Chen, R. J., Chen, B., Lu, M. Y., Barbieri, M., Shao, D., ... & Mahmood, F. (2022). Artificial intelligence for multimodal data integration in oncology. Cancer Cell, 40(10), 1095–1110. https://doi.org/10.1016/j.ccell.2022.09.012
Litjens, G., Kooi, T., Bejnordi, B. E., Setio, A. A. A., Ciompi, F., Ghafoorian, M., ... & Sánchez, C. I. (2017). A survey on deep learning in medical image analysis. Medical Image Analysis, 42, 60–88. https://doi.org/10.1016/j.media.2017.07.005
Lu, M. Y., Chen, R. J., Kong, D., Lipkova, J., Singh, R., Williamson, D. F. K., ... & Mahmood, F. (2022). Federated learning for computational pathology on gigapixel whole slide images. Medical Image Analysis, 76, 102298. https://doi.org/10.1016/j.media.2021.102298
McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ 2021;372:n71. doi:10.1136/bmj.n71. http://www.prisma-statement.org
Mir, B. A., Abbas, S. R., & Lee, S. W. (2026). Federated learning in healthcare ethics: A systematic review of privacy-preserving and equitable medical AI. Healthcare, 14(3), 306. https://doi.org/10.3390/healthcare14030306
Moor, M., Banerjee, O., Abad, Z. S. H., Krefl, H. M., Kahn, J. M., Topol, E. J., & Rajpurkar, P. (2023). Foundation models for generalist medical artificial intelligence. Nature, 616(7956), 259–265. https://doi.org/10.1038/s41586-023-05881-4
Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. https://doi.org/10.1126/science.aax2342
Park, Y. W., Park, J. E., Ahn, S. S., Han, K., Kim, N., Oh, J. Y., ... & Lee, S. K. (2024). Deep learning-based metastasis detection in patients with lung cancer to enhance reproducibility and reduce workload in brain metastasis screening with MRI: A multi-center study. Cancer Imaging, 24(1), 38. https://doi.org/10.1186/s40644-024-00669-9
Pierson, E., Bhatt, M., Bhatt, S., ... & Rajpurkar, P. (2024). Machine learning-enabled medical devices authorized by the US Food and Drug Administration in 2024: Regulatory characteristics, predicate lineage, and transparency reporting. NPJ Digital Medicine, (in press). PMC12730494.
Pomponio, R., Erus, G., Habes, M., Doshi, J., Srinivasan, D., Mamourian, E., ... & Davatzikos, C. (2020). Harmonization of large MRI datasets for the analysis of brain imaging patterns throughout the lifespan. NeuroImage, 208, 116450. https://doi.org/10.1016/j.neuroimage.2019.116450
Rajpurkar, P., Irvin, J., Ball, R. L., Zhu, K., Yang, B., Mehta, H., ... & Ng, A. Y. (2017). CheXNet: Radiologist-level pneumonia detection on chest X-rays with deep learning. arXiv preprint arXiv:1711.05225.
Rajpurkar, P., Chen, E., Banerjee, O., & Topol, E. J. (2022). AI in health and medicine. Nature Medicine, 28(1), 31–38. https://doi.org/10.1038/s41591-021-01614-0
Rieke, N., Hancox, J., Li, W., Milletarì, F., Roth, H. R., Albarqouni, S., ... & Cardoso, M. J. (2020). The future of digital health with federated learning. NPJ Digital Medicine, 3(1), 119. https://doi.org/10.1038/s41746-020-00323-1
Roest, C., Fransen, S. J., Kwee, T. C., & Yakar, D. (2022). Comparative performance of deep learning and radiologists for the diagnosis and localization of clinically significant prostate cancer at MRI: A systematic review. Life, 12(10), 1490. https://doi.org/10.3390/life12101490
RSNA Comments, FDA Docket FDA-2024-D-4488. (2025, April 7). Comments on AI-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations [Policy/gray literature].
Saha, A., Bosma, J. S., Twilt, J. J., van Ginneken, B., Bjartell, A., Padhani, A. R., ... & Litjens, G. (2024). Artificial intelligence and radiologists in prostate cancer detection on MRI (PI-CAI): an international, paired, non-inferiority, confirmatory study. The Lancet Oncology, 25(7), 879–887. https://doi.org/10.1016/S1470-2045(24)00235-9
Salim, M., Liu, Y., Sorkhei, M., Ntoula, D., Foukakis, T., Fredriksson, I., ... & Strand, F. (2024). AI-based selection of individuals for supplemental MRI in population-based breast cancer screening: The randomized ScreenTrustMRI trial. Nature Medicine, 30(9), 2623–2630. https://doi.org/10.1038/s41591-024-03093-5
Shen, D., Wu, G., & Suk, H. I. (2017). Deep learning in medical image analysis. Annual Review of Biomedical Engineering, 19, 221–248. https://doi.org/10.1146/annurev-bioeng-071516-044442
Shi, F., Wang, J., Shi, J., Wu, Z., Wang, Q., Tang, Z., ... & Shen, D. (2022). Review of artificial intelligence techniques in imaging data acquisition, segmentation, and diagnosis for COVID-19. IEEE Reviews in Biomedical Engineering, 14, 4–15. https://doi.org/10.1109/RBME.2020.2987975
Singh, Y., Hathaway, Q. A., Keishing, V., Salehi, S., Wei, Y., Horvat, N., ... & Andersen, J. B. (2025). Beyond post hoc explanations: A comprehensive framework for accountable AI in medical imaging through transparency, interpretability, and explainability. Bioengineering, 12(8), 879. https://doi.org/10.3390/bioengineering12080879
Srinidhi, C. L., Ciga, O., & Martel, A. L. (2021). Deep neural network models for computational histopathology: A survey. Medical Image Analysis, 67, 101813. https://doi.org/10.1016/j.media.2020.101813
Steyaert, S., Pizurica, M., Ye, H., Bhatt, P., Shen, J. Y., Wan, X., ... & Swaminathan, M. (2023). Multimodal data fusion for cancer biomarker discovery with deep learning. Nature Machine Intelligence, 5(4), 351–362. https://doi.org/10.1038/s42256-023-00633-5
Tjoa, E., & Guan, C. (2021). A survey on explainable artificial intelligence (XAI): Toward medical XAI. IEEE Transactions on Neural Networks and Learning Systems, 32(11), 4793–4813. https://doi.org/10.1109/TNNLS.2020.3027314
Topol, E. J. (2019). High-performance medicine: The convergence of human and artificial intelligence. Nature Medicine, 25(1), 44–56. https://doi.org/10.1038/s41591-018-0300-7
Wu, E., Wu, K., Daneshjou, R., Ouyang, D., Ho, D. E., & Zou, J. (2021). How medical AI devices are evaluated: Limitations and recommendations from an analysis of FDA approvals. Nature Medicine, 27(4), 582–584. https://doi.org/10.1038/s41591-021-01312-x
Xu, Y., Khan, T. M., Song, Y., & Meijering, E. (2025). Edge deep learning in computer vision and medical diagnostics: A comprehensive survey. Artificial Intelligence Review, 58(3), 93. https://doi.org/10.1007/s10462-024-11033-5
Zech, J. R., Badgeley, M. A., Liu, M., Costa, A. B., Titano, J. J., & Oermann, E. K. (2018). Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study. PLOS Medicine, 15(11), e1002686. https://doi.org/10.1371/journal.pmed.1002686
Zhang, X., Li, Q., Yu, H., et al. (2026). AI framework for multidisease detection via retinal imaging. Nature Medicine. https://doi.org/10.1038/s41591-026-04359-w
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Enoch Jacob Dodo, Amos Takai Yayock, Gregory Onwodi, Ruth Tukuluku (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
- Attribution — You must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.
- NonCommercial — You may not use the material for commercial purposes.
- No additional restrictions — You may not apply legal terms or technological measures that legally restrict others from doing anything the license permits.