Artificial intelligence-based clinical decision support systems (AI-CDSS) are increasingly promoted to improve diagnostic accuracy and clinical decision-making. However, questions remain about whether current evidence adequately demonstrates meaningful patient benefit. Existing research relies heavily on surrogate diagnostic metrics, provides limited insight into clinician–AI interaction, and rarely evaluates implementation in real-world healthcare settings. This review synthesizes evidence across these interconnected domains to identify critical gaps in AI-CDSS evaluation.
Methods:
A critical narrative review was conducted using the Grant and Booth (2009) review typology and thematic narrative synthesis. Literature published between January 2015 and May 2026 was searched in PubMed/MEDLINE, Scopus, CINAHL, and IEEE Xplore using three domains: clinical effectiveness, human–AI interaction, and implementation science. Relevant regulatory guidance and evaluation frameworks were identified through targeted searches of institutional and professional organization websites.
Results:
The search identified 1,247 records; following screening and eligibility assessment, 33 sources (28 peer-reviewed publications and five grey literature documents) were included. The strongest effectiveness evidence, a 2026 systematic review and meta-analysis of five randomized controlled trials involving 12,657 participants, demonstrated only a small pooled effect (standardized mean difference 0.182; 95% CI 0.003–0.362; GRADE: moderate), with all trials measuring diagnostic accuracy rather than patient-centred outcomes such as mortality or length of stay. Most studies originated from East Asia and focused on radiology, limiting generalizability. Human factors evidence identified significant vulnerabilities, including automation bias that reduced diagnostic accuracy among inexperienced radiologists when incorrect AI recommendations were presented, high false-positive alarm rates promoting alert fatigue, and limited clinician scrutiny of erroneous AI outputs. Although evaluation and implementation frameworks, including CONSORT-AI, SPIRIT-AI, DECIDE-AI, CFIR, and RE-AIM, have been developed, prospective validation remains limited, with almost no implementation evidence from sub-Saharan Africa or other low- and middle-income countries.
Conclusion:
Current AI-CDSS evidence inadequately captures real-world effectiveness and unintended harms. Future research should prioritize patient-centred outcomes, rigorous evaluation of clinician–AI interaction, and prospective validation of implementation frameworks across diverse healthcare settings. Strengthening evidence standards, including fairness and equity considerations, is essential for the safe, effective, and scalable deployment of AI-CDSS.
NB: This paper is currently under review. Journal:BMC Medical Informatics and Decision Making



