Artificial intelligence (AI) can classify skin lesions with an accuracy that, in curated evaluations, approaches or matches that of dermatologists. Yet pooled accuracy is only weakly informative for deployment. What determines whether a system is safe to place in a clinical pathway is whether it clears a pre-specified operating threshold, conventionally a melanoma rule-out sensitivity of at least 95%, and whether that attainment survives the move from internal test sets to external and real-world data, and whether it holds across skin tones. These deployment-relevant questions remain poorly quantified.
Objectives: To determine, among AI-based systems for cutaneous cancer detection evaluated against an acceptable reference standard: (i) pooled sensitivity and specificity by cancer type and the proportion of systems meeting pre-specified clinical sensitivity thresholds (RQ1); (ii) how accuracy and threshold attainment differ across data environments, internal validation, external validation, and prospective/real-world deployment (RQ2); and (iii) how performance differs across skin-tone strata where reported, and what proportion of studies report skin tone at all and report tone-stratified performance (RQ3).
NB: This paper is currently under review. Journal: PLOS MEDICINE



