Skip to main content
News

Foundation AI Model Surpasses Early-Career Dermatologists but Falls Short of Experts in Skin Cancer Diagnosis

Clinical Summary: 

  • Design/Population: Multicenter diagnostic study comparing 652 physician readers with 3 artificial intelligence (AI) models using 1,117 clinical and dermoscopic skin lesion cases representing real-world practice.
  • Key Outcomes: Expert dermatologists achieved the highest diagnostic accuracy (74.2%), outperforming all AI models. A modern unimodal foundation model (72.2%) exceeded physicians with fewer than 3 years of experience, while a first-generation CNN performed worst (56.7%).
  • Clinical Relevance: Findings suggest contemporary AI models may improve diagnostic accuracy among less experienced clinicians and support clinical decision-making, although expert dermatologists continue to provide the highest diagnostic performance.

A modern artificial intelligence (AI) foundation model outperformed less experienced physicians in skin cancer diagnosis but remained inferior to expert dermatologists, highlighting both the promise and current limitations of AI-assisted dermatologic care.

“AI systems for skin cancer detection perform well in controlled settings but frequently underperform in everyday clinical practice, raising critical questions about their readiness for deployment,” stated Julien Anriot, MD, Claude Bernard University Lyon 1, Lyon, France, and coauthors. 

Investigators conducted a multicenter diagnostic study comparing 3 AI algorithms with 652 physicians across varying levels of dermatology experience using a dataset of 1,117 clinical and dermoscopic skin lesion cases designed to reflect routine clinical practice, including rare and atypical lesions. Physician participants independently reviewed 100 randomly assigned cases, while AI performance was evaluated across the full dataset. The primary end point was multiclass diagnostic accuracy, with secondary analyses examining sensitivity, specificity, and balanced accuracy for distinguishing benign from malignant lesions.

Overall, physicians achieved a mean diagnostic accuracy of 65.9%, significantly outperforming the first-generation CNN model, which achieved 56.7%. The unimodal foundation model demonstrated 72.2% accuracy, exceeding the performance of clinicians with fewer than 3 years of dermatology experience (68.2%; P < .001). Dermatologists with more than 10 years of experience achieved the highest overall accuracy at 74.2%, outperforming all AI models, including the unimodal foundation model (72.2%) and the multimodal model (66.3%).

The findings also showed that the unimodal foundation model performed comparably to physicians with 3 to 10 years of experience but did not surpass expert dermatologists. Meanwhile, the multimodal foundation model failed to improve on the unimodal model and demonstrated lower overall diagnostic accuracy.

“Future practice should integrate human-AI collaboration, with AI supporting less experienced clinicians and providing expert triage assistance and help to minimize fatigue-related diagnostic errors,” concluded Dr Anriot et al. 


Source: 

Anriot J, Yan S, Coste C, et al. Limits of artificial intelligence models for skin cancer diagnosis in realistic settings. JAMA Dermatol. Published online: June 3, 2026. doi: 10.1001/jamadermatol.2026.1492

© 2026 HMP Global. All Rights Reserved.
Any views and opinions expressed are those of the author(s) and/or participants and do not necessarily reflect the views, policy, or position of Oncology Learning Network or HMP Global, their employees, and affiliates.