Skip to main content
Podcasts

Experienced Dermatologists Continue to Outperform Artificial Intelligence in Skin Cancer Diagnosis

Clinical Summary: 

  • Design/Population: This study compared artificial intelligence (AI) models with clinicians of varying expertise using a realistic dermoscopy-based diagnostic assessment that included benign and malignant skin lesions encountered in routine clinical practice. Performance was evaluated across both conventional neural network models and newer transformer-based foundation models.
  • Key Outcomes: Advanced AI models performed similarly to clinicians after dermatology training but remained less accurate than experienced dermoscopy experts. The findings suggest that AI performance declines in more clinically representative diagnostic settings compared with previously reported benchmark studies.
  • Clinical Relevance: These results support the integration of AI as a clinical decision-support tool while reinforcing the continued importance of physician training and expertise in skin cancer diagnosis.

Luc Thomas, MD, PhD, Claude Bernard University Lyon 1, Lyon, France, discusses findings from a study evaluating the performance of artificial intelligence (AI) models for skin cancer diagnosis under real-world clinical conditions. The investigators compared multiple AI platforms with clinicians across different levels of expertise using a comprehensive dermoscopy-based assessment designed to reflect routine dermatology practice.

The study demonstrated that although state-of-the-art AI performed comparably to trained dermatologists, experienced dermoscopy experts continued to outperform even the most advanced foundation models. Dr Thomas also discusses how AI can complement clinical practice and education while emphasizing that physician expertise remains essential for accurately diagnosing complex and uncommon skin lesions.

Transcript: 

Hello, my name is Luc Thomas. I'm living in France and working in France. I will do my best to explain why I think it is interesting to have a look at this paper we published a few weeks ago in the JAMA Dermatology journal.

The paper is entitled "Limits of artificial intelligence models in skin cancer diagnosis in realistic settings."

I am a professor of dermatology, and my main activity, in terms of both clinical activity and research activity, is the early diagnosis of skin cancer.

Most of my publications and books are about this topic, mainly with the help of dermoscopy, a noninvasive imaging system to better qualify pigmented and nonpigmented lesions on the skin and to more precisely diagnose which lesions should be excised and why we need to keep some others in place.

Dermoscopy is quite a young science. It was introduced at the end of the 20th century, but it became a very important tool in the diagnosis of skin cancer.

Of course, the most important of these is melanoma. I'm treating melanoma.

The difference between France and the US is the fact that we are also treating patients with metastatic disease. This is also part of my research, the development of new treatments in melanoma.

This paper was interesting because, as a teacher and a dermatology professor, I'm often asked by my students, "Well, is it really necessary to learn these very complex criteria for melanoma that you can see with dermoscopy, since artificial intelligence is nowadays able to do the same?"

In other words, "Well, I don't need to know how to spell words because I have spellcheck on my computer. Do I need to learn all these things that are very complicated, with many exceptions, many rules, since artificial intelligence is basically doing the same?"

The reason why we decided to launch this study is the fact that we got the impression that most of the testing—man against machine—in artificial intelligence for the diagnosis of skin cancer was kind of unfair to human beings.

In other words, the setting of the test was not very representative of what is happening during a real-time consultation with real patients.

For this reason, we decided to submit artificial intelligence systems to a more realistic way of diagnosing skin cancer in everyday practice in dermatology or in general medicine.

We had developed many years ago a test that is submitted to many students in dermatology, in general medicine, and in many other settings in order to evaluate the capacity of our students to use the dermoscopy criteria.

We were lucky enough to convince some very famous colleagues in the field of artificial intelligence to also test their algorithms with our clinical test.

The interesting part of this clinical test is two things.

First of all, it was not limited to melanoma and nevi but encompassed many different nosological entities, like other skin cancers, benign pigmented or nonpigmented tumors, and many confounders that you can have when you check the skin of a patient.

So there were many more possibilities to make a diagnosis of a malignant or benign condition, including some very rare tumors, like Merkel cell tumors, for example.

For this reason, we believe that it was much more representative of what a clinician is facing when looking at one patient's skin.

The second thing is that we were able to test the community of people interested in diagnosing skin cancer at different levels of expertise.

We were also able to test artificial intelligence systems at different stages of development.

As you probably know, the first systems that were used in artificial intelligence were classical neural network systems.

More recently, much more sophisticated transformer-based foundation models have been published, and these are believed to be much more efficient.

We were able to compare human beings with very sophisticated, advanced artificial intelligence systems, as well as classical neural network systems, in our clinical trial.

Basically, every artificial intelligence system was asked to make 100,000 tests, and we also had in our database the results of a large number of humans with different levels of expertise in dermoscopy.

Interestingly, under these more realistic conditions, we were able to show that human beings, after training—after three years of training, which is usually the duration of dermatology residency—achieved equivalent results to the most sophisticated artificial intelligence systems.

Of course, every human being outperformed the classical neural network artificial intelligence systems.

Interestingly also, all the very important and very well-trained experts who were included in our database performed better than even the most advanced artificial intelligence systems.

Of course, we can ask different questions about that.

The main question is, well, okay, artificial intelligence is trained on series of cases.

The most sophisticated artificial intelligence we used in this publication was trained with two million images, unselected, and the learning was not supervised in order not to introduce any kind of artificial results.

Probably, if we are able to increase the number of images used to train artificial intelligence, one could object to our conclusions and say that one day artificial intelligence will do better than human beings.

But this is just a hypothesis because introducing higher numbers of images to train artificial intelligence will require a very expensive gathering of images with known diagnoses.

This is why we think that it's a moving target.

Of course, we can increase the knowledge of the system, but at some point nosography is also changing very rapidly, and probably artificial intelligence will not achieve the results that experts are achieving in their clinical practice.

I will finish by explaining why I think this is important.

It's not only important to say that the classical experiments we have been publishing for many years are not valid.

Of course, artificial intelligence is a very important tool that we will all use in our everyday practice and that is very interesting in areas where access to doctors is difficult.

It's also a very interesting tool to be used during the training of our students.

It's kind of a safety net for beginners when they are facing a patient.

Sometimes they can be unsure of the diagnosis, and sometimes artificial intelligence can help with that.

But it's also important to stress that if we decide at some point not to use human skills for diagnosing skin cancer, we may end up with a regression in our ability to diagnose difficult conditions.

That is why it is still necessary to train experts and to train future experts.

It is very important to have medical students confident in their future by studying complex notions in medicine and continuing to study the important criteria needed to diagnose cancer.

So, in other words, of course, it's not a competition.

We will, of course, include artificial intelligence in the management of our patients more and more in the near future.

But knowing the basics, knowing the difficult and complex situations that every clinician may encounter with the next patient who opens the door of the clinic, is also very important in order to give our patients the very best care.

Thank you very much.


Source: 

Anriot J, Yan S, Coste C, et al. Limits of artificial intelligence models for skin cancer diagnosis in realistic settings. JAMA Dermatol. Published online: June 3, 2026. doi: 10.1001/jamadermatol.2026.1492

© 2026 HMP Global. All Rights Reserved.
Any views and opinions expressed are those of the author(s) and/or participants and do not necessarily reflect the views, policy, or position of Oncology Learning Network or HMP Global, their employees, and affiliates.