Are artificial intelligence large language models a reliable tool for difficult differential diagnosis? An a posteriori analysis of a peculiar case of necrotizing otitis externa

Pugliese, Giorgia; Maccari, Alberto; Felisati, Elena; Felisati, Giovanni; Giudici, Leonardo; Rapolla, Chiara; Pisani, Antonia; Saibene, Alberto Maria

doi:10.1002/ccr3.7933

Cited by 3 publications

(1 citation statement)

References 23 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…In particular, medical students have been exploring the potential of LLMs to reinforce learned information, provide clari cation on complex clinical topics, and aid in preparation for tests such as the United States Medical Licensing Examination (USMLE) series and related subject exams [1,6]. Initial studies have demonstrated that ChatGPT-4 can process board-level exam questions and provide useful clinical insights [7][8][9][10][11][12]. No studies, however, have strati ed these capabilities within the specialty-speci c domains appearing on the USMLE Step 2 examination and associated clinical subject exams.…”

Section: Introductionmentioning

confidence: 99%

Evaluating ChatGPT-4 in medical education: an assessment of subject exam performance reveals limitations in clinical curriculum support for students

Mackey,

Garabet,

Maule

et al. 2023

Preprint

View full text Add to dashboard Cite

This study evaluates the proficiency of ChatGPT-4 across various medical specialties and assesses its potential as a study tool for medical students preparing for the United States Medical Licensing Examination (USMLE) Step 2 and related clinical subject exams. ChatGPT-4 answered board-level questions with 89% accuracy, but showcased significant discrepancies in performance across specialties. Although it excelled in psychiatry, neurology, and obstetrics & gynecology, it underperformed in pediatrics, emergency medicine, and family medicine. These variations may be potentially attributed to the depth and recency of training data as well as the scope of the specialties assessed. Specialties with significant interdisciplinary overlap had lower performance, suggesting complex clinical scenarios pose a challenge to the AI. In terms of the future, the overall efficacy of ChatGPT-4 indicates a promising supplemental role in medical education, but performance inconsistencies across specialties in the current version lead us to recommend that medical students use AI with caution.

show abstract

Section: Introductionmentioning

confidence: 99%

Evaluating ChatGPT-4 in medical education: an assessment of subject exam performance reveals limitations in clinical curriculum support for students

Mackey,

Garabet,

Maule

et al. 2023

Preprint

View full text Add to dashboard Cite

show abstract

Evaluating ChatGPT-4 in medical education: an assessment of subject exam performance reveals limitations in clinical curriculum support for students

Mackey,

Garabet,

Maule

et al. 2024

Discov Artif Intell

View full text Add to dashboard Cite

This study evaluates the proficiency of ChatGPT-4 across various medical specialties and assesses its potential as a study tool for medical students preparing for the United States Medical Licensing Examination (USMLE) Step 2 and related clinical subject exams. ChatGPT-4 answered board-level questions with 89% accuracy, but showcased significant discrepancies in performance across specialties. Although it excelled in psychiatry, neurology, and obstetrics and gynecology, it underperformed in pediatrics, emergency medicine, and family medicine. These variations may be potentially attributed to the depth and recency of training data as well as the scope of the specialties assessed. Specialties with significant interdisciplinary overlap had lower performance, suggesting complex clinical scenarios pose a challenge to the AI. In terms of the future, the overall efficacy of ChatGPT-4 indicates a promising supplemental role in medical education, but performance inconsistencies across specialties in the current version lead us to recommend that medical students use AI with caution.

show abstract

Assessing unknown potential—quality and limitations of different large language models in the field of otorhinolaryngology

Buhr,

Smith,

Huppertz

et al. 2024

Acta Oto-Laryngologica

View full text Add to dashboard Cite

Are artificial intelligence large language models a reliable tool for difficult differential diagnosis? An a posteriori analysis of a peculiar case of necrotizing otitis externa

Cited by 3 publications

References 23 publications

Evaluating ChatGPT-4 in medical education: an assessment of subject exam performance reveals limitations in clinical curriculum support for students

Evaluating ChatGPT-4 in medical education: an assessment of subject exam performance reveals limitations in clinical curriculum support for students

Evaluating ChatGPT-4 in medical education: an assessment of subject exam performance reveals limitations in clinical curriculum support for students

Assessing unknown potential—quality and limitations of different large language models in the field of otorhinolaryngology

Contact Info

Product

Resources

About