EUVIMEDEuropean Health Evidence
Uhr 7/7Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

A Comparative Analysis of Large Language Model Performance on USMLE Step 1-Style Allergy/Immunology Questions: Evaluating Correctness and Consistency

Moshe Carroll, Sabrina Kentis, Hannah Kareff, Clyde Schechter, Sunit Jariwala

Applied Clinical Informatics · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Background: Large language models are rapidly transforming medical education, yet their performance in Allergy/Immunology remains insufficiently characterized. Furthermore, concerns regarding accuracy, consistency, and sensitivity to input format persist. Objectives: To evaluate and compare the accuracy and response consistency of three leading large language models—ChatGPT-5, Gemini 2.5, and Grok 4—on Allergy/Immunology United States Medical Licensing Examination Step 1-style questions under different prompt conditions. Methods: Thirty-five United States Medical Licensing Examination Step 1-style questions were selected. Questions were presented to each model in two formats: single-question prompts and a combined prompt containing all questions. Fifteen trials were conducted for each format per model. Performance was assessed using mean accuracy and variability was measured using Shannon entropy. Mixed effects models tested effects of model, prompt condition, and question difficulty. Results: Overall accuracy differed significantly (p < 0.001), with Gemini (80.7%) and Grok (80.5%) achieving higher mean scores than ChatGPT (74.3%). Single-item prompts yielded superior performance with Grok (93.1%) and Gemini (90.9%) demonstrating the highest accuracy. Transitioning to a combined prompt significantly reduced accuracy for all models. Accuracy also decreased with increasing question difficulty for all models. Grok demonstrated superior reliability, maintaining the lowest overall response entropy, whereas ChatGPT exhibited the highest variability. Conclusions: On Allergy/Immunology Step 1-style questions, Gemini and Grok demonstrated higher accuracy than ChatGPT, although their overall accuracies remained approximately 81%. Grok offered the most consistent performance. All models demonstrated substantial sensitivity to prompt complexity and inherent performance limitations. These findings underscore the importance of prompt optimization and support the supplementary role of these models in medical education.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Moshe Carroll, Sabrina Kentis, Hannah Kareff, Clyde Schechter, Sunit Jariwala
Quelle
Applied Clinical Informatics
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
1869-0327
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Moshe Carroll, Sabrina Kentis, Hannah Kareff, Clyde Schechter, Sunit Jariwala (2026). A Comparative Analysis of Large Language Model Performance on USMLE Step 1-Style Allergy/Immunology Questions: Evaluating Correctness and Consistency. Applied Clinical Informatics. https://doi.org/10.1055/a-2946-7393
RIS BibTeX CSL-JSON