EUVIMEDEuropean Health Evidence
Uhr 10/10Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

How Artificial, How Intelligent? LLMs and Medical Experts Face Off in TMA Diagnosis

Gizem Zorlu Görgülügil, Ceyda Kaçar, Elif Şen, Ayça İnci, Volkan Karakuş

Sakarya Medical Journal · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Objective: To compare physician specialists and large language models in the differential diagnosis and plasma exchange decision making of standardized thrombotic microangiopathy case vignettes.Methods: This case vignette based comparative study included three standardized TMA scenarios representing atypical hemolytic uremic syndrome, malignant hypertension associated TMA, and acquired thrombotic thrombocytopenic purpura. Nine evaluators participated: two nephrologists, two hematologists, two internal medicine specialists, and three large language models (ChatGPT-5.2, Gemini Pro, and Microsoft Copilot). For each case, participants provided a most likely diagnosis, three differential diagnoses, supporting clinical reasoning, and a plasma exchange decision. Diagnostic accuracy, plasma exchange appropriateness, and critical management errors were analyzed descriptively and comparatively.Results: A total of 27 independent evaluations were analyzed. Overall Top-1 diagnostic accuracy was 51.9%. Accuracy varied by case, with the lowest performance observed in the aHUS scenario (22.2%) and the highest in malignant hypertension–associated TMA (77.8%). Physicians achieved a Top-1 accuracy of 50.0% (9/18), whereas LLMs demonstrated 55.6% accuracy (5/9). Plasma exchange decision accuracy was 63.0% overall, with overtreatment more common in non-TTP cases. Ten critical management errors (37.0%) were identified. Agreement for plasma exchange decisions demonstrated fair concordance (κ = 0.29).Conclusions: Substantial variability exists in both diagnostic classification and therapeutic decision-making in TMA scenarios. While LLMs demonstrated competitive diagnostic performance in selected cases, management variability persisted across evaluator groups. These findings highlight the diagnostic complexity of TMA and support the role of artificial intelligence as an adjunctive clinical decision-support tool rather than an autonomous decision-maker.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Gizem Zorlu Görgülügil, Ceyda Kaçar, Elif Şen, Ayça İnci, Volkan Karakuş
Quelle
Sakarya Medical Journal
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2146-409X
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Gizem Zorlu Görgülügil, Ceyda Kaçar, Elif Şen, Ayça İnci, Volkan Karakuş (2026). How Artificial, How Intelligent? LLMs and Medical Experts Face Off in TMA Diagnosis. Sakarya Medical Journal. https://doi.org/10.31832/smj.1889486
RIS BibTeX CSL-JSON