Vollständiger Abstract
Worum geht es in dieser Arbeit?
Fetal magnetic resonance imaging (MRI) plays an important role in evaluating prenatal central nervous system (CNS) abnormalities, but expert interpretation and counseling remain challenging. Large language models (LLMs) may support text-based diagnostic reasoning, yet their performance in fetal CNS MRI has not been well characterized. The study aimed to compare ChatGPT, Gemini, and DeepSeek in text-based diagnostic reasoning for fetal CNS MRI cases. This retrospective single-center study included 85 fetal MRI cases with postnatal diagnostic confirmation. Each case was converted into a standardized text summary containing MRI findings and essential clinical information, without image input. The 3 LLMs generated ranked differential diagnoses, diagnostic reasoning, management recommendations, prognostic counseling points, and clinical cautions. Anonymized outputs were independently evaluated under blinded conditions for principal diagnosis matching, diagnostic accuracy, clinical reasoning quality, clinical suggestion utility, and communication style and safety. DeepSeek achieved the highest principal diagnosis matching rate (66/85, 77.6%), followed by Gemini (59/85, 69.4%) and ChatGPT (52/85, 61.2%). The overall difference was significant (Cochran’s Q = 6.2553, P = .0438), although no pairwise comparison remained significant after Holm correction. DeepSeek showed the highest diagnostic accuracy and reasoning quality, whereas Gemini performed best in clinical suggestion utility and communication style and safety. Exploratory correlation analyses suggested model-specific associations among evaluation metrics. LLM performance in fetal CNS MRI text-based reasoning was multidimensional and model-dependent. These findings support further supervised evaluation of LLMs as text-based decision-support tools but do not support autonomous clinical diagnosis or counseling.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Runze Yu, Miao Peng, Huiying Li, Simeng Liu, Tong Su, Hanjie Guan, Yufeng Shen, Yuhui Deng, Deli Zhao
- Quelle
- Medicine
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 0025-7974, 1536-5964
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Runze Yu, Miao Peng, Huiying Li, Simeng Liu, Tong Su, Hanjie Guan, Yufeng Shen, Yuhui Deng, Deli Zhao (2026). Comparative evaluation of large language models for text-based diagnostic reasoning in fetal central nervous system MRI. Medicine. https://doi.org/10.1097/md.0000000000050485
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1