Vollständiger Abstract
Worum geht es in dieser Arbeit?
Objective To compare reasoning vs. conventional large language models (LLMs) in generating answers with guideline-aligned explanations for Breast Imaging Reporting and Data System (BI-RADS) educational questions. Methods In this prospective study performed from February 6 to 12, 2025, 49 English-Chinese question pairs were extracted from BI-RADS Atlas Fifth Edition. Two reasoning LLMs (ChatGPT-o1, Deepseek-R1) and six conventional LLMs (Gemini2.0-Flash, Deepseek-V3, ChatGPT-4o, ChatGPT-3.5, Qwen-2.5, and WenXinYiYan-3.5) generated answers and explanations to the questions through structured prompts. Three radiologists specialized in breast imaging independently evaluated responses using a 5-point Likert scale, with reference to standard answers. Results The reasoning LLMs significantly outperformed conventional models (median [interquartile range (IQR)]: 3.7 [2.7–4.0] vs. 2.7 [2.0–3.7], P < 0.001), with ChatGPT-o1 and Deepseek-R1 demonstrating peak performance. Both categories of LLMs exhibited significant score reductions in handling questions with multifaceted clinical scenarios (reasoning models: median 4.0 [2.7–4.3] vs. 2.7 [2.3–2.7], Δ median = −1.3, P < 0.001; conventional models: 2.7 [2.0–3.7] vs. 2.3 [2.0–2.7], Δ median = −0.4, P < 0.001). While question language showed no significant impact on reasoning LLMs (ChatGPT-o1 and Deepseek-R1), it affected some conventional models (ChatGPT-3.5, Deepseek-V3 and Gemini2.0-Flash). LLMs performance remained independent of question section and question type. Conclusions Reasoning LLMs show significant potential for BI-RADS guideline explanation and education, but require specific optimization for complex clinical scenario instruction.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Yuxia Tang, Yuting Liu, Mengxuan Liu, Hao Ni, Siqi Wang, Shouju Wang
- Quelle
- Frontiers in Medicine
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2296-858X
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Yuxia Tang, Yuting Liu, Mengxuan Liu, Hao Ni, Siqi Wang, Shouju Wang (2026). Reasoning vs. conventional large language models for BI-RADS educational questions answering: a multi-model comparative evaluation. Frontiers in Medicine. https://doi.org/10.3389/fmed.2026.1926969
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1