EUVIMEDEuropean Health Evidence
Uhr 7/7Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Reasoning vs. conventional large language models for BI-RADS educational questions answering: a multi-model comparative evaluation

Yuxia Tang, Yuting Liu, Mengxuan Liu, Hao Ni, Siqi Wang, Shouju Wang

Frontiers in Medicine · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Objective To compare reasoning vs. conventional large language models (LLMs) in generating answers with guideline-aligned explanations for Breast Imaging Reporting and Data System (BI-RADS) educational questions. Methods In this prospective study performed from February 6 to 12, 2025, 49 English-Chinese question pairs were extracted from BI-RADS Atlas Fifth Edition. Two reasoning LLMs (ChatGPT-o1, Deepseek-R1) and six conventional LLMs (Gemini2.0-Flash, Deepseek-V3, ChatGPT-4o, ChatGPT-3.5, Qwen-2.5, and WenXinYiYan-3.5) generated answers and explanations to the questions through structured prompts. Three radiologists specialized in breast imaging independently evaluated responses using a 5-point Likert scale, with reference to standard answers. Results The reasoning LLMs significantly outperformed conventional models (median [interquartile range (IQR)]: 3.7 [2.7–4.0] vs. 2.7 [2.0–3.7], P < 0.001), with ChatGPT-o1 and Deepseek-R1 demonstrating peak performance. Both categories of LLMs exhibited significant score reductions in handling questions with multifaceted clinical scenarios (reasoning models: median 4.0 [2.7–4.3] vs. 2.7 [2.3–2.7], Δ median = −1.3, P < 0.001; conventional models: 2.7 [2.0–3.7] vs. 2.3 [2.0–2.7], Δ median = −0.4, P < 0.001). While question language showed no significant impact on reasoning LLMs (ChatGPT-o1 and Deepseek-R1), it affected some conventional models (ChatGPT-3.5, Deepseek-V3 and Gemini2.0-Flash). LLMs performance remained independent of question section and question type. Conclusions Reasoning LLMs show significant potential for BI-RADS guideline explanation and education, but require specific optimization for complex clinical scenario instruction.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Yuxia Tang, Yuting Liu, Mengxuan Liu, Hao Ni, Siqi Wang, Shouju Wang
Quelle
Frontiers in Medicine
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2296-858X
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Yuxia Tang, Yuting Liu, Mengxuan Liu, Hao Ni, Siqi Wang, Shouju Wang (2026). Reasoning vs. conventional large language models for BI-RADS educational questions answering: a multi-model comparative evaluation. Frontiers in Medicine. https://doi.org/10.3389/fmed.2026.1926969
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1