EUVIMEDEuropean Health Evidence
Uhr 10/10Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Detecting temporal clinical change in sequential chest radiography reports: a multi-model and multi-prompt evaluation of large language models

Ertürk Erdağı

Frontiers in Digital Health · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Background Radiology reports contain critical information for monitoring patients' clinical progress and evaluating treatment outcomes. Sequential radiological images are particularly valuable, providing not only current findings but also comparative assessments against previous reports. Automatically detecting temporal changes between free-text reports remains a significant challenge in Natural Language Processing. Methods This study evaluated the performance of Large Language Models in detecting temporal novelty using the LUNGUAGE dataset, which contains sequential chest radiography reports. The comparative multi-model and multi-prompt evaluation was carried out on a fixed, class-stratified subset of 1,500 of these examples. Seven language models from OpenAI, Anthropic, and Google Gemini were tested across five prompt strategies: zero-shot, few-shot, structured reasoning, memory-augmented, and temporal graph prompting. Each model classified target findings as new, improved, worsened, stable, or resolved by evaluating current and previous reports together. Performance was measured using accuracy, macro-F1, weighted-F1, and class-based metrics. The “new” class was analyzed separately due to its clinical importance. Model errors were categorized into clinically interpretable types, including missed new findings, missed resolved findings, temporal errors, and change direction errors. Bootstrap confidence intervals and paired McNemar significance tests assessed performance uncertainty. Results GPT-4.1 emerged as the top-performing model. Its few-shot prompt strategy achieved the best results: 0.8369 accuracy, 0.7888 macro-F1, 0.8322 weighted-F1, and 0.5385 recall for the “new” class. Conclusion These findings demonstrate that Large Language Models hold substantial promise for temporal clinical reasoning in sequential radiology reports, and that performance depends not only on model architecture but also on prompt strategy and the class being evaluated.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Ertürk Erdağı
Quelle
Frontiers in Digital Health
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2673-253X
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Ertürk Erdağı (2026). Detecting temporal clinical change in sequential chest radiography reports: a multi-model and multi-prompt evaluation of large language models. Frontiers in Digital Health. https://doi.org/10.3389/fdgth.2026.1919939
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1