Vollständiger Abstract
Worum geht es in dieser Arbeit?
Breast cancer remains one of the most common and life threatening cancers worldwide, and early detection is strongly associated with improved survival and reduced treatment burden. This study investigates the ability of Large Language Models to perform diagnostic prediction from structured breast cancer related data. We systematically evaluated 12 LLMs across three public datasets with different clinical characteristics: the Wisconsin Breast Cancer Dataset (WBCD) based on cytological features, the Breast Cancer Coimbra Dataset (BCCD) based on metabolic biomarkers, and the Mammographic Mass Dataset (MMD) based on mammographic attributes. The evaluation covered 13 prompting strategies, including three zero-shot variants, few-shot prompting, and multiple Chain-of-Thought (CoT) and knowledge-enhanced reasoning settings. Performance was assessed using confusion-matrix-based metrics, including accuracy, precision, recall, F1-score, specificity, and Matthews Correlation Coefficient. The results showed that performance was strongly dependent on both dataset type and prompting design, and no single model dominated all tasks. The best model–strategy pairvaried by dataset: Cogito-v1-preview-qwen-32B achieved the highest F1-score on WBCD with an F1-score of 92.00%, GPT 4.1 and GPT 4o on BCCD with an F1-score of 85.39%, and Gemini 2.5 Flash Lite on MMD with an F1-score of 82.91%. Prompt engineering had a substantial effect on outcomes, but its benefit varied across models, with some systems improving under knowledge-enhanced few shot prompting and others performing best under simpler strategies. Comparison with state-of-the-art traditional ML baselines showed that, while LLMs do not yet surpass supervised methods, the performance gap has narrowed substantially, particularly on MMD, where the best single-run gap was 2.04 percentage points in F1 and the mean gap under robustness analysis was approximately 5.3 points. Robustness analysis across multiple few-shot example sets confirmed stable performance on WBCD (F1 = 91.37 ± 0.71%) and BCCD (F1 = 87.46 ± 1.91%), while revealing moderate sensitivity on MMD (F1 = 79.62 ± 2.87%). Although the evaluated LLMs did not outperform traditional supervised models, the study provides a clear performance baseline for future research on structured clinical prediction with language models. The results show that LLMs may offer value as complementary exploratory tools, but their outputs should be interpreted only with expert oversight because clinically significant errors remain.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Habibe Karayiğit, Filiz Kalelioğlu
- Quelle
- International Journal of Interactive Multimedia and Artificial Intelligence
- Publikation
- 2026-08-28
- Band / Ausgabe
- 10 / 1
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 1989-1660
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Habibe Karayiğit, Filiz Kalelioğlu (2026). The Role of Advanced Language Models in MedicalDiagnostics: A Case Study on Breast Cancer Prediction. International Journal of Interactive Multimedia and Artificial Intelligence, 10 (1). https://doi.org/10.9781/hv6dk085
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1