Vollständiger Abstract
Worum geht es in dieser Arbeit?
Background Bladder cancer ranks among the most prevalent urological tumors worldwide, with its global incidence continuing to rise steadily. Although patient education materials (PEMs) play a crucial role in enhancing disease comprehension and supporting joint clinical decision-making, current online resources frequently surpass the readability thresholds recommended for the general public. Large language models (LLMs) hold promise for health communication, yet no systematic assessment has been conducted regarding their feasibility and trustworthiness specifically for bladder cancer patient education. Objective This study aimed to systematically benchmark five leading LLMs in producing question-and-answer content for bladder cancer science popularization, with a particular focus on readability, informational quality, and appropriateness for patient education. Methods In this cross-sectional simulation study, 20 common patient questions covering five disease domains were compiled. On January 15, 2026, each question was submitted identically to five publicly available LLMs (Doubao, DeepSeek, Kimi, Gemini, and ChatGPT). Readability was evaluated using seven conventional metrics. Two independent pharmacists, blinded to model identity, rated the responses using the Chinese version of the Patient Education Materials Assessment Tool for print materials (C-PEMAT-P) and the Global Quality Score (GQS). Additionally, two independent clinical specialists assessed factual accuracy and alignment with the Chinese Bladder Cancer Diagnosis and Treatment Guidelines (2024 edition) employing a 4-point scale. Cohen’s kappa was used to determine inter-rater reliability. Results ChatGPT, DeepSeek, and Doubao outperformed Kimi and Gemini on both C-PEMAT and GQS (all P < 0.001), indicating superior understandability, actionability, and overall quality. Across all models, median C-PEMAT scores ranged from 8 to 10, suggesting broadly acceptable suitability for patient education. Readability varied significantly by content domain, with treatment-oriented texts showing the highest complexity. ChatGPT achieved the best alignment with clinical guidelines. No model produced harmful advice or directly contradicted guideline recommendations. Traditional readability measures correlated weakly with GQS, whereas C-PEMAT showed a moderate positive correlation (r = 0.34). Conclusion Current mainstream LLMs demonstrate initial potential for generating educational content on bladder cancer, albeit with considerable heterogeneity across models. Disease-specific evaluation instruments for patient education materials are more effective than general readability formulas in reflecting perceived quality. Our results advocate for a prudent, assistive role of LLMs in health communication under a human-AI collaborative model.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Dian Wan, Youwen Li, Zheng Dong, Chen Dai, Jinghe Ye, Song Li, Sunlu Jiang
- Quelle
- Frontiers in Oncology
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2234-943X
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Dian Wan, Youwen Li, Zheng Dong, Chen Dai, Jinghe Ye, Song Li, Sunlu Jiang (2026). Performance evaluation of large language models in bladder cancer patient education Q&A: a cross-sectional study. Frontiers in Oncology. https://doi.org/10.3389/fonc.2026.1909227
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1