EUVIMEDEuropean Health Evidence
Uhr 7/7Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Safety, accuracy, empathy, information quality, and readability of publicly accessible LLM-based chatbots for traumatic brain injury and concussion questions: a cross-sectional comparative study

Xin Zuo, Huan Zuo, Min Zhang, Weihong Zheng

Frontiers in Public Health · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Background Large language model (LLM)-based chatbots are increasingly used by the public to obtain health information, but their performance in answering questions related to traumatic brain injury (TBI) and concussion remains unclear. This study evaluated five publicly accessible LLM-based chatbots across safety, accuracy, empathy, information quality, transparency, and readability. Methods Sixty-five English-language, public-facing questions about TBI and concussion were submitted to ChatGPT, Gemini, Copilot, DeepSeek, and Doubao, generating 325 responses. Five blinded independent raters assessed the generated responses using guideline-informed criteria and established tools, including DISCERN, EQIP, JAMA benchmark criteria, the Global Quality Score, and readability indices. Results Inter-rater agreement was high. The proportions of responses rated as safe ranged from 89.2 to 95.4%, and all models achieved a median accuracy score of 4.00. However, 27 responses were classified as potentially harmful, mainly because of under-triage, overly reassuring advice regarding imaging findings, premature return-to-activity or driving guidance, and insufficient pediatric caution. Accuracy differences were statistically significant but small, whereas empathy, information quality, transparency, and readability showed clearer model-level variation. DeepSeek produced the easiest-to-read responses. Conclusion LLM-based chatbots generated responses that were generally rated as safe and informative by expert evaluators, but potentially harmful advice and readability problems remained. These findings characterize expert-rated response performance and do not establish patient comprehension, educational effectiveness, or clinical benefit. LLM-based chatbots may have potential as supplementary sources of patient-facing health information, but they should not replace professional medical evaluation.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Xin Zuo, Huan Zuo, Min Zhang, Weihong Zheng
Quelle
Frontiers in Public Health
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2296-2565
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Xin Zuo, Huan Zuo, Min Zhang, Weihong Zheng (2026). Safety, accuracy, empathy, information quality, and readability of publicly accessible LLM-based chatbots for traumatic brain injury and concussion questions: a cross-sectional comparative study. Frontiers in Public Health. https://doi.org/10.3389/fpubh.2026.1926980
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1