Vollständiger Abstract
Worum geht es in dieser Arbeit?
Background Patients with diabetes increasingly consult artificial intelligence (AI) chatbots for medical advice, including guidance on antidiabetic medication management during Ramadan fasting, because AI can simplify and summarize long, complex guidelines. Also, in hospital settings, these tools are being used in hospitals much faster than it takes to establish formal regulations and guidelines for their use. Evaluations of the accuracy, completeness, and reproducibility of such advice across languages are still lacking. Therefore, the study aims to evaluate and compare the accuracy, completeness, safety, and reproducibility of three widely used AI chatbots—ChatGPT, Google Gemini, and Microsoft Copilot—when providing antidiabetic medication adjustment advice during Ramadan in both English and Arabic. Methods Twenty-three standardized clinical scenarios covering common antidiabetic regimens were presented to each chatbot in both English and Arabic. Each query was repeated to evaluate reproducibility, resulting in 276 responses scored. Responses were assessed against the International Diabetes Federation–Diabetes and Ramadan (IDF-DAR) Guidelines using a 0–2 accuracy scale, a 0–4 completeness scale, and a 0–3 safety scale. Results Overall, 77% of responses were fully consistent with the guideline, 12% were partially consistent, and 11% (30/276) contained clinically harmful or contradictory advice; harmful responses were about twice as common in Arabic as in English (14% vs. 8%). Completeness and safety were high, with medians at the observed ceiling. In the generalized linear mixed models, chatbots did not differ significantly in accuracy, completeness, or safety, and there was no significant main effect of language or chatbot × language interaction; the strongest signals were a chatbot effect on completeness ( p = 0.068) and a language effect on safety ( p = 0.064), both non-significant. Two-week reproducibility was fair for accuracy (weighted κ = 0.20, p = 0.009) and completeness (κ = 0.29, p = 0.001) and showed a very low κ in the safety scale (κ = 0.02, p = 0.81). Conclusions AI chatbots demonstrated comparable performance in delivering guideline-based advice for diabetes management during Ramadan, with no significant differences in accuracy, completeness, or safety. While most responses aligned with the IDF-DAR guideline, some harmful recommendations persisted, and response consistency fluctuated over time. These results suggest that AI chatbots should serve as a supplementary resource rather than a substitute for professional medical advice.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Sufyan Alomair, Maryam Alsuwayq, Walla Alabbad, Khawlah Albeladi, Zainab Alhassan
- Quelle
- Frontiers in Medicine
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2296-858X
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Sufyan Alomair, Maryam Alsuwayq, Walla Alabbad, Khawlah Albeladi, Zainab Alhassan (2026). AI Chatbots as a source of Ramadan medication management advice for patients with diabetes: a multilingual comparative evaluation. Frontiers in Medicine. https://doi.org/10.3389/fmed.2026.1934432
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1