Vollständiger Abstract
Worum geht es in dieser Arbeit?
Objective: To compare the accuracy of 2 artificial intelligence models, ChatGPT and NotebookLM, in answering clinical questions regarding narcolepsy management. Methods: A set of 30 clinical questions was developed based on 2 reference documents: the 2021 European guideline on the management of narcolepsy in adults and children, and the American Academy of Sleep Medicine clinical practice guideline for the treatment of central disorders of hypersomnolence. Both models were queried with the question set. Three independent scorers evaluated the responses across 4 categories (Accuracy, Evidence Reasoning, Additional Information, and Information Integration) using a binary scale (1 = criterion met, 0 = criterion not met). Interrater reliability was assessed using Cohen kappa and Fleiss kappa. Performance comparisons between the models were analyzed using the McNemar test, Wilcoxon signed-rank test, and Mann-Whitney U test, while scorer consistency was checked via the Friedman test. Results: ChatGPT achieved 261 of 360 positive ratings (72.5%), whereas NotebookLM achieved 179 of 360 (49.7%). ChatGPT scored significantly higher than NotebookLM in Accuracy (P < .001) and Evidence Reasoning (P < .001). No statistically significant differences were observed between the 2 models in Additional Information or Information Integration. Conclusion: ChatGPT demonstrated superior accuracy and internal consistency in answering narcolepsy-related clinical questions compared with NotebookLM. However, neither model showed high proficiency in providing necessary additional clinical information.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Volkan Tekin, Mehmet Koçer
- Quelle
- The European Research Journal
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2149-3189
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Volkan Tekin, Mehmet Koçer (2026). Benchmarking Large Language Models in Complex Hypersomnolence Disorders: A Comparative Clinical Analysis of ChatGPT and NotebookLM on Narcolepsy Guidelines. The European Research Journal. https://doi.org/10.18621/eurj.1363
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1