EUVIMEDEuropean Health Evidence
Uhr 7/7Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Benchmarking Large Language Models in Complex Hypersomnolence Disorders: A Comparative Clinical Analysis of ChatGPT and NotebookLM on Narcolepsy Guidelines

Volkan Tekin, Mehmet Koçer

The European Research Journal · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Objective: To compare the accuracy of 2 artificial intelligence models, ChatGPT and NotebookLM, in answering clinical questions regarding narcolepsy management. Methods: A set of 30 clinical questions was developed based on 2 reference documents: the 2021 European guideline on the management of narcolepsy in adults and children, and the American Academy of Sleep Medicine clinical practice guideline for the treatment of central disorders of hypersomnolence. Both models were queried with the question set. Three independent scorers evaluated the responses across 4 categories (Accuracy, Evidence Reasoning, Additional Information, and Information Integration) using a binary scale (1 = criterion met, 0 = criterion not met). Interrater reliability was assessed using Cohen kappa and Fleiss kappa. Performance comparisons between the models were analyzed using the McNemar test, Wilcoxon signed-rank test, and Mann-Whitney U test, while scorer consistency was checked via the Friedman test. Results: ChatGPT achieved 261 of 360 positive ratings (72.5%), whereas NotebookLM achieved 179 of 360 (49.7%). ChatGPT scored significantly higher than NotebookLM in Accuracy (P < .001) and Evidence Reasoning (P < .001). No statistically significant differences were observed between the 2 models in Additional Information or Information Integration. Conclusion: ChatGPT demonstrated superior accuracy and internal consistency in answering narcolepsy-related clinical questions compared with NotebookLM. However, neither model showed high proficiency in providing necessary additional clinical information.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Volkan Tekin, Mehmet Koçer
Quelle
The European Research Journal
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2149-3189
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Volkan Tekin, Mehmet Koçer (2026). Benchmarking Large Language Models in Complex Hypersomnolence Disorders: A Comparative Clinical Analysis of ChatGPT and NotebookLM on Narcolepsy Guidelines. The European Research Journal. https://doi.org/10.18621/eurj.1363
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1