Vollständiger Abstract
Worum geht es in dieser Arbeit?
Abstract Background This study aimed to evaluate the performance of contemporary Large Language Models (LLMs) on the clinical sciences component of the Turkish Dental Specialization Examination (DUS) by comparing their accuracy across disciplines, examination years, and question types. Methods The source dataset contained 1,040 scheduled clinical-science questions from the official DUS examinations (2012/1–2021). Thirteen officially cancelled items were excluded, leaving 1,027 evaluable questions across eight disciplines (871 MCQs, 119 CMCQs, and 37 IBMCQs). The same questions were submitted once to seven consumer AI interfaces. Overall and exploratory subgroup comparisons were performed using Cochran’s Q test. Only the overall comparison was followed by 21 exact pairwise McNemar tests with Holm adjustment. Wilson 95% confidence intervals were calculated for overall accuracy. Results Overall accuracy differed among the seven interfaces (Cochran’s Q = 669.10; p < 0.001). Gemini 2.5 Pro achieved the highest accuracy (92.9%; 954/1,027), whereas Qwen3-Max achieved the lowest (59.5%; 611/1,027). Exploratory Cochran’s Q tests indicated omnibus differences among interfaces within every dental discipline (all p < 0.001) and within MCQ (Q = 669.91; p < 0.001), CMCQ (Q = 59.29; p < 0.001), and IBMCQ groups (Q = 14.68; p = 0.023). The IBMCQ result should be interpreted cautiously because only 37 items were available. Conclusions Under standardized, single-attempt Turkish DUS testing conditions, the evaluated consumer AI interfaces showed materially different accuracy across the full dataset and exploratory discipline- and format-based subgroups. These findings characterize examination performance only and require confirmation with novel, unpublished, clinically contextualized, and independently assessed questions before broader educational or clinical utility can be inferred.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Fatih Karaaslan, Muhammed Halil Yılan, Merve Tirimoğulları
- Quelle
- BMC Oral Health
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 1472-6831
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Fatih Karaaslan, Muhammed Halil Yılan, Merve Tirimoğulları (2026). Accuracy of large language models in the Turkish dental specialization examination (DUS): a multidimensional evaluation across disciplines and question formats. BMC Oral Health. https://doi.org/10.1186/s12903-026-09758-6
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1