Vollständiger Abstract
Worum geht es in dieser Arbeit?
Background The rapid development of large language models (LLMs) has raised questions about the continued educational value of conventional assessment formats in health professions education. This study evaluated how examination type, search access, image-based question characteristics, and item format influenced LLM accuracy in Japanese national licensing examinations for physicians, dentists, and pharmacists. Methods Questions from the 2025 academic-year Japanese national licensing examinations administered in early 2026 were analyzed, including 400 medical, 360 dental, and 345 pharmacist questions. Multiple LLMs were assessed under search-on and search-off conditions. Questions were classified by examination type, image presence, image type, and item format. Accuracy was determined using official answer keys, and item-level comparisons were performed using McNemar and chi-square tests. Holm-adjusted sensitivity analyses were conducted to account for multiple comparisons. Results Accuracy was highest in the medical examination, intermediate in the pharmacist examination, and lowest in the dental examination. Search access did not consistently improve performance and significantly reduced accuracy in selected model–examination combinations. Image-based questions substantially decreased accuracy, particularly in the dental examination, with reductions of approximately 17 to 28 percentage points; these dental image effects remained significant after Holm adjustment. In contrast, most item-format differences did not remain statistically significant after correction. Conclusions LLM performance in licensing examinations was strongly influenced by domain, search access, and visual characteristics, whereas associations with response format were less consistent and model-dependent. These findings may inform assessment design and AI literacy education by clarifying how domain, visual demands, search access, and response structure influence LLM performance. Future assessments should prioritize construct-relevant professional competencies and evaluate both independent reasoning and the critical, ethical use of AI.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Toshitsugu Sakurai, Daichi Aizawa, Kazuyoshi Okawa, Ryo Kofuchi, Masatsugu Hirota, Hiroya Gotouda, Takatsugu Yamamoto, Chikahiro Ohkubo
- Quelle
- Journal of Medical Education and Curricular Development
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2382-1205, 2382-1205
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Toshitsugu Sakurai, Daichi Aizawa, Kazuyoshi Okawa, Ryo Kofuchi, Masatsugu Hirota, Hiroya Gotouda, Takatsugu Yamamoto, Chikahiro Ohkubo (2026). Assessment Design in the Era of Large Language Models: Evidence From Japanese Health Professions Licensing Examinations. Journal of Medical Education and Curricular Development. https://doi.org/10.1177/23821205261483568
Kontext