EUVIMEDEuropean Health Evidence
Uhr 7/7Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Refusal-aware evaluation of frontier AI models available in June 2026 using the Japanese National License Examination for Pharmacists: a comparative study

Hiroyasu Sato, Katsuhiko Ogasawara, Hidehiko Sakurai

Journal of Educational Evaluation for Health Professions · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Purpose: Conventional single-run accuracy may be insufficient for evaluating frontier generative artificial intelligence (AI) models when safety-related refusals occur. This 4-model benchmark examined the need for repeated, refusal-aware evaluation using the Japanese National License Examination for Pharmacists (JNLEP).Methods: ChatGPT GPT-5.5, Gemini 3.5 Flash, Claude Opus 4.8, and Claude Fable 5 were evaluated using all 345 questions from the 107th JNLEP. The original Japanese questions, including image-containing items, were submitted through application programming interfaces (APIs) in 3 independent runs. Refusals were treated as incorrect when overall accuracy was calculated. For Fable 5, accuracy excluding refusals, refusal rate, refusal consistency across runs, the subject-wise distribution of refusals, and system-assigned refusal categories were also evaluated.Results: The mean overall accuracies were 98.7% for GPT-5.5, 98.3% for Gemini 3.5 Flash, 96.1% for Claude Opus 4.8, and 70.3% for Claude Fable 5. Fable 5 had a mean refusal rate of 29.0%, whereas its mean accuracy excluding refusals was 99.0%. Among the 345 items, 93 were refused in all 3 runs, 14 were refused inconsistently across runs, and 238 were never refused. All 300 refusal responses were assigned to the bio category. Refusals were most frequent in Biology (90.0%) and Pharmacology (69.2%) but uncommon in Practice (2.5%).Conclusion: Near-saturation benchmark performance coexisted with frequent and partly run-dependent refusals in a safeguard-equipped model. Overall accuracy, accuracy excluding refusals, refusal rate, and refusal consistency describe complementary aspects of performance; therefore, repeated, refusal-aware evaluation is needed to interpret frontier AI models in pharmacy education.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Hiroyasu Sato, Katsuhiko Ogasawara, Hidehiko Sakurai
Quelle
Journal of Educational Evaluation for Health Professions
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
1975-5937
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Hiroyasu Sato, Katsuhiko Ogasawara, Hidehiko Sakurai (2026). Refusal-aware evaluation of frontier AI models available in June 2026 using the Japanese National License Examination for Pharmacists: a comparative study. Journal of Educational Evaluation for Health Professions. https://doi.org/10.3352/jeehp.2026.23.29
RIS BibTeX CSL-JSON