Journal of Educational Evaluation for Health Professions
Refusal-aware evaluation of frontier AI models available in June 2026 using the Japanese National License Examination for Pharmacists: a comparative study
Purpose: Conventional single-run accuracy may be insufficient for evaluating frontier generative artificial intelligence (AI) models when safety-related refusals occur. This 4-model benchmark examined the need for repeated, refusal-aware evaluation using the Japanese National License Examination for Pharmacists (JNLEP).Methods: ChatGPT GPT-5.5, Gemini 3.5 Flash, Claude Opus 4.8, and Claude Fable 5 were evaluated using all 345 questions from the 107th JNLEP. The original Japanese questions, including image-containin …