Vollständiger Abstract
Worum geht es in dieser Arbeit?
<h4>Background</h4>While large language models (LLMs) have been widely adopted in nursing education and several studies have evaluated their performance on the Chinese National Nursing Licensing Examination (NNLE), systematic comparisons focusing on the latest generation of Chinese LLMs such as DeepSeek-V3, Doubao, ERNIE 4.5 Turbo, and iFLYTEK Spark remain limited. Furthermore, no study has specifically examined the multimodal capabilities of these models using image-based nursing questions.<h4>Objective</h4>This study aimed to compare the performance of DeepSeek-V3, Doubao, ERNIE 4.5 Turbo, and iFLYTEK Spark LLMs on the NNLE and evaluate their potential for nursing education.<h4>Methods</h4>This cross-sectional study assessed DeepSeek-V3, Doubao, ERNIE 4.5 Turbo, and iFLYTEK Spark using questions from the 2025 NNLE, including both authentic text-based items (<i>n</i> = 240) and image-based simulation items (<i>n</i> = 80). We evaluated model performance across several dimensions, including overall accuracy, accuracy by question type (A1, A2, A3, A4), accuracy on case analysis versus non-case analysis questions, accuracy on text-based versus image-based questions and accuracy on disciplinary background (nursing vs. non-nursing). All reported accuracies are presented with 95% confidence intervals. Additional metrics included generation speed, session capacity, accessibility, and content modality support (text, image).<h4>Results</h4>On text-based NNLE items, all four models achieved accuracy rates above 90% (DeepSeek-V3: 92.9, 95% CI [89.0, 95.6%]; Doubao: 94.6, 95% CI [91.0, 96.9%]; ERNIE 4.5 Turbo: 93.3, 95% CI [89.5, 95.9%]; iFLYTEK Spark: 92.1, 95% CI [88.0, 95%]), meeting the NNLE passing requirement. No significant differences in overall accuracy were observed among models (all pairwise comparisons <i>p</i> > 0.05). On image-based simulation items, performance declined substantially across all multimodal-capable models (Doubao: 48, 95% CI [37.2, 58.9%]; ERNIE 4.5 Turbo: 48, 95% CI [37.2, 58.9%]; iFLYTEK Spark: 61, 95% CI [50, 71.4%]), with DeepSeek-V3 unable to process image inputs due to its text-only architecture. For the image-based simulation test items, the performance difference among the three models was not statistically significant (χ 2 = 4.040, <i>p</i> = 0.133). The difference between text-based and image-based performance was statistically significant for all models (<i>p</i> < 0.01). Subgroup analyses based on question type, case status, and disciplinary content revealed variable performance patterns, though these findings should be interpreted with caution given limited sample sizes in certain categories.<h4>Conclusion</h4>This study demonstrates that DeepSeek-V3, Doubao, ERNIE 4.5 Turbo, and iFLYTEK Spark perform well on Chinese nursing examinations questions in text format, indicating their potential as auxiliary resources for nursing education and exam preparation. However, their performance on image-based items is substantially lower, and the findings do not support claims of clinical applicability. Further research is needed to assess their utility in real-world clinical reasoning or patient care.
Abstract: PubMed · Datensatz
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Nicht angegeben
- Quelle
- CrossRef Listing of Deleted DOIs
- Publikation
- 2000-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 0849-6757
- Zitationen
- 14 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
(2000). 10.3389/fpsyg.2012.00132. CrossRef Listing of Deleted DOIs. https://doi.org/10.3389/frai.2026.1870847