EUVIMEDEuropean Health Evidence
Uhr 7/7Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Comparing Chinese-language large language models for caregiver questions about developmental dysplasia of the hip

Fangyuan Wang, Xianhong Li, Yinghui Guan, Jiaqi Tian, Le Xu, Bing Huang, Jianhui Xie

Frontiers in Pediatrics · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Background Large language models (LLMs) are increasingly used for caregiver-facing health information, but their reliability in Chinese-language pediatric orthopaedics remains uncertain. This study evaluated whether responses to developmental dysplasia of the hip (DDH) questions were clinically accurate, aligned with Chinese guidance, and educationally usable. Methods We compared ChatGPT-4o and DeepSeek-R1 using 53 Chinese-language DDH questions, including 31 caregiver-oriented frequently asked questions and 22 guideline-derived items. Each model generated one response per question using a standardized single-turn, five-sentence prompt. Six blinded pediatric orthopaedic surgeons rated clinical accuracy and guideline concordance. Paired model comparisons, inter-rater reliability, and exploratory formula-based readability were assessed. Results DeepSeek-R1 had higher clinical accuracy than ChatGPT-4o across 31 caregiver-oriented questions (mean question-level score 4.70 [SD 0.15] vs. 3.69 [0.28]; P < 0.001) and higher guideline concordance across 22 items (4.44 [0.22] vs. 3.82 [0.33]; P < 0.001). Average-rating reliability was good for clinical accuracy (ICC(2,6) = 0.826, 95% bootstrap CI 0.770−0.863) and moderate for guideline concordance (ICC(2,6) = 0.519, 95% CI 0.377−0.621). Formula-based readability findings were mixed: unadjusted tests indicated lower FKGL, FOG, and SMOG values for ChatGPT-4o, but only FOG and SMOG remained statistically significant after Holm adjustment. Surgery-related observations were exploratory and descriptive. Conclusions Under the tested single-query and five-sentence conditions, DeepSeek-R1 achieved higher expert-rated clinical accuracy and Chinese-guideline concordance. The findings describe specific model-platform configurations rather than a permanent model ranking. Professionally reviewed LLM responses may help clinicians reinforce routine DDH education, but individualized diagnostic and treatment guidance, especially for surgery-related questions, should remain clinician-led. Caregivers should use chatbot information only as a supplementary resource.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Fangyuan Wang, Xianhong Li, Yinghui Guan, Jiaqi Tian, Le Xu, Bing Huang, Jianhui Xie
Quelle
Frontiers in Pediatrics
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2296-2360
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Fangyuan Wang, Xianhong Li, Yinghui Guan, Jiaqi Tian, Le Xu, Bing Huang, Jianhui Xie (2026). Comparing Chinese-language large language models for caregiver questions about developmental dysplasia of the hip. Frontiers in Pediatrics. https://doi.org/10.3389/fped.2026.1919701
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1