Vollständiger Abstract
Worum geht es in dieser Arbeit?
Background Large language models (LLMs) are increasingly used for caregiver-facing health information, but their reliability in Chinese-language pediatric orthopaedics remains uncertain. This study evaluated whether responses to developmental dysplasia of the hip (DDH) questions were clinically accurate, aligned with Chinese guidance, and educationally usable. Methods We compared ChatGPT-4o and DeepSeek-R1 using 53 Chinese-language DDH questions, including 31 caregiver-oriented frequently asked questions and 22 guideline-derived items. Each model generated one response per question using a standardized single-turn, five-sentence prompt. Six blinded pediatric orthopaedic surgeons rated clinical accuracy and guideline concordance. Paired model comparisons, inter-rater reliability, and exploratory formula-based readability were assessed. Results DeepSeek-R1 had higher clinical accuracy than ChatGPT-4o across 31 caregiver-oriented questions (mean question-level score 4.70 [SD 0.15] vs. 3.69 [0.28]; P < 0.001) and higher guideline concordance across 22 items (4.44 [0.22] vs. 3.82 [0.33]; P < 0.001). Average-rating reliability was good for clinical accuracy (ICC(2,6) = 0.826, 95% bootstrap CI 0.770−0.863) and moderate for guideline concordance (ICC(2,6) = 0.519, 95% CI 0.377−0.621). Formula-based readability findings were mixed: unadjusted tests indicated lower FKGL, FOG, and SMOG values for ChatGPT-4o, but only FOG and SMOG remained statistically significant after Holm adjustment. Surgery-related observations were exploratory and descriptive. Conclusions Under the tested single-query and five-sentence conditions, DeepSeek-R1 achieved higher expert-rated clinical accuracy and Chinese-guideline concordance. The findings describe specific model-platform configurations rather than a permanent model ranking. Professionally reviewed LLM responses may help clinicians reinforce routine DDH education, but individualized diagnostic and treatment guidance, especially for surgery-related questions, should remain clinician-led. Caregivers should use chatbot information only as a supplementary resource.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Fangyuan Wang, Xianhong Li, Yinghui Guan, Jiaqi Tian, Le Xu, Bing Huang, Jianhui Xie
- Quelle
- Frontiers in Pediatrics
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2296-2360
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Fangyuan Wang, Xianhong Li, Yinghui Guan, Jiaqi Tian, Le Xu, Bing Huang, Jianhui Xie (2026). Comparing Chinese-language large language models for caregiver questions about developmental dysplasia of the hip. Frontiers in Pediatrics. https://doi.org/10.3389/fped.2026.1919701
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1