Vollständiger Abstract
Worum geht es in dieser Arbeit?
Background/Objectives: Artificial intelligence (AI)-based chatbots are increasingly used as sources of health information, yet the quality and semantic stability of their responses may vary across systems, languages, repeated queries, and prompt formulations. This study compared expert-rated general information quality and embedding-based semantic stability of responses generated by seven AI chatbot systems to expert-derived, patient-oriented complete denture questions in Turkish and English. Methods: Twenty-two complete denture-related questions were submitted to seven chatbot systems on three study days across scheduled morning, afternoon, and evening sessions. Two additional semantically equivalent Turkish variants were generated for each original question. Five prosthodontists assessed the original Turkish and English responses using the 5-point Global Quality Score (GQS). Semantic stability was quantified using the multilingual SentenceTransformer checkpoint paraphrase-multilingual-MiniLM-L12-v2 and cosine similarity. To incorporate complete responses, responses were divided into non-overlapping token-based chunks, chunk embeddings were combined into one normalized full-response representation, and the 36 pairwise similarities from the nine repeated responses were averaged to one stability estimate per question–model–language/variant condition. Linear mixed-effects models with question-level clustering were used, with Bonferroni adjustment and partial eta squared effect sizes with 95% confidence intervals. Results: AI model significantly affected embedding-based semantic stability (p < 0.001; ηp2 = 0.820, 95% CI 0.795–0.837) and GQS (p < 0.001; ηp2 = 0.249, 95% CI 0.152–0.316). Grok had the numerically highest overall semantic stability mean (0.952 ± 0.015), whereas GPT-4o had the numerically highest overall GQS (4.48 ± 0.85). English responses showed higher overall GQS than the original Turkish responses (3.96 ± 0.94 vs. 3.66 ± 1.23; p = 0.004) and higher embedding-based similarity than the Turkish conditions overall (p < 0.001). No significant overall effect of scheduled study day was observed (p = 0.504; ηp2 = 0.001), whereas semantic stability differed across scheduled query sessions (p < 0.001; ηp2 = 0.055), with means of 0.863 ± 0.106, 0.874 ± 0.092, and 0.841 ± 0.140 for morning, afternoon, and evening sessions, respectively. Inter-rater reliability was moderate for English GQS ratings (ICC = 0.675, 95% CI 0.581–0.752) and good for Turkish ratings (ICC = 0.771, 95% CI 0.687–0.833). Conclusions: The evaluated chatbot systems differed in expert-rated information quality and full-response embedding-based semantic stability. English responses showed higher overall values than Turkish responses, although cross-language measurement effects of the embedding model cannot be excluded. Differences across scheduled query sessions should be interpreted as run-to-run output variability rather than intrinsic temporal behavior. These findings characterize comparative chatbot performance under the tested conditions but do not establish clinical accuracy, safety, or patient education effectiveness.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Sinan Coşkun, Fatma Nur Karaman, Hüseyin Ardıl Uytun, Gülcan Coşkun Akar
- Quelle
- Healthcare
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2227-9032
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Sinan Coşkun, Fatma Nur Karaman, Hüseyin Ardıl Uytun, Gülcan Coşkun Akar (2026). Information Quality and Semantic Stability of AI Chatbot Responses to Complete Denture Questions: A Turkish–English Benchmarking Study. Healthcare. https://doi.org/10.3390/healthcare14172797
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1