Vollständiger Abstract
Worum geht es in dieser Arbeit?
Abstract Background Mental health chatbots are increasingly used to support people with depressive symptoms, and large language models make these systems more flexible than rule-based chatbots. However, it remains unclear how well large language model–based chatbots deliver structured psychological interventions. Objective This study examined how well a GPT-4o–based chatbot delivered a behavioral activation intervention for young people with depression using sessions with artificial users and clinical expert assessment. It also identified limitations and potential refinements. Methods We implemented a GPT-4o (gpt-4o-2024-08-06; OpenAI)–based chatbot using a structured system prompt to deliver a single-session behavioral activation intervention for people with depression aged 14 to 29 years. We generated 48 sessions with GPT-4o–based artificial users derived from clinical vignettes varying across 7 characteristics. Ten clinical experts, either licensed psychotherapists or advanced psychotherapy trainees, independently assessed the sessions using the 14-item Quality of Behavioral Activation Scale (Q-BAS), rated from 0 to 6, supplemented by rating therapeutic capabilities, artificial user authenticity and difficulty, and qualitative feedback. Results The chatbot completed all 7 intervention phases in every session. The mean holistic session quality rating was 3.94 (SD 1.23), and the mean Q-BAS rating was 4.03 (SD 1.18). Thirteen of 14 Q-BAS components exceeded the satisfactory threshold of 3. Ratings were highest for mood assessment (mean 5.42, SD 1.09) and activity planning (mean 4.98, SD 1.41) and lowest for explaining positive reinforcement (mean 2.92, SD 2.30) and supporting activity-mood monitoring (mean 3.02, SD 2.04). Therapeutic capability ratings were highest for message safety (mean 5.90, SD 0.37), message clarity (mean 5.56, SD 0.77), and objective, nonjudgmental communication (mean 5.17, SD 1.04) and lowest for therapeutic rapport (mean 4.12, SD 1.45) and natural conversation flow (mean 4.25, SD 1.42). Artificial users were rated below the scale midpoint for authenticity (mean 2.75, SD 1.41) and difficulty (mean 1.23, SD 1.46). Clinical experts described the chatbot as structured, clear, and safe but identified insufficient clinical reasoning as the main limitation, particularly in evaluating the therapeutic suitability and feasibility of activities, barriers, solution strategies, and rewards. Artificial users were often highly compliant, especially when identifying positive activities. Conclusions In expert-rated sessions with artificial users, the chatbot delivered the behavioral activation intervention as intended and performed strongest on procedural components. It performed less well on positive reinforcement and activity-mood monitoring, indicating refinement needs in clinical reasoning, follow-up questioning, and evaluating whether proposed activities, plans, barriers, solution strategies, and rewards are therapeutically appropriate and feasible. The findings identify targets for improvement before testing with human users, while the artificial user design and expert ratings limit conclusions about real therapeutic interactions.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Florian Onur Kuhlmeier, Leon Hanschmann, Melina Rabe, Stefan Lüttke, Eva-Lotta Brakemeier, Alexander Maedche
- Quelle
- JMIR Mental Health
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2368-7959
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Florian Onur Kuhlmeier, Leon Hanschmann, Melina Rabe, Stefan Lüttke, Eva-Lotta Brakemeier, Alexander Maedche (2026). Large Language Model–Based Behavioral Activation Chatbot for Young People With Depression Using Artificial Users and Clinical Experts: Mixed Methods Evaluation. JMIR Mental Health. https://doi.org/10.2196/94781