Vollständiger Abstract
Worum geht es in dieser Arbeit?
Abstract Depression is a leading contributor to the global burden of disease and a significant barrier to both personal well-being and societal development. Yet, it remains underdiagnosed, primarily due to social stigma and limited access to clinical resources. This study proposes a dual-modality framework for automated depression detection using speech signals, capturing both what individuals say (text) and how they say it (vocal cues) from speech recordings. The method integrates contextual embeddings and acoustic features extracted from clinical interview recordings. Transcriptions are generated using the Whisper automatic speech recognizer, and text embeddings are derived from a fine-tuned BERT model, while prosodic and spectral features serve as vocal cues. To improve classification sensitivity under class imbalance, random oversampling is applied during fine-tuning. Both feature-level and decision-level fusion strategies are evaluated using fully connected and Long Short-Term Memory (LSTM) networks. On the E-DAIC dataset, the decision-level OR fusion strategy achieved 96.68% accuracy, 90.41% precision, 95.17% recall, and a 92.73% F1-score. Evaluation on the Chinese MODMA dataset assesses the applicability of the framework to an additional language-specific dataset after dataset-specific training, while indicating that modality contribution may vary across datasets. The reported results are segment-level methodological findings and should not be interpreted as participant-independent clinical performance. The speech-only design avoids visual-data collection but does not eliminate the privacy risks associated with identifiable speech recordings and transcripts.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Na Wang, Jing Li, Weijia Zhang, Raja Kamil, Ian Renner, Syed Abdul Rahman Al-Haddad, Normala Ibrahim, Zhen Zhao
- Quelle
- Scientific Reports
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2045-2322
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Na Wang, Jing Li, Weijia Zhang, Raja Kamil, Ian Renner, Syed Abdul Rahman Al-Haddad, Normala Ibrahim, Zhen Zhao (2026). Dual-modality modeling for depression detection using speech signals. Scientific Reports. https://doi.org/10.1038/s41598-026-68971-z
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1