Vollständiger Abstract
Worum geht es in dieser Arbeit?
Abstract Automated auscultation using wearable devices is essential for remote respiratory monitoring, but deep learning models often struggle to generalize due to the severe scarcity of annotated abnormal respiratory sounds. Since directly using synthetic data for supervised training risks learning artifacts instead of true pathological features, we propose a robust three-phase pipeline leveraging massive synthetic data without synthetic labels. First, a modified StyleGAN2 natively synthesizes rectangular Mel-spectrograms to preserve high-temporal-resolution acoustic characteristics, validated by kernel audio distance. Second, 100,000 synthetic spectrograms are used for unsupervised variational autoencoder pre-training, introducing a parallel asymmetric convolutional block to independently capture distinct time and frequency semantics. Finally, the encoder is repurposed as a feature extractor to train a lightweight classifier on limited clinical data. Empirical evaluations demonstrate that while a serial asymmetric kernel geometry yields the highest classification accuracy under frequency-dominant pathologies, our parallel architecture achieves highly competitive performance using only 43% of conventional square baseline parameters and rivals the massive pretrained audio neural networks CNN14 model using merely 0.5% of its convolutional footprint. This framework provides a potential scalable pathway for domains with critically limited annotated data.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Takehiro Hirasawa, Yasumasa Tamura, Kaoruko Shimizu, Satoshi Konno, Masahito Yamamoto
- Quelle
- Artificial Life and Robotics
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 1433-5298, 1614-7456
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Takehiro Hirasawa, Yasumasa Tamura, Kaoruko Shimizu, Satoshi Konno, Masahito Yamamoto (2026). Domain-specific unsupervised pre-training for robust respiratory sound classification. Artificial Life and Robotics. https://doi.org/10.1007/s10015-026-01151-4
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1