Vollständiger Abstract
Worum geht es in dieser Arbeit?
Abstract In this paper, we propose an advanced speech enhancement model capable of effectively separating clean speech from noisy audio signals. The primary objective here is to improve speech intelligibility and quality in noisy environments while preserving critical speech components. We propose a GAN based novel residual learning architecture with long range temporal context using complex spectrograms. This is achieved by combining residual neural networks with dilated temporal processing. The model processes both magnitude and phase components in the time-frequency domain, implementing residual learning for stable training and a dilated temporal module to capture long-range dependencies in speech signals. Experiments were conducted on benchmark of VoiceBank-DEMAND dataset under various noise conditions. During experimentation, the proposed model demonstrated superior performance compared to existing ones, achieving 3.43 dB improvement in perceptual evaluation of speech quality measurement. The novelties include a phase-aware complex masking system which jointly optimizes magnitude and phase for more natural-sounding output and incorporation of a Dilated temporal module. This module further expands receptive fields exponentially in order to model speech dynamics without increasing parameters drastically. Residual architecture also balances deep feature extraction, computational efficiency and the composite loss function that uniquely combine in delivering superior results. Analysis of our results indicates that this multi-objective optimization framework achieves superior speech enhancement by simultaneously addressing magnitude and phase reconstruction, waveform-level distortion, and psychoacoustic naturalness, a balance not achieved in any prior works so far.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Debabrata Gogoi, Sushanta Kabir Dutta
- Quelle
- Engineering Research Express
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2631-8695
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Debabrata Gogoi, Sushanta Kabir Dutta (2026). A deep residual complex learning framework with long-range temporal context for phase-aware speech enhancement. Engineering Research Express. https://doi.org/10.1088/2631-8695/ae9cf1
Kontext