EUVIMEDEuropean Health Evidence
Uhr 10/10Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

A deep residual complex learning framework with long-range temporal context for phase-aware speech enhancement

Debabrata Gogoi, Sushanta Kabir Dutta

Engineering Research Express · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Abstract In this paper, we propose an advanced speech enhancement model capable of effectively separating clean speech from noisy audio signals. The primary objective here is to improve speech intelligibility and quality in noisy environments while preserving critical speech components. We propose a GAN based novel residual learning architecture with long range temporal context using complex spectrograms. This is achieved by combining residual neural networks with dilated temporal processing. The model processes both magnitude and phase components in the time-frequency domain, implementing residual learning for stable training and a dilated temporal module to capture long-range dependencies in speech signals. Experiments were conducted on benchmark of VoiceBank-DEMAND dataset under various noise conditions. During experimentation, the proposed model demonstrated superior performance compared to existing ones, achieving 3.43 dB improvement in perceptual evaluation of speech quality measurement. The novelties include a phase-aware complex masking system which jointly optimizes magnitude and phase for more natural-sounding output and incorporation of a Dilated temporal module. This module further expands receptive fields exponentially in order to model speech dynamics without increasing parameters drastically. Residual architecture also balances deep feature extraction, computational efficiency and the composite loss function that uniquely combine in delivering superior results. Analysis of our results indicates that this multi-objective optimization framework achieves superior speech enhancement by simultaneously addressing magnitude and phase reconstruction, waveform-level distortion, and psychoacoustic naturalness, a balance not achieved in any prior works so far.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Debabrata Gogoi, Sushanta Kabir Dutta
Quelle
Engineering Research Express
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2631-8695
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Debabrata Gogoi, Sushanta Kabir Dutta (2026). A deep residual complex learning framework with long-range temporal context for phase-aware speech enhancement. Engineering Research Express. https://doi.org/10.1088/2631-8695/ae9cf1
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1 · Lizenz 2