EUVIMEDEuropean Health Evidence
Uhr 7/7Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

TF-Mamba-DPHNet: a time-frequency state-space and dual-path interaction framework with hybrid cross-scale feature calibration for speech enhancement

Shaik AreefaBegam, Sunnydayal Vanambathina

Frontiers in Signal Processing · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Introduction Speech enhancement plays a fundamental role in improving the perceptual quality and intelligibility of speech signals degraded by noise. Conventional U-Net architectures effectively capture local spectral information but struggle to model long-range dependencies and may allow residual noise to propagate through skip connections. Transformer-based networks can model global contextual dependencies and generate high-quality enhanced speech; however, their limited capacity to preserve fine-grained high-frequency spectral details can limit their suitability for real-time applications. Methods To address these limitations, this paper proposes TF-Mamba-DPHNet, a novel encoder-decoder speech enhancement framework integrating Multi-Scale Feature Extraction (MSFE), Time-Frequency Mamba (TF-Mamba), Dual-Path Higher-Order Information Interaction with Time-Frequency Attention (DPH-TFA), and Hybrid Cross-Scale Feature Calibration (H-CS-FC) modules. The MSFE blocks in the encoder and decoder extract local patterns across multiple receptive fields, capturing fine-grained and global time-frequency cues. The TF-Mamba module models global time-frequency dependencies using selective state-space modeling for effective long-term sequence understanding. At the bottleneck, four stacked DPH-TFA blocks capture long-range dependencies in both the temporal and spectral domains. These global features are fused with hierarchical encoder outputs through H-CS-FC modules, which perform cross-scale-guided feature recalibration to suppress noise leakage in the skip pathways and improve decoder reliability. Results Evaluations conducted on the Common Voice and LibriSpeech datasets demonstrate that TF-Mamba-DPHNet effectively improves speech quality and robustness compared with earlier models. Discussion The results indicate that combining multi-scale local feature extraction, selective state-space modeling, dual-path time-frequency interaction, and cross-scale feature calibration enables effective modeling of complementary local and global information for robust speech enhancement.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Shaik AreefaBegam, Sunnydayal Vanambathina
Quelle
Frontiers in Signal Processing
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2673-8198
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Shaik AreefaBegam, Sunnydayal Vanambathina (2026). TF-Mamba-DPHNet: a time-frequency state-space and dual-path interaction framework with hybrid cross-scale feature calibration for speech enhancement. Frontiers in Signal Processing. https://doi.org/10.3389/frsip.2026.1911510
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1