EUVIMEDEuropean Health Evidence
Uhr 10/10Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

DMAWNet: Multi-Objective Based Distributed Attention Enabled WNet Framework for Speech Enhancement

Anil Garg, Ankur Singhal, Deepti Chaudhary, Priyanka Jangra, Poonam Rani, Ajay Jangra

Fluctuation and Noise Letters · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Speech enhancement is significant with the advancement of communication systems, and in real-world circumstances, ambient noise is often present in speech data, emphasizing the vital role of enhancement techniques. The existing methods employed for enhancement are susceptible to various drawbacks, including the lack of interpretation of noise characteristics, failure to extract spectral details in different environments, and the required high time consumption, which leads to high cost. Moreover, the models either focus on noise suppression or speech restoration. However, the joint handling is often limited due to the complexities in tonal variations. Therefore, the research proposes a distributed attention-enabled multi-objective autoencoder-based WNet (DMAWNet) framework, which aims to address the limitations in existing methods as well as exhibit better performance. The application of the Complex Rectangular Bandwidth features (CReBF) extraction method provides better extraction of the subtle transitions and nuances of the complex channels, irrespective of background noise. Additionally, the incorporation of the Distributed attention mechanism (DisAT) layer in the WNet model improves the capability of the model to concentrate on the spatial details along the channel axis, thereby allowing it to process significant signal elements. The incorporation of all these components provides enhanced speech signals, which are further transformed into the time domain using Inverse Short-time Fourier Transform (ISTFT). By evaluation results, the DMAWNet achieves a Hearing-Aid Speech Perception Index of 0.80, Hearing Aid Speech Quality Index of 0.90, Log Spectral Distance of 2.15, 4.49 Perceptual Evaluation of Speech Quality, Relative Root Mean Square Error of 0.10, Signal-to-Distortion Ratio of 19.49, and Short-Time Objective Intelligibility score of 0.90 in the LibriSpeech dataset.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Anil Garg, Ankur Singhal, Deepti Chaudhary, Priyanka Jangra, Poonam Rani, Ajay Jangra
Quelle
Fluctuation and Noise Letters
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
0219-4775, 1793-6780
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Anil Garg, Ankur Singhal, Deepti Chaudhary, Priyanka Jangra, Poonam Rani, Ajay Jangra (2026). DMAWNet: Multi-Objective Based Distributed Attention Enabled WNet Framework for Speech Enhancement. Fluctuation and Noise Letters. https://doi.org/10.1142/s0219477526500549
RIS BibTeX CSL-JSON