Noise Modeling to Build Training Sets for Robust Speech Enhancement

preprint OA: closed CC-BY-4.0
📄 Open PDF View at publisher

Abstract

The performance of Deep Neural Network (DNN)-based speech enhancement models degrades significantly in real recordings because the synthetic training sets are mismatched with real test sets. To solve this problem, we propose a new Generative Adversarial Network framework for Noise Modeling (NM-GAN) that can build training sets by imitating real noise distribution. The framework combines a novel U-Net with two bidirectional Long Short-Term Memory (LSTM) layers that act as a generator to construct complex noise. The Gaussian distribution is adapted and used as conditional information to direct the noise generation. A discriminator then learns to determine whether a noise sample is from the model distribution or from a real noise distribution. By adversarial and alternate training, NM-GAN can generate enough recall (diversity) and precision (quality of noise) in its samples for it to look like real noise. Afterwards, realistic-looking paired training sets are composed. Extensive experiments were carried out and qualitative and quantitative evaluation of the generated noise samples and training sets demonstrate that potential of the framework. An Speech enhancement model trained on our synthetic training sets and on real training sets was found to be capable of good noise suppression for real speech-related noise.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0