Abstract

Multi-channel speech enhancement based on deep learning has achieved significant progress. While keeping low complexity, lightweight dual-channel models exploit spatial cues to outperform single-channel models on edge devices. However, their performance often degrades for closely-spaced sources due to unconstrained spatial reliance. To mitigate this, we propose a Dynamic Spatial-aware Knowledge Distillation (D-SKD) framework that enhances spatial discrimination during training. Specifically, D-SKD introduces a dynamic arbitration into the distillation process with a spatially invariant single-channel teacher, enabling the dual-channel model to adaptively regulate spatial reliance without explicit angular supervision. Experiments on dual-channel GTCRN and other models show that D-SKD consistently improves performance for closely-spaced sources while preserving quality elsewhere. Requiring no additional inference overhead or auxiliary inputs, it is well suited for lightweight deployment.

This page is for research demonstration purposes only.

Framework Overview

Overall architecture of the D-SKD framework
Figure 1. Overall architecture of D-SKD. (a) Training stage with teacher-student distillation via DAM. (b) Inference stage with only the student branch.

Audio Demo

PESQ, eSTOI, and SI-SDR are computed against the clean speech in each sample. SC-GTCRN is the single-channel teacher, DC-GTCRN is the dual-channel model trained without D-SKD, and Proposed is the DC-GTCRN student trained with D-SKD.

Sample 1

Method Audio PESQ eSTOI SI-SDR (dB)
Clean - - -
Noisy 1.5484 0.8263 9.7669
SC-GTCRN 2.1589 0.8577 14.9212
DC-GTCRN 1.7217 0.7911 11.9548
Proposed 2.0274 0.8503 14.1026
Sample 1 clean spectrogram
Clean
Sample 1 noisy spectrogram
Noisy
Sample 1 SC-GTCRN spectrogram
SC-GTCRN
Sample 1 DC-GTCRN spectrogram
DC-GTCRN
Sample 1 proposed spectrogram
Proposed

Sample 2

Method Audio PESQ eSTOI SI-SDR (dB)
Clean - - -
Noisy 1.3059 0.8355 7.1870
SC-GTCRN 2.1691 0.8793 13.5816
DC-GTCRN 1.5797 0.8145 7.5665
Proposed 2.1732 0.8728 10.6656
Sample 2 clean spectrogram
Clean
Sample 2 noisy spectrogram
Noisy
Sample 2 SC-GTCRN spectrogram
SC-GTCRN
Sample 2 DC-GTCRN spectrogram
DC-GTCRN
Sample 2 proposed spectrogram
Proposed