Mon-2-8-5 Dual Stage Learning based Dynamic Time-Frequency Mask Generation for Audio Event Classification

Donghyeon Kim(Korea university), Jaihyun Park(Korea University), David Han(US Army Research Laboratory) and Hanseok Ko(Korea University)

Abstract: Audio based event recognition becomes quite challenging in real world noisy environments. To alleviate the noise issue, time-frequency mask based feature enhancement methods have been proposed. While these methods with fixed filter settings have been shown to be effective in familiar noise backgrounds, they become brittle when exposed to unexpected noise. To address the unknown noise problem, we develop an approach based on dynamic filter generation learning. In particular, we propose a dual stage dynamic filter generator networks that can be trained to generate a time-frequency mask specifically created for each input audio. Two alternative approaches of training the mask generator network are developed for feature enhancements in high noise environments. Our proposed method shows improved performance and robustness in both clean and unseen noise environments.

Paper

prev Mon-2-8-4 Memory Controlled Sequential Self Attention for Sound Recognition

next Mon-2-8-6 An Effective Perturbation based Semi-Supervised Learning Method for Sound Event Detection

About

About the Conference

Welcome from the Chair

Conference Committees

Calls