University of Surrey

Test tubes in the lab Research in the ATI Dance Research

Large-scale weakly supervised audio classification using gated convolutional neural network

Xu, Yong, Kong, Qiuqiang, Wang, Wenwu and Plumbley, Mark (2018) Large-scale weakly supervised audio classification using gated convolutional neural network In: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 15 - 20 April 2018, Calgary, Alberta, Canada.

icassp2018_dcase2017_final_paper.pdf - Accepted version Manuscript

Download (327kB) | Preview


In this paper, we present a gated convolutional neural network and a temporal attention-based localization method for audio classification, which won the 1st place in the large-scale weakly supervised sound event detection task of Detection and Classification of Acoustic Scenes and Events (DCASE) 2017 challenge. The audio clips in this task, which are extracted from YouTube videos, are manually labelled with one or more audio tags, but without time stamps of the audio events, hence referred to as weakly labelled data. Two subtasks are defined in this challenge including audio tagging and sound event detection using this weakly labelled data. We propose a convolutional recurrent neural network (CRNN) with learnable gated linear units (GLUs) non-linearity applied on the log Mel spectrogram. In addition, we propose a temporal attention method along the frames to predict the locations of each audio event in a chunk from the weakly labelled data. The performances of our systems were ranked the 1st and the 2nd as a team in these two sub-tasks of DCASE 2017 challenge with F value 55.6% and Equal error 0.73, respectively.

Item Type: Conference or Workshop Item (Conference Paper)
Divisions : Faculty of Engineering and Physical Sciences > Electronic Engineering
Authors :
Date : 13 September 2018
DOI : 10.1109/ICASSP.2018.8461975
Copyright Disclaimer : © 2018 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Uncontrolled Keywords : DCASE2017 challenge, weakly supervised sound event detection, audio tagging, attention, gated linear unit
Related URLs :
Depositing User : Melanie Hughes
Date Deposited : 06 Feb 2018 10:16
Last Modified : 10 Dec 2018 13:12

Actions (login required)

View Item View Item


Downloads per month over past year

Information about this web site

© The University of Surrey, Guildford, Surrey, GU2 7XH, United Kingdom.
+44 (0)1483 300800