Robust acoustic scene classification using a multi-spectrogram encoder-decoder framework

Pham, Lam Dang, Phan, Huy, Nguyen, Truc, Palaniappan, Ramaswamy, Mertins, Afred, McLoughlin, Ian Vince (2020) Robust acoustic scene classification using a multi-spectrogram encoder-decoder framework. Digital Signal Processing, . Article Number 102943. ISSN 1051-2004. (doi:10.1016/j.dsp.2020.102943) (KAR id:85290)

PDF Author's Accepted Manuscript Language: English This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Download this file (PDF/576kB)	Preview
Request a format suitable for use with assistive technology e.g. a screenreader
Official URL: https://doi.org/10.1016/j.dsp.2020.102943

Abstract

This article proposes an encoder-decoder network model for Acoustic Scene Classification (ASC), the task of identifying the scene of an audio recording from its acoustic signature. We make use of multiple low-level spectrogram features at the front-end, transformed into higher level features through a well-trained CNN-DNN front-end encoder. The high-level features and their combination (via a trained feature combiner) are then fed into different decoder models comprising random forest regression, DNNs and a mixture of experts, for back-end classification. We conduct extensive experiments to evaluate the performance of this framework on various ASC datasets, including LITIS Rouen and IEEE AASP Challenge on Detection and Classification of Acoustic Scenes and Events (DCASE) 2016 Task 1, 2017 Task 1, 2018 Tasks 1A & 1B and 2019 Tasks 1A & 1B. The experimental results highlight two main contributions; the first is an effective method for high-level feature extraction from multi-spectrogram input via the novel CNN-DNN architecture encoder network, and the second is the proposed decoder which enables the framework to achieve competitive results on various datasets. The fact that a single framework is highly competitive for several different challenges is an indicator of its robustness for performing general ASC tasks.

Item Type:	Article
DOI/Identification number:	10.1016/j.dsp.2020.102943
Uncontrolled keywords:	Acoustic scene classification;Encoder-decoder network;Low-level features;High-level features;Multi-spectrogram
Institutional Unit:	Schools > School of Computing
Former Institutional Unit:	Divisions > Division of Computing, Engineering and Mathematical Sciences > School of Computing
Depositing User:	Palaniappan Ramaswamy
Date Deposited:	03 Jan 2021 23:38 UTC
Last Modified:	22 Jul 2025 09:04 UTC
Resource URI:	https://kar.kent.ac.uk/id/eprint/85290 (The current URI for this page, for reference purposes)

University of Kent Author Information

Pham, Lam Dang.

Creator's ORCID:
CReDIT Contributor Roles:

Phan, Huy.

Creator's ORCID:	https://orcid.org/0000-0003-4096-785X
CReDIT Contributor Roles:

Palaniappan, Ramaswamy.

Creator's ORCID:	https://orcid.org/0000-0001-5296-8396
CReDIT Contributor Roles:

McLoughlin, Ian Vince.

Creator's ORCID:	https://orcid.org/0000-0001-7111-2008
CReDIT Contributor Roles:

Depositors only (login required):

Altmetric

Total Views

Total unique views of this page since July 2020. For more details click on the image.