End-to-End Language Identification Using High-Order Utterance Representation with Bilinear Pooling

Jin, Ma, Song, Yan, McLoughlin, Ian Vince, Guo, Wu, Dai, Li-Rong (2017) End-to-End Language Identification Using High-Order Utterance Representation with Bilinear Pooling. In: The proceedings of Interspeech 2017. . pp. 2571-2575. International Speech Communication Society (doi:10.21437/Interspeech.2017-44) (KAR id:61814)

PDF Author's Accepted Manuscript Language: English
Download this file (PDF/896kB)
Request a format suitable for use with assistive technology e.g. a screenreader
Official URL: http://dx.doi.org/10.21437/Interspeech.2017-44

Abstract

A key problem in spoken language identification (LID) is how to design effective representations which are specific to language information. Recent advances in deep neural networks have led to significant improvements in results, with deep end-to-end methods proving effective. This paper proposes a novel network which aims to model an effective representation for high (first and second)-order statistics of LID-senones, defined as being LID analogues of senones in speech recognition. The high-order information extracted through bilinear pooling is robust to speakers, channels and background noise.

Evaluation with NIST LRE 2009 shows improved performance compared to current state-of-the-art DBF/i-vector systems, achieving over 33% and 20% relative equal error rate (EER) improvement for 3s and 10s utterances and over 40% relative Cavg improvement for all durations.

Item Type:	Conference or workshop item (Paper)
DOI/Identification number:	10.21437/Interspeech.2017-44
Subjects:	T Technology > T Technology (General)
Institutional Unit:	Schools > School of Computing
Former Institutional Unit:	Data Science Divisions > Division of Computing, Engineering and Mathematical Sciences > School of Computing
Depositing User:	Ian McLoughlin
Date Deposited:	23 May 2017 08:32 UTC
Last Modified:	20 May 2025 10:20 UTC
Resource URI:	https://kar.kent.ac.uk/id/eprint/61814 (The current URI for this page, for reference purposes)

University of Kent Author Information

McLoughlin, Ian Vince.

Creator's ORCID:	https://orcid.org/0000-0001-7111-2008
CReDIT Contributor Roles:

Depositors only (login required):

Altmetric

Total Views

Total unique views of this page since July 2020. For more details click on the image.