Conferences >ICASSP 2023 - 2023 IEEE Inter...

Improving Spoken Language Identification with Map-Mix

Download PDF
Download References
Request Permissions
Save to
Alerts

Abstract:

The pre-trained multi-lingual XLSR model generalizes well for language identification after fine-tuning on unseen languages. However, the performance significantly degrad...Show More

Metadata

Abstract:

The pre-trained multi-lingual XLSR model generalizes well for language identification after fine-tuning on unseen languages. However, the performance significantly degrades when the languages are not very distinct from each other, for example, in the case of dialects. Low resource dialect classification remains a challenging problem to solve. We present a new data augmentation method that leverages model training dynamics of individual data points to improve sampling for the latent mixup. The method works well in low-resource settings where generalization is paramount. Our datamaps-based mixup technique, which we call Map-Mix, improves weighted F1 scores by 2% compared to the random mixup baseline and results in a significantly well-calibrated model. The code for our method is open-sourced on github.

Published in: ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

Date of Conference: 04-10 June 2023

Date Added to IEEE Xplore: 05 May 2023

ISBN Information:

ISSN Information:

DOI: 10.1109/ICASSP49357.2023.10095765

Conference Location: Rhodes Island, Greece

Contents

1. INTRODUCTION

Spoken Language Identification(SLID) is the problem of classifying the language spoken by a speaker in an audio clip. SLID is useful in personalized voice assistants, automatic speech translation systems, and multi-lingual speech recognition systems and has been used in call centers to route calls to a specific language operator automatically. Earlier studies [1] [2] have used the phonetic, phonotactic, prosodic, and lexical features for SLID. Classical SLID models first extract the i-vectors [3], or x-vectors [4] and then train an independent classifier model on top. Acoustic features such as MFCCs and filter banks are also commonly used as input features [5].

References is not available for this document.

Improving Spoken Language Identification with Map-Mix

Abstract:

Metadata

Abstract:

ISSN Information:

1. INTRODUCTION

References

IEEE Account

Purchase Details

Profile Information

Need Help?

Improving Spoken Language Identification with Map-Mix

Alerts

Abstract:

Metadata

Abstract:

ISSN Information:

1. INTRODUCTION

References