Please use this identifier to cite or link to this item:
|Title:||Variants of mel-frequency cepstral coefficients for improved whispered speech speaker verification in mismatched conditions|
|Authors:||Sarria Paja, Milton|
Falk, Tiago H.
|Keywords:||Whispered speech;Speaker verification;Fusion;I-vector extraction;MFCC|
|Publisher:||Institute of Electrical and Electronics Engineers Inc.|
|Abstract:||In this paper, automatic speaker verification using normal and whispered speech is explored. Typically, for speaker verification systems, varying vocal effort inputs during the testing stage significantly degrades system performance. Solutions such as feature mapping or addition of multi-style data during training and enrollment stages have been proposed but do not show similar advantages for the involved speaking styles. Herein, we focus attention on the extraction of invariant speaker-dependent information from normal and whispered speech, thus allowing for improved multi vocal effort speaker verification. We base our search on previously reported perceptual and acoustic insights and propose variants of the mel-frequency cepstral coefficients (MFCC). We show the complementarity of the proposed features via three fusion schemes. Gains as high as 39% and 43% can be achieved for normal and whispered speech, respectively, relative to the existing systems based on conventional MFCC features.|
|Appears in Collections:||Artículos Científicos|
Files in This Item:
|Variants of mel-frequency cepstral coefficients for improved.jpg||254,09 kB||JPEG||View/Open|
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.