Development of efficient techniques for enhancing speaker recognition from forensic speech with multiple distortions

Loading...
Thumbnail Image

Date

item.page.authors

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Forensic speaker recognition (FSR) or forensic automatic speaker recognition (FASR) newlinerefers to the process of automatically recognizing the speaker of a speech utterance newlineobtained from a crime scenario with the assistance of machines. Such speech newlineevidence may be unpredictably distorted due to noise, codec distortions, channel newlineeffects, multiple speakers, voice disguise etc. The different stages in a typical forensic newlinespeaker recognition system are pre-processing, feature extraction, modelling and newlinescore matching/interpretation. Robust methods at these different stages of FASR must newlinebe developed to reduce the recognition error. This research work mainly focusses on newlinethe distorted speech collected from mobile phone based communication. newlineIn mobile communications, noise can interfere with speech. Speech enhancement newlineusing adaptive filtering methods is known to provide good signal recovery, but these newlinealgorithms have a constraint that correlating noise should be given as the reference newlinesignal for denoising. A novel method for identifying the best correlating part of newlinethe noise signal with respect to noise in noisy speech is proposed to address the newlineaforementioned constraint. This is used as a reference for speech enhancement in newlinevariable step size least mean square (VSSLMS) and recursive least squares (RLS) newlinealgorithms for speech enhancement. Prior to this, noise classification is done. The newlineproposed system performs very well even when speech is mixed with noise under newlinevery low SNR conditions, as commonly encountered in forensic applications. newlineSpeech codecs used in mobile phone communication may remove or distort some newlinespeaker-specific features, reducing speaker verification accuracy. To improve the newlineverification accuracy, power normalized cepstral coefficient (PNCC) features are newlineslightly modified in this work. A series combination of Gaussian Mixture Model- newlineUniversal Background Model (GMM-UBM) and SVM classifiers is also proposed here to enhance the speaker verification accuracy even further. newline

Description

Keywords

Citation

item.page.endorsement

item.page.review

item.page.supplemented

item.page.referenced