Development of efficient techniques for enhancing speaker recognition from forensic speech with multiple distortions
Loading...
Date
item.page.authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Forensic speaker recognition (FSR) or forensic automatic speaker recognition (FASR)
newlinerefers to the process of automatically recognizing the speaker of a speech utterance
newlineobtained from a crime scenario with the assistance of machines. Such speech
newlineevidence may be unpredictably distorted due to noise, codec distortions, channel
newlineeffects, multiple speakers, voice disguise etc. The different stages in a typical forensic
newlinespeaker recognition system are pre-processing, feature extraction, modelling and
newlinescore matching/interpretation. Robust methods at these different stages of FASR must
newlinebe developed to reduce the recognition error. This research work mainly focusses on
newlinethe distorted speech collected from mobile phone based communication.
newlineIn mobile communications, noise can interfere with speech. Speech enhancement
newlineusing adaptive filtering methods is known to provide good signal recovery, but these
newlinealgorithms have a constraint that correlating noise should be given as the reference
newlinesignal for denoising. A novel method for identifying the best correlating part of
newlinethe noise signal with respect to noise in noisy speech is proposed to address the
newlineaforementioned constraint. This is used as a reference for speech enhancement in
newlinevariable step size least mean square (VSSLMS) and recursive least squares (RLS)
newlinealgorithms for speech enhancement. Prior to this, noise classification is done. The
newlineproposed system performs very well even when speech is mixed with noise under
newlinevery low SNR conditions, as commonly encountered in forensic applications.
newlineSpeech codecs used in mobile phone communication may remove or distort some
newlinespeaker-specific features, reducing speaker verification accuracy. To improve the
newlineverification accuracy, power normalized cepstral coefficient (PNCC) features are
newlineslightly modified in this work. A series combination of Gaussian Mixture Model-
newlineUniversal Background Model (GMM-UBM) and SVM classifiers is also proposed here to enhance the speaker verification accuracy even further.
newline