Development of efficient techniques for enhancing speaker recognition from forensic speech with multiple distortions

dc.contributor.guideSathidevi, P S
dc.coverage.spatial
dc.creator.researcherM S, Athulya
dc.date.accessioned2022-12-17T02:02:13Z
dc.date.available2022-12-17T02:02:13Z
dc.date.awarded2022
dc.date.completed2022
dc.date.registered2015
dc.description.abstractForensic speaker recognition (FSR) or forensic automatic speaker recognition (FASR) newlinerefers to the process of automatically recognizing the speaker of a speech utterance newlineobtained from a crime scenario with the assistance of machines. Such speech newlineevidence may be unpredictably distorted due to noise, codec distortions, channel newlineeffects, multiple speakers, voice disguise etc. The different stages in a typical forensic newlinespeaker recognition system are pre-processing, feature extraction, modelling and newlinescore matching/interpretation. Robust methods at these different stages of FASR must newlinebe developed to reduce the recognition error. This research work mainly focusses on newlinethe distorted speech collected from mobile phone based communication. newlineIn mobile communications, noise can interfere with speech. Speech enhancement newlineusing adaptive filtering methods is known to provide good signal recovery, but these newlinealgorithms have a constraint that correlating noise should be given as the reference newlinesignal for denoising. A novel method for identifying the best correlating part of newlinethe noise signal with respect to noise in noisy speech is proposed to address the newlineaforementioned constraint. This is used as a reference for speech enhancement in newlinevariable step size least mean square (VSSLMS) and recursive least squares (RLS) newlinealgorithms for speech enhancement. Prior to this, noise classification is done. The newlineproposed system performs very well even when speech is mixed with noise under newlinevery low SNR conditions, as commonly encountered in forensic applications. newlineSpeech codecs used in mobile phone communication may remove or distort some newlinespeaker-specific features, reducing speaker verification accuracy. To improve the newlineverification accuracy, power normalized cepstral coefficient (PNCC) features are newlineslightly modified in this work. A series combination of Gaussian Mixture Model- newlineUniversal Background Model (GMM-UBM) and SVM classifiers is also proposed here to enhance the speaker verification accuracy even further. newline
dc.description.note
dc.format.accompanyingmaterialDVD
dc.format.dimensions
dc.format.extent
dc.identifier.urihttp://hdl.handle.net/10603/425941
dc.languageEnglish
dc.publisher.institutionDepartment of Electronics and Communication Engineering
dc.publisher.placeCalicut
dc.publisher.universityNational Institute of Technology Calicut
dc.relation
dc.rightsuniversity
dc.source.universityUniversity
dc.subject.keywordEngineering and Technology
dc.subject.keywordEngineering
dc.subject.keywordEngineering Electrical and Electronic
dc.subject.keywordForensic automatic speaker recognition
dc.titleDevelopment of efficient techniques for enhancing speaker recognition from forensic speech with multiple distortions
dc.title.alternative
dc.type.degreePh.D.

Files

Original bundle

Now showing 1 - 5 of 11
Loading...
Thumbnail Image
Name:
01_title.pdf
Size:
61.14 KB
Format:
Adobe Portable Document Format
Description:
Attached File
Loading...
Thumbnail Image
Name:
02_prelim pages.pdf
Size:
1.05 MB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
03_content.pdf
Size:
38.02 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
04_abstract.pdf
Size:
33.63 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
05_chapter 1.pdf
Size:
256.36 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.79 KB
Format:
Plain Text
Description: