Children s Speech Recognition and Speaker Characterization through Raw Speech Driven Deep Learning Models

dc.contributor.guideBansal, Mohan
dc.coverage.spatial
dc.creator.researcherRadha, Kodali
dc.date.accessioned2024-02-08T06:03:47Z
dc.date.available2024-02-08T06:03:47Z
dc.date.awarded2024
dc.date.completed2024
dc.date.registered2021
dc.description.abstractAutomatic speech recognition (ASR) in children is a rapidly evolving field, as children become more accustomed to interacting with virtual assistants, such as Amazon Echo, Cortana, and other smart speakers, and it has advanced human-computer interaction in recent generations. The intricate vocal patterns, intonations, and linguistic nuances present in children s speech hold substantial potential for diverse practical applications. These applications span across domains like interactive voice response systems, personalized service delivery, seamless human-machine interaction, speaker pathology assessment, and the intricate processes of forensic investigation. Recognizing the growing need for harnessing children s speech for these purposes, this thesis delves into exploring accurate methods through raw waveform-driven deep learning models to unlock valuable insights and enhance the effectiveness of speech analysis in children. newlineThis research undertakes a rigorous exploration of the distinct challenges associated with processing children s speech, acknowledging the inherent complexities aris- newlineing from their evolving speech characteristics. To confront these challenges, the study advocates for the adoption of raw waveform modeling a transformative approach that newlineallows direct analysis of speech signals without preliminary feature extraction. This newlineapproach eliminates the limitations of traditional methods that tend to discard subtle details and might induce information loss. To make raw waveform modeling practical,this research work taps into the capabilities of advanced deep learning architectures,notably convolutional neural networks (CNNs) and its variant, SincNet. newlineThe research initially starts with an in-depth exploration of gender identification newlineand speaker recognition in children, employing a diverse array of methodologies. In the realm of gender identification, the study extensively examines comprehensive fea- newlineture engineering techniques through fusion and ablation experiments, including mel newlinefrequency c
dc.description.note
dc.format.accompanyingmaterialDVD
dc.format.dimensions29x19
dc.format.extentxviii,126
dc.identifier.urihttp://hdl.handle.net/10603/544339
dc.languageEnglish
dc.publisher.institutionDepartment of Electronics Engineering
dc.publisher.placeAmaravati
dc.publisher.universityVellore Institute of Technology (VIT-AP)
dc.relation
dc.rightsuniversity
dc.source.universityUniversity
dc.subject.keywordAutomatic speech recognition
dc.subject.keywordChildren speech
dc.subject.keywordGender identification
dc.titleChildren s Speech Recognition and Speaker Characterization through Raw Speech Driven Deep Learning Models
dc.title.alternative
dc.type.degreePh.D.

Files

Original bundle

Now showing 1 - 5 of 11
Loading...
Thumbnail Image
Name:
01_title.pdf
Size:
115.5 KB
Format:
Adobe Portable Document Format
Description:
Attached File
Loading...
Thumbnail Image
Name:
02_ prelim pages.pdf
Size:
299.76 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
03_ contents.pdf
Size:
51.48 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
04_abstract.pdf
Size:
63.13 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
05_ chapter-1.pdf
Size:
461.78 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.79 KB
Format:
Plain Text
Description: