Children s Speech Recognition and Speaker Characterization through Raw Speech Driven Deep Learning Models

Abstract

Automatic speech recognition (ASR) in children is a rapidly evolving field, as children become more accustomed to interacting with virtual assistants, such as Amazon Echo, Cortana, and other smart speakers, and it has advanced human-computer interaction in recent generations. The intricate vocal patterns, intonations, and linguistic nuances present in children s speech hold substantial potential for diverse practical applications. These applications span across domains like interactive voice response systems, personalized service delivery, seamless human-machine interaction, speaker pathology assessment, and the intricate processes of forensic investigation. Recognizing the growing need for harnessing children s speech for these purposes, this thesis delves into exploring accurate methods through raw waveform-driven deep learning models to unlock valuable insights and enhance the effectiveness of speech analysis in children. newlineThis research undertakes a rigorous exploration of the distinct challenges associated with processing children s speech, acknowledging the inherent complexities aris- newlineing from their evolving speech characteristics. To confront these challenges, the study advocates for the adoption of raw waveform modeling a transformative approach that newlineallows direct analysis of speech signals without preliminary feature extraction. This newlineapproach eliminates the limitations of traditional methods that tend to discard subtle details and might induce information loss. To make raw waveform modeling practical,this research work taps into the capabilities of advanced deep learning architectures,notably convolutional neural networks (CNNs) and its variant, SincNet. newlineThe research initially starts with an in-depth exploration of gender identification newlineand speaker recognition in children, employing a diverse array of methodologies. In the realm of gender identification, the study extensively examines comprehensive fea- newlineture engineering techniques through fusion and ablation experiments, including mel newlinefrequency c

Description

Keywords

Citation

item.page.endorsement

item.page.review

item.page.supplemented

item.page.referenced