Multilevel Prosodic Features for Automatic Emotion and Speaker Recognition from Speech using Deep Learning Techniques

Abstract

Speech is the primary mode of communication for human beings. Apart from the intended newlinemessage, the speech signal contains several implicit characteristics, including that of newlinethe speaker and the conveyed emotion.Speech processing tries to extract the desired newlineinformation from the speech signal to facilitate better human-machine interactions. newlineHuman speech signal carries many attributes specific to the underlying emotion or speaker newlineidentity,occurring at multiple levels. However,it is difficult to isolate the attributes related to a speaker or emotion. The proper choice of these attributes would improve recognition newlineaccuracy.Prosodic characteristics are important in human speech communication and newlinelend naturalness and intelligibility to speech. Prosodic features such as duration, energy and fundamental frequency (F0) vary among speakers/emotions. This thesis is an attempt to identify, extract and characterize prosodic features at multiple levels of speech signal for emotion and speaker recognition. newlineAnother goal of this thesis is to develop an efficient classification scheme for making newlineprediction or classification decisions. Recently, the focus of research in various speech processing areas has changed to classifiers incorporating deep learning. This thesis, newlinetherefore, also examines the applicability of various deep learning methods to model the newlineemotional/speaker-specific information present in the speech signal. newlineThe first part of this work concentrates on exploring emotion-specific features at newlinethree different levels of the speech signal, namely, utterance, syllable and frame levels, newlinefor automatic emotion recognition (AER). At the syllable and utterance levels, prosodic newlinecharacteristics using features derived from duration, energy and F0 contour are used to newlinerepresent emotion-specific variations. At the frame level, spectral features represented newlineby Mel frequency cepstral coefficients (MFCCs) are used. The effectiveness of these newlinemultilevel features for emotion recognition are experimentally evaluated

Description

Keywords

Citation

item.page.endorsement

item.page.review

item.page.supplemented

item.page.referenced