Multilevel Prosodic Features for Automatic Emotion and Speaker Recognition from Speech using Deep Learning Techniques
Loading...
Date
item.page.authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Speech is the primary mode of communication for human beings. Apart from the intended
newlinemessage, the speech signal contains several implicit characteristics, including that of
newlinethe speaker and the conveyed emotion.Speech processing tries to extract the desired
newlineinformation from the speech signal to facilitate better human-machine interactions.
newlineHuman speech signal carries many attributes specific to the underlying emotion or speaker
newlineidentity,occurring at multiple levels. However,it is difficult to isolate the attributes related to a speaker or emotion. The proper choice of these attributes would improve recognition
newlineaccuracy.Prosodic characteristics are important in human speech communication and
newlinelend naturalness and intelligibility to speech. Prosodic features such as duration, energy and fundamental frequency (F0) vary among speakers/emotions. This thesis is an attempt to identify, extract and characterize prosodic features at multiple levels of speech signal for emotion and speaker recognition.
newlineAnother goal of this thesis is to develop an efficient classification scheme for making
newlineprediction or classification decisions. Recently, the focus of research in various speech processing areas has changed to classifiers incorporating deep learning. This thesis,
newlinetherefore, also examines the applicability of various deep learning methods to model the
newlineemotional/speaker-specific information present in the speech signal.
newlineThe first part of this work concentrates on exploring emotion-specific features at
newlinethree different levels of the speech signal, namely, utterance, syllable and frame levels,
newlinefor automatic emotion recognition (AER). At the syllable and utterance levels, prosodic
newlinecharacteristics using features derived from duration, energy and F0 contour are used to
newlinerepresent emotion-specific variations. At the frame level, spectral features represented
newlineby Mel frequency cepstral coefficients (MFCCs) are used. The effectiveness of these
newlinemultilevel features for emotion recognition are experimentally evaluated