An Aggrandized Framework for Improving Large Vocabulary Continuous Speech Recognition LVCSR of Lecture Speech in Indian English
Loading...
Date
item.page.authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Automatic Speech Recognition (ASR) is concerned about converting spoken utterances in audio signal into text. It has a wide variety of applications like voice user interfaces which support voice dialing, domotic appliance control and voice search. During the past few decades, drastic developments have been reported in ASR for many languages such as English, Finnish, German, etc. However, development of Indian English (IE) speech recognition models seems to be quite untended. IE is one of the varieties of English spoken in Indian subcontinent showing idiosyncrasy in terms of pronunciation, vocabulary, dialect and accent from English spoken in other parts of world. Also, there is dearth of resources in terms of transcribed speech data and spoken language corpora in IE. In spite of these challenges, an aggrandized framework has been proposed to improve
newlinethe speech recognition accuracy in this thesis. In this work, (i) an Indian English
newlineacoustic model has been developed that gives 35% lessWord Error Rate (WER) in comparison to existing English acoustic models such as HUB4; (ii) the acoustic-phonetic analysis of vowels and consonants has been carried out to understand the characteristic differences within Indian English varieties; (iii) an effective methodology has been proposed by using Wikipedia dump corpus along with Google search to interpolate and adapt the language models closer to the topic of the spoken lecture which has further reduced the Word Error Rate to 14% with a major decrease in the perplexity of language model; (iv) an effective retrieval method is proposed using Elasticsearch framework to retrieve documents using the key phrases identified from the ASR output to create domain-specific language model that reduces the search time by more than 90% in comparison to conventional search and retrieval mechanism. Finally, from the reported results, one can conclude that the framework proposed drastically reduces the perplexity as well as WER and improves the performance of speech recognition of Indian Englis