Robust detection of vowels in speech signal with application to children s ASR

dc.contributor.guidePradhan, Gayadhar
dc.coverage.spatial
dc.creator.researcherKumar, Avinash
dc.date.accessioned2022-12-09T08:04:27Z
dc.date.available2022-12-09T08:04:27Z
dc.date.awarded2020
dc.date.completed2019
dc.date.registered2014
dc.description.abstractThis thesis proposes acoustic modeling as well as signal processing approaches for robustly detecting vowels and corresponding onset points newline(VOPs) and offset points (VEPs) in a given speech signal. The VOP newlineand VEP are defined as the instant of starting and ending of a vowel, newlinerespectively. The knowledge of vowel and non-vowel regions is then newlineexploited for non-uniformly suppressing the pitch induced mismatch newlinein children s automatic speech recognition (ASR) system. newlineAt first, using mel-frequency cepstral coefficients (MFCCs) as the newlinefront-end features, three-class classifiers (vowels, non-vowels and silences) are developed using recently reported state-of-the-art acoustic newlinemodeling methods for the task of detecting vowels, VOPs and VEPs newlinein a given speech signal. Among the explored acoustic modeling techniques, best performance is observed for pre-trained deep neural networks (DNN). To further enhance the performance, a novel front-end newlinefeature exploiting the temporal and spectral characteristics of the excitation source information in speech signal is proposed. The use of newlinethe proposed feature results in the detection of vowel regions that are newlinequite different from those obtained through the MFCCs. Exploiting newlinethose differences in the obtained evidences taking two different kinds newlineof features, a technique to combine the evidences is also proposed. The newlinestatistical learning based approaches provides significantly improved newlineperformance when compared with the explicit signal processing methods reported in the literature. However, performance of statistical classifiers degrades significantly when speech signal is corrupted by newlineambient noises. newlineIn order to enhance the robustness towards ambient noises, a signal newlineprocessing approach based on non-local means (NLM) estimation is newlinethen proposed for the detection of vowels, VOPs and VEPs. In the newlineNLM algorithm, the signal value at each sample point is estimated newlineas the weighted sum of signal values at other sample points within a newlinesearch neighborhood.
dc.description.note
dc.format.accompanyingmaterialCD
dc.format.dimensions29cm.
dc.format.extentxxxi, 155p.
dc.identifier.urihttp://hdl.handle.net/10603/423552
dc.languageEnglish
dc.publisher.institutionElectronics and Communications Engineering
dc.publisher.placePatna
dc.publisher.universityNational Institute of Technology Patna
dc.relation
dc.rightsuniversity
dc.source.universityUniversity
dc.subject.keywordEngineering
dc.subject.keywordEngineering and Technology
dc.subject.keywordEngineering Electrical and Electronic
dc.titleRobust detection of vowels in speech signal with application to children s ASR
dc.title.alternative
dc.type.degreePh.D.

Files

Original bundle

Now showing 1 - 5 of 13
Loading...
Thumbnail Image
Name:
01_title.pdf
Size:
102.09 KB
Format:
Adobe Portable Document Format
Description:
Attached File
Loading...
Thumbnail Image
Name:
02_prelim pages.pdf
Size:
121.95 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
03_content.pdf
Size:
242.54 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
04_abstract.pdf
Size:
51.43 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
05_chapter 1.pdf
Size:
126.85 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.79 KB
Format:
Plain Text
Description: