Design and Development of Enhanced Multi modal Deep Learning Frameworks for Visual Question Answering VQA

dc.contributor.guideKovoor, Binsu C
dc.coverage.spatial
dc.creator.researcherManmadhan, Sruthy
dc.date.accessioned2023-09-04T08:18:04Z
dc.date.available2023-09-04T08:18:04Z
dc.date.awarded2023
dc.date.completed2022
dc.date.registered2017
dc.description.abstractThis thesis studies a multi-modal AI task called Visual Question Answering newline(VQA). It covers two different areas of computer science research; Computer Vision newline(CV) and Natural Language Processing (NLP). Due to its expansive set of newlineapplications including assistance to visually impaired people, surveillance data newlineanalysis etc., many researchers attracted to this AI-complete task for the last few newlineyears. Most of the existing works have given attention to the multi-modal feature newlinefusion phase of VQA ignoring the effect of individual input features. Thus, despite newlinerapid improvements in VQA algorithm efficiency, there is still a substantial gap newlinebetween the best methods and humans. The proposed research aims to design and newlinedevelop deep learning models for the AI-complete task of Visual Question newlineAnswering with enhanced multi-modal representations and thereby reducing the gap newlinebetween human and machine intelligence. The proposed research focus on each task in the established three phase newlinepipeline of VQA; image and question feature extractions, the multi-modal newlineembedding of visual and textual features and answer generation. The methodologies newlineused to tackle image featurization include a ranking and feature fusion framework to newlinefuse feature vectors from pre-trained CNN image feature extractors for a dataset, and newlinea dedicated Convolutional Denoising Auto-encoder (CDAE) design for extracting newlineimage features from domain-specific VQA images. newline
dc.description.note
dc.format.accompanyingmaterialDVD
dc.format.dimensions
dc.format.extentxvi,240
dc.identifier.urihttp://hdl.handle.net/10603/510334
dc.languageEnglish
dc.publisher.institutionDepartment of Information Technology
dc.publisher.placeCochin
dc.publisher.universityCochin University of Science and Technology
dc.relation
dc.rightsuniversity
dc.source.universityUniversity
dc.subject.keywordComputer Vision
dc.subject.keywordEngineering and Technology
dc.subject.keywordInformation Technology
dc.subject.keywordNatural Language Processing (NLP)
dc.subject.keywordVisual Question Answering
dc.titleDesign and Development of Enhanced Multi modal Deep Learning Frameworks for Visual Question Answering VQA
dc.title.alternative
dc.type.degreePh.D.

Files

Original bundle

Now showing 1 - 5 of 14
Loading...
Thumbnail Image
Name:
01_title.pdf
Size:
194.87 KB
Format:
Adobe Portable Document Format
Description:
Attached File
Loading...
Thumbnail Image
Name:
02 -preliminary pages.pdf
Size:
836.83 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
03_content.pdf
Size:
190.5 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
04_abstract.pdf
Size:
178.42 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
05_chapter1.pdf
Size:
908.35 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.79 KB
Format:
Plain Text
Description: