Design and Development of Enhanced Multi modal Deep Learning Frameworks for Visual Question Answering VQA
Loading...
Date
item.page.authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
This thesis studies a multi-modal AI task called Visual Question Answering
newline(VQA). It covers two different areas of computer science research; Computer Vision
newline(CV) and Natural Language Processing (NLP). Due to its expansive set of
newlineapplications including assistance to visually impaired people, surveillance data
newlineanalysis etc., many researchers attracted to this AI-complete task for the last few
newlineyears. Most of the existing works have given attention to the multi-modal feature
newlinefusion phase of VQA ignoring the effect of individual input features. Thus, despite
newlinerapid improvements in VQA algorithm efficiency, there is still a substantial gap
newlinebetween the best methods and humans. The proposed research aims to design and
newlinedevelop deep learning models for the AI-complete task of Visual Question
newlineAnswering with enhanced multi-modal representations and thereby reducing the gap
newlinebetween human and machine intelligence. The proposed research focus on each task in the established three phase
newlinepipeline of VQA; image and question feature extractions, the multi-modal
newlineembedding of visual and textual features and answer generation. The methodologies
newlineused to tackle image featurization include a ranking and feature fusion framework to
newlinefuse feature vectors from pre-trained CNN image feature extractors for a dataset, and
newlinea dedicated Convolutional Denoising Auto-encoder (CDAE) design for extracting
newlineimage features from domain-specific VQA images.
newline