Design and Development of Enhanced Multi modal Deep Learning Frameworks for Visual Question Answering VQA

Abstract

This thesis studies a multi-modal AI task called Visual Question Answering newline(VQA). It covers two different areas of computer science research; Computer Vision newline(CV) and Natural Language Processing (NLP). Due to its expansive set of newlineapplications including assistance to visually impaired people, surveillance data newlineanalysis etc., many researchers attracted to this AI-complete task for the last few newlineyears. Most of the existing works have given attention to the multi-modal feature newlinefusion phase of VQA ignoring the effect of individual input features. Thus, despite newlinerapid improvements in VQA algorithm efficiency, there is still a substantial gap newlinebetween the best methods and humans. The proposed research aims to design and newlinedevelop deep learning models for the AI-complete task of Visual Question newlineAnswering with enhanced multi-modal representations and thereby reducing the gap newlinebetween human and machine intelligence. The proposed research focus on each task in the established three phase newlinepipeline of VQA; image and question feature extractions, the multi-modal newlineembedding of visual and textual features and answer generation. The methodologies newlineused to tackle image featurization include a ranking and feature fusion framework to newlinefuse feature vectors from pre-trained CNN image feature extractors for a dataset, and newlinea dedicated Convolutional Denoising Auto-encoder (CDAE) design for extracting newlineimage features from domain-specific VQA images. newline

Description

Keywords

Citation

item.page.endorsement

item.page.review

item.page.supplemented

item.page.referenced