Development of Interpretable and Efficient Visual Question Answering Models

dc.contributor.guideSoni, Badal
dc.coverage.spatial
dc.creator.researcherChowdhury, Souvik
dc.date.accessioned2025-08-05T09:36:13Z
dc.date.available2025-08-05T09:36:13Z
dc.date.awarded2025
dc.date.completed2025
dc.date.registered2021
dc.description.abstractVisual Question Answering (VQA) combines computer vision and natural language processing to enable machines to answer questions about images. This task, essential for applications like personalized education, medical diagnostics, and interactive gaming, requires both image understanding and language comprehension. Despite advances in deep learning, VQA models struggle with dataset biases, diverse question types, and limited annotations, often relying on shortcuts from training data that hinder generalization. This thesis aims to address these challenges, improving the efficiency and power of VQA models. Firstly, we introduce two VQA frameworks to reduce training time and computational requirements without compromising accuracy. The first technique employs automatic mixed precision, using 32-bit precision (FP32) for loss calculation and 16-bit precision (FP16) for gradient computation during backpropagation. The second technique involves question segregation to minimize inference time. By classifying questions and training on reduced, category-specific datasets, we optimize model performance and accuracy measurement using weighted averages based on dataset composition. Secondly, to develop efficient VQA models, we propose a novel model that integrates both local and global image features. Unlike existing models that often overlook critical foreground and background information, our approach uses an ensemble of attention methods including traditional co-attention, encoder-decoder attention, and spatial-channel attention to extract visual features. This integration aims to improve understanding by capturing comprehensive contextual information from images.
dc.description.note
dc.format.accompanyingmaterialDVD
dc.format.dimensions
dc.format.extent220
dc.identifier.researcherid
dc.identifier.urihttp://hdl.handle.net/10603/656061
dc.languageEnglish
dc.publisher.institutionComputer Science and Engineering
dc.publisher.placeSilchar
dc.publisher.universityNational Institute of Technology Silchar
dc.relation
dc.rightsuniversity
dc.source.universityUniversity
dc.subject.keywordComputer Science
dc.subject.keywordComputer Science Artificial Intelligence
dc.subject.keywordEngineering and Technology
dc.titleDevelopment of Interpretable and Efficient Visual Question Answering Models
dc.title.alternative
dc.type.degreePh.D.

Files

Original bundle

Now showing 1 - 5 of 13
Loading...
Thumbnail Image
Name:
80_recommendation.pdf
Size:
405.01 KB
Format:
Adobe Portable Document Format
Description:
Attached File
Loading...
Thumbnail Image
Name:
abstract.pdf
Size:
12.8 MB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
chapter 1.pdf
Size:
387.17 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
chapter 2.pdf
Size:
1.14 MB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
chapter 3.pdf
Size:
4.12 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.79 KB
Format:
Plain Text
Description: