Beyond VQA: Generating Multi-word Answers and Rationales to Visual Questions

Radhika Dua,Sai Srinivas Kancheti,Vineeth N Balasubramanian

2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGITION WORKSHOPS (CVPRW 2021)（2021）

引用 20|浏览21

暂无评分

摘要

Visual Question Answering is a multi-modal task that aims to measure high-level visual understanding. Contemporary VQA models are restrictive in the sense that answers are obtained via classification over a limited vocabulary (in the case of open-ended VQA), or via classification over a set of multiple-choice-type answers. In this work, we present a completely generative formulation where a multi-word answer is generated for a visual query. To take this a step forward, we introduce a new task: ViQAR (Visual Question Answering and Reasoning), wherein a model must generate the complete answer and a rationale that seeks to justify the generated answer. We propose an end-to-end architecture to solve this task and describe how to evaluate it. We show that our model generates strong answers and rationales through qualitative and quantitative evaluation, as well as through a human Turing Test.

查看译文

关键词

generating multiword answers,rationale,Visual questions,Visual Question Answering,multimodal task,high-level visual understanding,contemporary VQA models,open-ended VQA,multiple-choice-type answers,completely generative formulation,multiword answer,visual query,complete answer,generated answer,end-to-end architecture,strong answers

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要