Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question Answering

Authors

  • Jianwen Jiang Tsinghua University
  • Ziqiang Chen Tsinghua University
  • Haojie Lin Tsinghua University
  • Xibin Zhao Tsinghua University
  • Yue Gao Tsinghua University

DOI:

https://doi.org/10.1609/aaai.v34i07.6766

Abstract

Understanding questions and finding clues for answers are the key for video question answering. Compared with image question answering, video question answering (Video QA) requires to find the clues accurately on both spatial and temporal dimension simultaneously, and thus is more challenging. However, the relationship between spatio-temporal information and question still has not been well utilized in most existing methods for Video QA. To tackle this problem, we propose a Question-Guided Spatio-Temporal Contextual Attention Network (QueST) method. In QueST, we divide the semantic features generated from question into two separate parts: the spatial part and the temporal part, respectively guiding the process of constructing the contextual attention on spatial and temporal dimension. Under the guidance of the corresponding contextual attention, visual features can be better exploited on both spatial and temporal dimensions. To evaluate the effectiveness of the proposed method, experiments are conducted on TGIF-QA dataset, MSRVTT-QA dataset and MSVD-QA dataset. Experimental results and comparisons with the state-of-the-art methods have shown that our method can achieve superior performance.

Downloads

Published

2020-04-03

How to Cite

Jiang, J., Chen, Z., Lin, H., Zhao, X., & Gao, Y. (2020). Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question Answering. Proceedings of the AAAI Conference on Artificial Intelligence, 34(07), 11101-11108. https://doi.org/10.1609/aaai.v34i07.6766

Issue

Section

AAAI Technical Track: Vision