Blog(2)
VQA [1,2] is the field of research that aims to develop methods for answering natural language questions based on the information provided in corresponding images.
The widespread adoption of mobile devices has led to a rapid growth of video content that is captured, transmitted and shared on various social media platforms.
Research Areas(0)
Publications(3)
LEARNING TO JOINTLY SHARE AND PRUNE WEIGHTS FOR GROUNDING BASED VISION AND LANGUAGE MODELS
AuthorShangqian Gao,Burak Uzkent,Yilin Shen,Hongxia Jin
PublishedInternational Conference on Learning Representation (ICLR)
Date2023-05-01
Progressive Attention Memory Network for Movie Story Question Answering
AuthorJunyeong Kim, Minuk Ma, Kyungsu Kim, Sungjin Kim, Chang D. Yoo
PublishedComputer Vision and Pattern Recognition (CVPR)
Date2019-06-21
Gaining Extra Supervision via Multi-task learning
PublishedInternational Joint Conference on Neural Networks (IJCNN)
Date2019-02-08
News(4)
The competition was a Visual Question Answering (VQA) challenge focused on autonomous driving safety, using dashcam videos from the 2COOOL dataset. Participants had to develop models that accurately answer structured, numeric questions about incident-related details (e.g., weather, road conditions, primary/secondary entities, prevention measures, and points of impact).
Samsung demonstrated technical superiority in video VQA by securing 4th place in the 'VQA - AUTOPILOT CVPR' competition at the prestigious CVPR 2026
Large vision-language models (VLMs) have achieved impressive performance across a wide range of multimodal tasks, from visual question answering (VQA) to reasoning over images and text [1, 2]. However, these models often suffer from hallucinations and poor grounding when faced with knowledge-intensive queries.
Others(0)