SkillAgentSearch skills...

Video Question Answering Resources

Video Question Answering | Video QA | VQA

Install / Use

npx skills add chakravarthi589/Video-Question-Answering_Resources

Installs into whichever agent you are using.

README

<h1 align="center" id="video-qa"> Video-Question-Answering (VideoQA) Resources </h1>

The Video-Question-Answering-Resources repository is a curated guide for beginners and researchers interested in the Video Question Answering (VQA) field. It provides an organized collection of the most relevant papers, models, datasets, and additional resources to help users understand and contribute to this evolving area. The repository focuses on the intersection of computer vision and natural language processing, particularly how video data can be used to answer complex questions, offering a range of materials from introductory guides to advanced research. (Last Update on 09/22/2025)

Keywords:

Video question answering (VideoQA), LLMs, Long video understanding, Spatial Reasoning, Temporal Reasoning, Multi-Choice QA, Open-Ended QA;

Curators:

Bharatesh Chakravarthi, Ph.D </br> Joseph Raj Vishal



Beginners Guide to Video Question Answering

  1. Answering Questions from YouTube Videos with OpenAI Whisper and GPT-4 (Medium article)

  2. Try a quick example on how to use LLMs for Video Question Answering here (Check Additional Resources for API key)

  3. Community Computer Vision Course (Unit 4) MultiModal Models


Publications

Survey/Review Papers

  • Video Understanding with Large Language Models: A Survey (2025) <a href="https://ieeexplore.ieee.org/abstract/document/10982110?casa_token=6fBIWiRaaB4AAAAA:ccI-pIG5EyznAUCBLxPlyzJRZ58cwKffAlqtVGDexj-JTPIhCTKwv1PI53IQy7C9jrID6pY" target="_blank"/>[Paper]
  • VideoQA in the Era of LLMs:An Empirical Study (2025) <a href="https://link.springer.com/article/10.1007/s11263-025-02385-8" target="_blank">[Paper]
  • A Survey on Generative AI and LLM for Video Generative Understanding, and Streaming (2024) <a href="https://arxiv.org/abs/2404.16038" target="_blank">[Paper]
  • Video Question-Answering Techniques, Benchmark Datasets and Evaluation Metrics Leveraging Video Captioning: A Comprehensive Survey (2021) <a href="https://ieeexplore.ieee.org/abstract/document/9350580" target="_blank">[Paper]
  • Video Question Answering: a Survey of Models and Datasets (2021) <a href="https://link.springer.com/article/10.1007/s11036-020-01730-0#ref-CR57" target="_blank">[Paper]
  • A survey on VQA: Datasets and approaches (2020, ITCA) <a href="https://doi.org/10.1109/ITCA52113.2020.00069" target="_blank">[Paper]

Conference/Journal Papers

2026

  • GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents (ACL) <a href="https://arxiv.org/abs/2603.24329" target="_blank"/>[Paper]

2025

  • RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives <a href="https://arxiv.org/pdf/2503.21459" target="_blank"/>[Paper]
  • MovieChat+: Question-aware Sparse Memory for Long Video Question Answering <a href="https://ieeexplore.ieee.org/abstract/document/11146594?casa_token=wfEY3s6gMfcAAAAA:82X04yDX3HQV7oexWLDP9e8VrffBsEUsaJrIhappD7cNnSkcVG7LuKGHtgtsquprKcFYtA&signout=success" target="_blank"/>[Paper]
  • Object-centric Video Question Answering with Visual Grounding and Referring <a href="https://arxiv.org/abs/2507.19599" target="_blank"/>[Paper]
  • VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering <a href="https://arxiv.org/abs/2508.03039" target="_blank"/>[Paper]
  • LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering <a href="https://arxiv.org/abs/2507.14784" target="_blank"/>[Paper]
  • Enhancing Long Video Question Answering with Scene-Localized Frame Grouping <a href="https://arxiv.org/abs/2508.03009" target="_blank"/>[Paper]
  • VisiQuest:Video-Question Answering with Advanced Vision-Language AI <a href="https://ieeexplore.ieee.org/abstract/document/11118375?casa_token=Su8gwzgyc1YAAAAA:vvdSGRNuljHW-J8a80LbsO_KwdOLg6WyuInTxerl3OzIT91EsmIYPlWuyjgMtF3IrvflCw" target="_blank"/>[Paper]
  • Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning (CVPR) <a href="https://openaccess.thecvf.com/content/CVPR2025/html/Liu_Commonsense_Video_Question_Answering_through_Video-Grounded_Entailment_Tree_Reasoning_CVPR_2025_paper.html" target="_blank"/>[Paper]
  • Question-Answering Dense Video Events (ACM) <a href="https://dl.acm.org/doi/pdf/10.1145/3726302.3729945" target="_blank"/>[Paper]
  • CogStream: Context-guided Streaming Video Question Answering <a href="https://arxiv.org/pdf/2506.10516" target="_blank"/>[Paper]
  • VRAG: Retrieval-Augmented Video Question Answering for Long-Form Videos <a href="https://openaccess.thecvf.com/content/CVPR2025W/IViSE/papers/Gia_VRAG_Retrieval-Augmented_Video_Question_Answering_for_Long-Form_Videos_CVPRW_2025_paper.pdf" target="_blank"/>[Paper]
  • Advancing Egocentric Video Question Answering with Multimodal Large Language Models <a href="https://arxiv.org/pdf/2504.04550" target="_blank"/>[Paper]
  • Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering <a href="https://arxiv.org/abs/2503.14957" target="_blank"/>[Paper]
  • MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering <a href="https://arxiv.org/abs/2506.18071" target="_blank"/>[Paper]
  • MELA: Multi-Event Localization Answering Framework for Video Question Answering <a href="https://dl.acm.org/doi/pdf/10.1145/3672608.3707973" target="_blank"/>[Paper]
  • HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding <a href="https://arxiv.org/abs/2501.01645" target="_blank"/>[Paper]
  • BIMBA: Selective-Scan Compression for Long-Range Video Question Answering <a href="https://arxiv.org/abs/2503.09590" target="_blank"/>[Paper]
  • Semantic Distance-Aware Cross-Modal Attention Mechanism for Video Question Answering <a href="https://ieeexplore.ieee.org/abstract/document/11007207" target="_blank"/>[Paper]
  • Empowering LLMs with pseudo-untrimmed videos for audio-visual temporal understanding <a href="https://ojs.aaai.org/index.php/AAAI/article/view/32784" target="_blank"/>[Paper]
  • Leveraging LLMs with Iterative Loop Structure for Enhanced Social Intelligence in Video Question Answering <a href="https://arxiv.org/abs/2503.21190" target="_blank"/>[Paper]
  • Agentic Keyframe Search for Video Question Answering <a href="https://arxiv.org/abs/2503.16032" target="_blank"/>[Paper]
  • Admitting Ignorance Helps the Video Question Answering Models to Answer <a href="https://arxiv.org/abs/2501.08771" target="_blank"/>[Paper]
  • VQALS: A Video Question Answering Method in Low-Light Scenes Based on Illumination Correction and Feature Enhancement <a href="https://cje.ejournal.org.cn/article/doi/10.23919/cje.2023.00.403?viewType=HTML" target="_blank"/>[Paper]
  • Keyframe-oriented vision token pruning: Enhancing efficiency of large vision language models on long-form video processing<a href="https://arxiv.org/abs/2503.10742" target="_blank"/>[Paper]
  • TUMTraffic-Videoqa: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes<a href="https://arxiv.org/abs/2502.02449" target="_blank"/>[Paper]
  • EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering <a href="https://arxiv.org/abs/2502.07411" target="_blank"/>[Paper]
  • ReasVQA:Advancing VideoQA with Imperfect Reasoning Process<a href="https://arxiv.org/abs/2501.13536" target="_blank"/>[Paper]
  • Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMs <a href="https://arxiv.org/abs/2501.04336" target="_blank"/>[Paper]
  • Unhackable temporal rewarding for scalable video MLLMs <a href="https://arxiv.org/abs/2502.12081" target="_blank"/>[Paper]
  • Grounded multi-hop videoqa in long-form egocentric videos<a href="https://ojs.aaai.org/index.php/AAAI/article/view/32214" target="_blank"/>[Paper]
  • VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding <a href="https://arxiv.org/pdf/2501.13106" target="blank" />[Paper]
  • Videoqa-SC: Adaptive Semantic Communication for Video Question Answering <a href="https://ieeexplore.ieee.org/abstract/document/10960438?casa_token=fsw9cFFeC90AAAAA:fYx9MTLw7Vm8IWMJpO1k7B9i74-ULwvjtXm93oM0vEUvtiHCEhWYC9TMT04XHI4c5ppjTWc" target="blank"/>[Paper]
  • Granularity-Adaptive Spatial Evidence Tokenization for Video Question Answering <a href="https://ojs.aaai.org/index.php/AAAI/article/view/32416" target="blank" />[Paper]
  • Towards Fine-Grained Video Question Answering <a href="https://arxiv.org/abs/2503.06820" target=""/>[Paper]
  • Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models <a href="https://ojs.aaai.org/index.php/AAAI/article/view/34183" target=""/>[Paper]
  • REVEAL: Relation-based Video Representation Learning for Video-Question-Answering <a href="https://arxiv.org/abs/2504.05463" target=""/>[Paper]
  • VideoQA-TA: Temporal-Aware Multi-Modal Video Question Answering <a href="https://aclanthology.org/2025.coling-main.483/" target="blank"/>[Paper]
  • VideoMultiAgents: A Multi-Agent Framework for Video Question Answering <a href="https://arxiv.org/abs/2504.20091" target="blank" />[Paper]
  • Open-Ended and Knowledge-Intensive Video Question Answering <a href="https://arxiv.org/abs/2502.11747" target="blank"/>[Paper]
  • A CLIP-base

Related Skills

View on GitHub
GitHub Stars97
CategoryContent
Updated1d ago
Forks13

Security Score

100/100

Audited on Aug 6, 2026

No findings