Video Question Answering Resources
Video Question Answering | Video QA | VQA
Install / Use
npx skills add chakravarthi589/Video-Question-Answering_ResourcesInstalls into whichever agent you are using.
README
The Video-Question-Answering-Resources repository is a curated guide for beginners and researchers interested in the Video Question Answering (VQA) field. It provides an organized collection of the most relevant papers, models, datasets, and additional resources to help users understand and contribute to this evolving area. The repository focuses on the intersection of computer vision and natural language processing, particularly how video data can be used to answer complex questions, offering a range of materials from introductory guides to advanced research. (Last Update on 09/22/2025)
Keywords:
Video question answering (VideoQA), LLMs, Long video understanding, Spatial Reasoning, Temporal Reasoning, Multi-Choice QA, Open-Ended QA;
Curators:
Bharatesh Chakravarthi, Ph.D </br> Joseph Raj Vishal
- Beginners Guide to Video-Question-Answering <br>
- Publications <br/>
- Survey/Review Papers
- Conference/Journal Papers
- Datasets <br>
- Models <br>
- Additional Resources <br>
Beginners Guide to Video Question Answering
-
Answering Questions from YouTube Videos with OpenAI Whisper and GPT-4 (Medium article)
-
Try a quick example on how to use LLMs for Video Question Answering here (Check Additional Resources for API key)
Publications
Survey/Review Papers
- Video Understanding with Large Language Models: A Survey (2025) <a href="https://ieeexplore.ieee.org/abstract/document/10982110?casa_token=6fBIWiRaaB4AAAAA:ccI-pIG5EyznAUCBLxPlyzJRZ58cwKffAlqtVGDexj-JTPIhCTKwv1PI53IQy7C9jrID6pY" target="_blank"/>[Paper]
- VideoQA in the Era of LLMs:An Empirical Study (2025) <a href="https://link.springer.com/article/10.1007/s11263-025-02385-8" target="_blank">[Paper]
- A Survey on Generative AI and LLM for Video Generative Understanding, and Streaming (2024) <a href="https://arxiv.org/abs/2404.16038" target="_blank">[Paper]
- Video Question-Answering Techniques, Benchmark Datasets and Evaluation Metrics Leveraging Video Captioning: A Comprehensive Survey (2021) <a href="https://ieeexplore.ieee.org/abstract/document/9350580" target="_blank">[Paper]
- Video Question Answering: a Survey of Models and Datasets (2021) <a href="https://link.springer.com/article/10.1007/s11036-020-01730-0#ref-CR57" target="_blank">[Paper]
- A survey on VQA: Datasets and approaches (2020, ITCA) <a href="https://doi.org/10.1109/ITCA52113.2020.00069" target="_blank">[Paper]
Conference/Journal Papers
2026
- GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents (ACL) <a href="https://arxiv.org/abs/2603.24329" target="_blank"/>[Paper]
2025
- RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives <a href="https://arxiv.org/pdf/2503.21459" target="_blank"/>[Paper]
- MovieChat+: Question-aware Sparse Memory for Long Video Question Answering <a href="https://ieeexplore.ieee.org/abstract/document/11146594?casa_token=wfEY3s6gMfcAAAAA:82X04yDX3HQV7oexWLDP9e8VrffBsEUsaJrIhappD7cNnSkcVG7LuKGHtgtsquprKcFYtA&signout=success" target="_blank"/>[Paper]
- Object-centric Video Question Answering with Visual Grounding and Referring <a href="https://arxiv.org/abs/2507.19599" target="_blank"/>[Paper]
- VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering <a href="https://arxiv.org/abs/2508.03039" target="_blank"/>[Paper]
- LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering <a href="https://arxiv.org/abs/2507.14784" target="_blank"/>[Paper]
- Enhancing Long Video Question Answering with Scene-Localized Frame Grouping <a href="https://arxiv.org/abs/2508.03009" target="_blank"/>[Paper]
- VisiQuest:Video-Question Answering with Advanced Vision-Language AI <a href="https://ieeexplore.ieee.org/abstract/document/11118375?casa_token=Su8gwzgyc1YAAAAA:vvdSGRNuljHW-J8a80LbsO_KwdOLg6WyuInTxerl3OzIT91EsmIYPlWuyjgMtF3IrvflCw" target="_blank"/>[Paper]
- Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning (CVPR) <a href="https://openaccess.thecvf.com/content/CVPR2025/html/Liu_Commonsense_Video_Question_Answering_through_Video-Grounded_Entailment_Tree_Reasoning_CVPR_2025_paper.html" target="_blank"/>[Paper]
- Question-Answering Dense Video Events (ACM) <a href="https://dl.acm.org/doi/pdf/10.1145/3726302.3729945" target="_blank"/>[Paper]
- CogStream: Context-guided Streaming Video Question Answering <a href="https://arxiv.org/pdf/2506.10516" target="_blank"/>[Paper]
- VRAG: Retrieval-Augmented Video Question Answering for Long-Form Videos <a href="https://openaccess.thecvf.com/content/CVPR2025W/IViSE/papers/Gia_VRAG_Retrieval-Augmented_Video_Question_Answering_for_Long-Form_Videos_CVPRW_2025_paper.pdf" target="_blank"/>[Paper]
- Advancing Egocentric Video Question Answering with Multimodal Large Language Models <a href="https://arxiv.org/pdf/2504.04550" target="_blank"/>[Paper]
- Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering <a href="https://arxiv.org/abs/2503.14957" target="_blank"/>[Paper]
- MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering <a href="https://arxiv.org/abs/2506.18071" target="_blank"/>[Paper]
- MELA: Multi-Event Localization Answering Framework for Video Question Answering <a href="https://dl.acm.org/doi/pdf/10.1145/3672608.3707973" target="_blank"/>[Paper]
- HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding <a href="https://arxiv.org/abs/2501.01645" target="_blank"/>[Paper]
- BIMBA: Selective-Scan Compression for Long-Range Video Question Answering <a href="https://arxiv.org/abs/2503.09590" target="_blank"/>[Paper]
- Semantic Distance-Aware Cross-Modal Attention Mechanism for Video Question Answering <a href="https://ieeexplore.ieee.org/abstract/document/11007207" target="_blank"/>[Paper]
- Empowering LLMs with pseudo-untrimmed videos for audio-visual temporal understanding <a href="https://ojs.aaai.org/index.php/AAAI/article/view/32784" target="_blank"/>[Paper]
- Leveraging LLMs with Iterative Loop Structure for Enhanced Social Intelligence in Video Question Answering <a href="https://arxiv.org/abs/2503.21190" target="_blank"/>[Paper]
- Agentic Keyframe Search for Video Question Answering <a href="https://arxiv.org/abs/2503.16032" target="_blank"/>[Paper]
- Admitting Ignorance Helps the Video Question Answering Models to Answer <a href="https://arxiv.org/abs/2501.08771" target="_blank"/>[Paper]
- VQALS: A Video Question Answering Method in Low-Light Scenes Based on Illumination Correction and Feature Enhancement <a href="https://cje.ejournal.org.cn/article/doi/10.23919/cje.2023.00.403?viewType=HTML" target="_blank"/>[Paper]
- Keyframe-oriented vision token pruning: Enhancing efficiency of large vision language models on long-form video processing<a href="https://arxiv.org/abs/2503.10742" target="_blank"/>[Paper]
- TUMTraffic-Videoqa: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes<a href="https://arxiv.org/abs/2502.02449" target="_blank"/>[Paper]
- EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering <a href="https://arxiv.org/abs/2502.07411" target="_blank"/>[Paper]
- ReasVQA:Advancing VideoQA with Imperfect Reasoning Process<a href="https://arxiv.org/abs/2501.13536" target="_blank"/>[Paper]
- Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMs <a href="https://arxiv.org/abs/2501.04336" target="_blank"/>[Paper]
- Unhackable temporal rewarding for scalable video MLLMs <a href="https://arxiv.org/abs/2502.12081" target="_blank"/>[Paper]
- Grounded multi-hop videoqa in long-form egocentric videos<a href="https://ojs.aaai.org/index.php/AAAI/article/view/32214" target="_blank"/>[Paper]
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding <a href="https://arxiv.org/pdf/2501.13106" target="blank" />[Paper]
- Videoqa-SC: Adaptive Semantic Communication for Video Question Answering <a href="https://ieeexplore.ieee.org/abstract/document/10960438?casa_token=fsw9cFFeC90AAAAA:fYx9MTLw7Vm8IWMJpO1k7B9i74-ULwvjtXm93oM0vEUvtiHCEhWYC9TMT04XHI4c5ppjTWc" target="blank"/>[Paper]
- Granularity-Adaptive Spatial Evidence Tokenization for Video Question Answering <a href="https://ojs.aaai.org/index.php/AAAI/article/view/32416" target="blank" />[Paper]
- Towards Fine-Grained Video Question Answering <a href="https://arxiv.org/abs/2503.06820" target=""/>[Paper]
- Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models <a href="https://ojs.aaai.org/index.php/AAAI/article/view/34183" target=""/>[Paper]
- REVEAL: Relation-based Video Representation Learning for Video-Question-Answering <a href="https://arxiv.org/abs/2504.05463" target=""/>[Paper]
- VideoQA-TA: Temporal-Aware Multi-Modal Video Question Answering <a href="https://aclanthology.org/2025.coling-main.483/" target="blank"/>[Paper]
- VideoMultiAgents: A Multi-Agent Framework for Video Question Answering <a href="https://arxiv.org/abs/2504.20091" target="blank" />[Paper]
- Open-Ended and Knowledge-Intensive Video Question Answering <a href="https://arxiv.org/abs/2502.11747" target="blank"/>[Paper]
- A CLIP-base
Related Skills
qqbot-channel
385.5kQQ channel management skill. Use qqbot_channel_api for explicit QQ channel-management requests; confirm write, delete, and bulk actions before calling authenticated QQ Open Platform endpoints.
docs-writer
106.4kAlways use this skill when the task involves writing, reviewing, or editing files in the `/docs` directory or any `.md` files in the repository.
cpp
40.5kGuide Cursor to write modern C++ and CMake code with clear structure, RAII, const-correctness, and safe error handling.
gamemaker-gml
40.5kGameMaker Language (GML) rules for scripts, objects, events, rooms, data structures, and performance-minded game code
Security Score
Audited on Aug 6, 2026
