[CVPR 2021] SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning over Traffic Events
-
Updated
Aug 31, 2026 - JavaScript
[CVPR 2021] SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning over Traffic Events
[EMNLP 2023] TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding
[ICLR2024] Codes and Models for COSA: Concatenated Sample Pretrained Vision-Language Foundation Model
Unifying the Video and Question Attentions for Open-Ended Video Question Answering
Video Question Answering via Hierarchical Spatio-Temporal Attention Networks
These #LangChain-powered apps include a Research Assistant that generates reports using web scraping and GPT-4o-mini, and Chat With Video, a #Streamlit app that #transcribes videos and enables content-based Q&A via #embeddings.
AI-powered Bilibili video Q&A agent — ask anything about any video in the comment section. 在 B 站评论区 @机器人,即可对任意视频提问。
These #LangChain-powered apps include a Research Assistant that generates reports using web scraping and GPT-4o-mini, and Chat With Video, a #Streamlit app that #transcribes videos and enables content-based Q&A via #embeddings.
[EMNLP 2025] D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
[EMNLP 2026] DynaTokens: Controlling Token Dynamics for Continual Video-Language Understanding
AI 科技口播与演示视频 Codex Skill:真实界面、官方素材、连续口播、分镜与逐镜头 QA | Evidence-first AI video workflow
视频问答助手(歪比巴布 WabbyBabu):B站 + YouTube 视频字幕问答、跟随式翻译、课堂笔记、重点问题、成就系统,零后端自配 OpenAI 兼容 API Key
pip install vidqa-cli — local, deterministic video QA CLI + MCP for tests, CI, and agents: 28 commands — frame timing, golden-frame diff, failure auto-locate, on-screen text index, color-aware run-to-run diff, Playwright/adb capture, local-VLM rubric review. One compact JSON + exit code per call. $0, fully on-device.
An open-source multimodal Video RAG framework for long-video understanding, action-level retrieval, ASR/OCR fusion and VLM reasoning.
These #LangChain-powered apps include a Research Assistant that generates reports using web scraping and GPT-4o-mini, and Chat With Video, a #Streamlit app that #transcribes videos and enables content-based Q&A via #embeddings.
The teaches you to integrate text, images, and videos into applications using Gemini's state-of-the-art multimodal models. Learn advanced prompting techniques, cross-modal reasoning, and how to extend Gemini's capabilities with real-time data and API integration.
Frame-first cross-platform video review tool for Windows x64 and Apple Silicon macOS.
[ACL BioNLP2026] LAMAR-2 at MedGenVidQA@BioNLP2026: Visual Answer Localization in Medical Videos via Multimodal LLM and Context-Augmented Prompting. Shared task website:
To associate your repository with the video-qa topic, visit your repo's landing page and select "manage topics."