ASSEMBLE: Atomic Skills for Evidence-Grounded Video Reasoning

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of scattered evidence and uninterpretable decision-making in long video reasoning by proposing a structured reasoning framework with explicit evidence grounding. The method constructs a timestamped evidence catalog and generates citation-supported reasoning trajectories through atomic skill composition. It innovatively introduces correctness-gated citation alignment as a grounding signal, jointly optimizing answer accuracy and citation consistency via supervised fine-tuning with teacher supervision and Group Relative Policy Optimization (GRPO). Experimental results demonstrate that the proposed approach achieves a macro-average accuracy of 59.2% across three benchmarks, surpassing Gemini-2.5-Pro, while improving grounding accuracy by 6.7%.
📝 Abstract
Complex video reasoning often depends on evidence scattered across distant moments, entities, and events, yet a correct answer alone does not reveal whether a model relied on the right parts of the video. We introduce ASSEMBLE, a framework that makes supporting evidence explicit throughout long-video reasoning. ASSEMBLE organizes local observations and cross-clip narratives into timestamped evidence catalogs traceable to the source video. A grounding-aware reader then composes question-specific atomic skills whose structured outputs contain explicit evidence references and support assessments. We use correctness-gated citation alignment as a direct grounding signal: after teacher-supervised fine-tuning, Group Relative Policy Optimization (GRPO) jointly optimizes answer correctness and citation alignment. This produces inspectable intermediate traces while keeping final predictions linked to explicit supporting evidence. Using a 9B reader supervised by a 235B teacher and shared precomputed evidence catalogs, ASSEMBLE achieves 59.2% macro-averaged answer accuracy across three long-video reasoning benchmarks, compared with 58.3% for Gemini-2.5-Pro, while improving macro-averaged overlap-based Grounded accuracy by 6.7%, with gains on all three benchmarks. Ablations further show that, with the same post-trained reader and inference budget, structured skill inference improves Grounded accuracy over free-form reasoning. Together, these results show that explicit evidence grounding can be integrated directly into long-video reasoning without sacrificing answer accuracy.
Problem

Research questions and friction points this paper is trying to address.

Long-video reasoning
Evidence grounding
Video reasoning
Citation alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evidence-Grounded Video Reasoning
Atomic Skills
Group Relative Policy Optimization (GRPO)
Citation Alignment
Long-Video Reasoning
🔎 Similar Papers
No similar papers found.