OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of omni-modal large models in passively processing long audiovisual sequences without actively acquiring cross-modal evidence. To this end, we propose an active inference agent framework that empowers models to dynamically determine when to look or listen, retrieving critical audiovisual evidence through multi-turn interactions to support reasoning. We innovatively introduce an audiovisual necessity objective to suppress unimodal shortcuts and construct a data engine for synthesizing multi-hop chains-of-thought. The model is optimized via supervised fine-tuning on the OmniTraj-170K dataset, followed by a two-stage reinforcement learning strategy with verifiable rewards. Experimental results demonstrate that our approach achieves adaptive cross-modal evidence search and significantly enhances audiovisual reasoning performance.
📝 Abstract
We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process: it dynamically decides whether to look or listen, and over which temporal window, to retrieve sparse but critical evidence across different modalities within long contexts. Through an iterative multi-turn protocol, the retrieved raw audio or visual segments are appended back into the context to support subsequent reasoning. To cold-start this capability, we build a data engine that synthesizes OmniTraj-170K, a corpus of multi-hop Chain-of-Thought trajectories with interleaved audio and visual evidence. We first supervise the model on these trajectories to instill multi-turn tool-use behavior, and then further optimize the policy via a two-stage reinforcement learning with verifiable rewards. Moreover, we introduce an Audio-Visual Necessity objective that explicitly rewards successful trajectories whose reasoning depends on both modalities, discouraging single-modality shortcuts. Extensive experiments across a wide range of benchmarks demonstrate that OmniSeek learns adaptive cross-modal evidence seeking and consistently improves audio-visual reasoning performance.
Problem

Research questions and friction points this paper is trying to address.

audio-visual reasoning
multi-turn interaction
long-context understanding
cross-modal evidence seeking
omni large language model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Framework
Multi-turn Audio-Visual Reasoning
Native Tool Integration
Reinforcement Learning
Chain-of-Thought
🔎 Similar Papers
No similar papers found.