🤖 AI Summary
This study addresses the challenge of determining optimal response timing in streaming video-language models, which struggle to assess whether accumulated visual evidence is sufficient for accurate answering. We discover that frozen models implicitly encode a linearly decodable "evidence readiness" signal and propose Readiness Gating, a strategy leveraging linear probes to extract this state and guide response timing without additional training or computational overhead. This approach significantly improves response accuracy at zero extra cost. Empirically, it achieves AUROC scores ranging from 0.733 to 0.905, outperforming conventional uncertainty estimation baselines. Furthermore, under equivalent waiting durations, our method yields accuracy improvements of up to 9.75 percentage points.
📝 Abstract
Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it. We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output. It decodes in all seven models of a shared byte-identical evaluation (AUROC 0.733-0.905 under the strictest not-ready sampling, where a fitted clock is near chance), and a probe fitted without any of a benchmark family's footage still reads that family. It is question-conditioned: on byte-identical windows, changing only the question reverses the readout on 66.1% of pairs, while every question-blind control is at chance by construction. The model can answer incorrectly and still encode readiness: AUROC remains 0.722 among wrong answers. Readiness also beats uncertainty estimators and their supervised combination on latency-matched answer selection, and tracks independent human judgments more closely than confidence. Released streaming triggers are also linear readouts, yet a trained trigger read on its own base model's activations is approximately orthogonal to readiness and decodes it far less accurately than a probe. We turn the readout into Readiness Gating, an answer-timing policy that improves accuracy by up to +9.75 pp at matched video duration with negligible computational overhead. How much it gains varies with the accuracy headroom the task makes available: across 26 configurations the gain tracks that headroom, and an intervention that moves it over identical pixels moves the gain with it.