DynaPix: Can Vision-Language Models Identify the Exact Future?

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current vision-language models often produce plausible yet imprecise predictions of future physical states, lacking the ability to discriminate the actual future from alternatives. To address this limitation, this work introduces DynaPix, a novel benchmark that establishes the first verifiable evaluation paradigm for future prediction grounded in physics-based simulation. DynaPix leverages high-fidelity simulations to generate ground-truth future images with precise timestamps and evaluates models on their capacity to identify these authentic frames among highly similar distractors or within large-scale image galleries. The benchmark reveals a substantial performance gap between “event-triggered” and “pure time-anchored” prediction tasks—the latter performing near chance level. Experiments show that fine-tuning with simulation-derived ground truth partially mitigates this “time-anchoring gap,” though long-horizon prediction remains challenging.
📝 Abstract
Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking image, so the predicted state is never checked against the true one. We introduce DynaPix (Dynamic Pixels), a benchmark that makes prediction checkable. Given a video clip that stops before a key event and a question about a later moment, a model must pick the true future image from close candidates or a large gallery. The scenes come from a physics simulator, so the correct image and its time are known exactly and the wrong options are deliberately similar. Models often succeed when a visible event marks the target moment, but are near chance when only elapsed time marks it. Gallery search is harder still, as the true image rarely ranks first. People handle the elapsed-time items well, so the difficulty lies with the models, not the questions. Training on scene accounts drawn from the simulator's true record, not a teacher's guess, repairs much of this but not the longer elapsed time case. DynaPix thus exposes a temporal-anchoring gap: models attach a prediction to an event far better than to time itself.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
future prediction
temporal anchoring
physical scene understanding
model evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

DynaPix
temporal anchoring
vision-language models
future prediction
physics simulation