Scaling Video Generation for Reasoning: At What Cost?

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how scaling video generation models affects latent-state reasoning capabilities and computational costs. Leveraging an autoregressive architecture and a Rubik’s Cube prediction benchmark, we systematically evaluate the trade-off between reasoning performance and compute consumption across varying model scales. Our findings reveal that reductions in mean squared error do not necessarily enhance reasoning ability, that smaller models offer high efficiency under limited compute budgets, and that symbolic supervision complements scaling effectively. Experimentally, a 70M-parameter model achieves 44.6% accuracy using only 0.1 PF-days of compute. Furthermore, introducing symbolic supervision boosts the accuracy of a 20M-parameter model from 31.1% to 67.3%, establishing a new paradigm for computationally efficient reasoning.
📝 Abstract
We study whether scaling video generation enables models to reason about hidden information from the past frames, and at what computational cost. Our controlled benchmark requires predicting nine prescribed moves of an initially solved 2x2x2 Rubik's Cube from a fixed view of three faces. Correct predictions require inferring how actions change hidden states, and the simulator provides exact ground truth for evaluation. Models learn plausible cube geometry early, while correct sticker configurations require substantially more training. Although validation MSE follows approximate power-law scaling, lower MSE loss does not reliably indicate downstream reasoning capabilities. Smaller autoregressive models achieve higher state accuracy with limited compute, while larger models reach higher accuracy after more training. At roughly 0.1 PF-days, the 70M-parameter model correctly predicts the visible sticker configuration in 44.6% of post-action frames, compared with 0.3% for the 1B model, which reaches 83.7% at 3.14 PF-days. Symbolic state supervision raises the 20M model's frame accuracy from 31.1% to 67.3% at the same training-data budget, suggesting that learning representations of state changes can complement scaling.
Problem

Research questions and friction points this paper is trying to address.

video generation
reasoning
scaling
computational cost
hidden state inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Generation
Reasoning
Scaling Laws
Symbolic State Supervision
Autoregressive Models