DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing Vision-Language-Action (VLA) verifiers, which overlook the varying importance of state-decision pairs within trajectories, resulting in inefficient action selection during test-time scaling. To this end, we propose DiVeR, a framework that reweights verifier learning by estimating decision criticality. Without requiring additional annotations or environment interactions, our method innovatively leverages the dispersion of action representations to identify critical states, thereby substantially enhancing the discriminator capability of the verifier. Extensive evaluations on benchmarks such as LIBERO, alongside real-world robot experiments, demonstrate that DiVeR consistently improves task success rates with negligible inference overhead.
📝 Abstract
Scaling robot data and model capacity has improved Vision-Language-Action (VLA) policies, but further progress is constrained by the high cost of robotic data. Verifier-guided test-time scaling offers an efficient alternative by sampling multiple action candidates and selecting the one most likely to lead to task success at inference time. Existing classification-based verifiers learn from trajectory-level outcomes but treat all visited states equally, even though their value for candidate discrimination can vary across a trajectory. At many states, plausible actions are similar and provide limited discrimination signal, while only a sparse set of decision-critical states admits meaningfully different actions that can substantially affect downstream outcomes. To address this, we propose DiVeR, which estimates decision criticality from the dispersion of sampled action representations. DiVeR then uses this signal to reweight verifier learning toward states where action selection is most consequential, without requiring step-level annotations or additional environment interaction. Across LIBERO, RoboCasa, and real-world experiments on a Franka Research 3 robot, DiVeR consistently improves task success through more effective verifier-guided action selection, while adding negligible verifier inference overhead.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action (VLA)
test-time scaling
verifier learning
decision-critical states
robotics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Scaling
Vision-Language-Action
Verifier Learning
Decision Criticality
Action Selection
🔎 Similar Papers
No similar papers found.