OGAM: Connecting Systematic Testing to Runtime Assurance through Object-Grounded Attention Monitoring for VLA Policies

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited coverage of existing Vision-Language-Action (VLA) policy benchmarks, which compromises deployment safety. We propose a safety framework bridging offline system testing and online runtime assurance. By integrating gradient-weighted visual attention projection, dynamic time warping, conformal prediction calibration, and scene instruction generation, the framework systematically reveals attention divergence phenomena. Furthermore, it leverages object grounding monitoring to enable failure-label-free anomaly detection and early intervention. Experimental results demonstrate that our approach successfully intercepts 87%–100% of failure cases across four policies while maintaining a false-stop rate below 5%, substantially mitigating undetected risks during real-world deployment.
📝 Abstract
Benchmarks expose vision-language-action (VLA) policies to few canonical instructions, while exhaustive deployment testing is impossible. We introduce Object-Grounded Attention Monitoring (OGAM), connecting systematic testing to runtime assurance: testing reveals attention divergence between successful and failed executions, and OGAM uses this signal to stop failures beyond the finite suite. We generate scene-grounded instructions through pairwise combinations of action templates and objects, and separately test meaning-preserving paraphrases. All 87 out-of-benchmark cases reveal problematic behavior across OpenVLA, OpenVLA-OFT, UniVLA, and $\pi_{0.5}$: none completes any of the 24 feasible instructions, while infeasible or hazardous requests also trigger behavior substitution. At each action query, we project gradient-weighted visual attention through object masks and group it by instruction role for comparison across tasks and policies. Dynamic time warping aligns this course with a successful reference despite speed differences; conformal calibration on successful episodes sets the early-stopping threshold for sustained deviations, with a nominal false-stop target of $\alpha=0.05$. Across four policies, OGAM stops 87-100% of failed episodes at median times of 5-12s within a 20s budget, with observed false-stop rates of 3-5%, without failure-labeled training. Finite testing thus identifies attention patterns that support online intervention before failure fully unfolds.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action (VLA) policies
Runtime assurance
Systematic testing
Out-of-distribution instructions
Safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

Object-Grounded Attention Monitoring
Vision-Language-Action Policies
Runtime Assurance
Dynamic Time Warping
Conformal Calibration
💼 Related Jobs
No related jobs found.