PRIME: Perception Feedback with Situational Memory Embeddings in VLA Models

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出PRIME,通过引入情境记忆来改进VLA模型的感知反馈机制,从而提升自动驾驶中的决策和导航性能。
📝 Abstract
Current Vision-Language-Action (VLA) models for autonomous driving operate primarily through feedforward inference across the perception--reasoning--planning hierarchy. While modern architectures maintain temporal recurrence within the perceptual module, early perception remains blind to downstream reasoning and navigation goals, processing visual inputs agnostically without prioritizing cues informed by prior decisions. To bridge this gap, this paper introduces PRIME, a learned feedback mechanism that conditions the VLA perceptual queries on a novel Situational Memory. By aggregating latent representations of past perception, reasoning, navigation goals, and predicted behaviors across an L-step window via cross-attention, PRIME enables intent-driven perceptual attention at minimal computational cost, adding only a maximum of 29.7M parameters (0.41% of the 7.3B-parameter base model). Evaluated on the Bench2Drive closed-loop benchmark, PRIME achieves a state-of-the-art Driving Score of 82.47 (+4.73 over ORION) and a Success Rate of 60.00% (+5.38 percentage points), the highest reported Driving Score among published VLAs trained on Think2Drive demonstrations.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
autonomous driving
perception
reasoning
navigation goals
Innovation

Methods, ideas, or system contributions that make the work stand out.

PRIME
Situational Memory
cross-attention
intent-driven perceptual attention
autonomous driving
🔎 Similar Papers
No similar papers found.