ProAct-VLM: Pre-Failure Vision-Language Task Replanning with Continuous Perception Feedback

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of long-horizon robotic tasks to failures caused by abrupt environmental changes by proposing an adaptive physically grounded task planning framework. The approach integrates vision-language models (VLMs) into a real-time perceptual feedback loop, overcoming the limitations of conventional discrete verification to establish a continuous perception mechanism. This enables a paradigm shift from reactive recovery to proactive prevention. By continuously monitoring the environment and performing instantaneous replanning, the framework ensures robust manipulation in dynamic scenarios. Extensive experiments across multiple baselines and diverse VLM backbones demonstrate that the proposed method significantly improves both the success rate and efficiency of dynamic, long-horizon manipulation tasks.
📝 Abstract
Long-horizon robotic tasks are vulnerable to unexpected environmental changes that can render planned actions ineffective or unsafe. To address this, robots must detect such changes as they occur, interpret their impact, and adjust their actions accordingly. Traditional rule-based decision-making pipelines are brittle in open-world conditions, as they are hand-tuned for specific scenarios and lack generalization. Vision-Language Models (VLMs) offer a promising alternative as they combine broad world knowledge with unified visual--text reasoning, enabling them to generalize across diverse scenarios and generate accurate, grounded task plans. However, for effective deployment in dynamic real-world settings, VLMs must be embedded into frameworks capable of handling uncertainty and environmental changes. Existing frameworks broadly address this reactively, triggering replanning only after execution failures or post-task checks, risking failed actions. Some methods verify conditions before actions, but these discrete checks miss changes occurring during execution. To address this, we present ProAct-VLM, an adaptive, physically grounded task planning framework that integrates VLMs within a real-time perception--feedback loop. ProAct-VLM continuously monitors the environment and re-plans as soon as relevant changes are detected, enabling adaptation before failure occurs. Evaluations against multiple baselines and across different VLM backbones show that our framework improves both success rates and efficiency in dynamic, long-horizon manipulation tasks. Project page: https://github.com/moured/ProAct-VLM
Problem

Research questions and friction points this paper is trying to address.

long-horizon robotic tasks
vision-language models
task replanning
dynamic environments
continuous perception feedback
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Task Replanning
Continuous Perception Feedback
Long-horizon Robotic Tasks
Pre-Failure Adaptation