FASA: Feedback-Aware Sampling Adaptation for Efficient Diffusion-Based VLA Models

📅 2026-09-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决基于扩散的VLA模型在边缘平台实时部署时的高计算和内存访问成本问题,提出FASA框架,通过动态调整采样步骤以适应实际交互反馈,从而提高推理速度。
📝 Abstract
Diffusion-based Vision-Language-Action (VLA) models achieve strong performance in embodied tasks, but their iterative sampling imposes heavy computational and memory-access cost, blocking real-time deployment on edge platforms. Existing acceleration methods either require expensive training (e.g., distillation, flow matching) or degrade perception via statically scheduled pruning and caching, ignoring the dynamic workload variance of robotic interactions. This paper presents FASA (Feedback-Aware Sampling Adaptation), a training-free runtime framework that treats real-time multimodal feedback as a control signal for the denoising pipeline: an interaction-driven range adaptor modulates the global sampling-step budget based on visual and gripper-force feedback, and a proprioception-aware step adaptor pinpoints the optimized step within the adapted range. This co-designed framework allows the underlying hardware architecture to adaptively match the workload demands of different execution phases. Comparative evaluations across several benchmarks show that the inference speed can be increased by up to 1.45$\times$ while maintaining competitive success rates, providing a novel dynamic runtime architecture paradigm for deploying heavy generative embodied AI workloads onto resource-constrained computing platforms.
Problem

Research questions and friction points this paper is trying to address.

diffusion-based VLA models
real-time deployment
computational cost
memory-access cost
dynamic workload variance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Feedback-Aware
Sampling Adaptation
Real-time Multimodal Feedback
Denoising Pipeline
💼 Related Jobs
No related jobs found.
Y
Yuchen Han
South China University of Technology, Guangzhou, China; Ping An Technology (Shenzhen) Co., Ltd., Shenzhen, China
J
Jianhan Wu
Ping An Technology (Shenzhen) Co., Ltd., Shenzhen, China
X
Xiaoyang Qu
Ping An Technology (Shenzhen) Co., Ltd., Shenzhen, China
Lingwei Kong
Lingwei Kong
Ping An Technology (Shenzhen) Co., Ltd.
federated learningprivacy preserving machine learninglarge language models
S
Shiyi Li
Harbin Institute of Technology, Shenzhen, China
Jianzong Wang
Jianzong Wang
Postdoctoral Researcher of Department of Electrical and Computer Engineering, University of Florida
Big DataStorage SystemCloud Computing