Discrete Forcing: Infusing Discrete Guidance into Continuous Denoising for Few-Step Action Experts

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off in Vision-Language-Action (VLA) models between the low precision of discrete actions and the slow inference of continuous actions. To resolve this, we propose a discrete forcing framework that employs an explicit coarse-to-fine generation mechanism, first predicting discrete structures to guide subsequent continuous refinement. Built upon a shared backbone network for single-pass forward propagation, the approach integrates flow matching, diffusion transformers, and hybrid discrete-continuous representation learning. Extensive evaluations across multiple benchmarks demonstrate that our method significantly enhances performance while accelerating inference. Furthermore, it exhibits exceptional effectiveness in real-world, high-precision dynamic manipulation tasks, successfully unifying computational efficiency with action accuracy.
📝 Abstract
Efficient action generation in vision-language-action (VLA) models requires capturing both coarse action structure and fine-grained details. Discrete action tokens provide compact structural representations but sacrifice precision, while continuous action tokens offer high precision but often require multiple denoising steps. We introduce Discrete Forcing, a flow-matching framework that combines these representations through an explicit coarse-to-fine generation process. It first predicts discrete action tokens to establish a coarse action structure, then uses them to guide continuous action refinement. The discrete and continuous components share a common diffusion transformer backbone with specialized branches, maintaining a parameter count comparable to a conventional single-branch model while requiring only one forward pass per branch. Extensive evaluations across multiple benchmarks demonstrate improved performance and faster inference over a parameter-matched continuous action expert, with consistent performance gains as model capacity increases. Real-world experiments further demonstrate improvements on high-precision and dynamic manipulation tasks.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
action generation
discrete action tokens
continuous denoising
few-step inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Discrete Forcing
Flow Matching
Vision-Language-Action Models
Coarse-to-Fine Generation
Diffusion Transformer
🔎 Similar Papers
No similar papers found.