Text-to-3D Policy: Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of text-to-3D policies in generalizing to unseen fine-grained behavioral specifications by proposing the T3DP framework. T3DP integrates point cloud-based 3D diffusion policies with fine-grained language-action alignment techniques, achieving precise mapping between linguistic instructions and actions through local structure preservation. Furthermore, it introduces a novel bidirectional token-level correspondence mechanism that effectively prevents the collapse of similar specifications within the representation space. Experimental results demonstrate that T3DP improves success rates by 11–14 percentage points in simulation tasks and significantly increases real-world performance from 47.5% to 65.0%, validating its superior generalization capability in embodied manipulation tasks.
📝 Abstract
3D visuomotor policies provide a strong foundation for spatially precise manipulation, yet current text-to-3D policies struggle to follow unseen fine-grained behavioral specifications beyond those covered by demonstrations. We study this challenge as unseen specification generalization, where language specifies behaviorally significant variations, such as target position, displacement, or articulated state, that are absent from policy training. We find that pretrained language representations and conventional global behavior-language alignment capture coarse task semantics but often blur nearby specifications that require distinct behaviors. We introduce T3DP, a Text-to-3D Policy framework for fine-grained language-behavior alignment. Rather than compressing each instruction and demonstration into a single global embedding, T3DP preserves their local structures and establishes bidirectional token-level correspondence between linguistic elements and behavioral segments. This directly grounds subtle linguistic variations in the behavior components they affect, preventing closely related specifications from collapsing in the representation space. The resulting specification-sensitive language representation conditions a point-cloud-based 3D diffusion policy, enabling more precise control over unseen behavioral specifications without modifying the underlying policy architecture. Across Meta-World, ManiSkill, and RoboTwin, T3DP improves average held-out-specification success over global language-behavior alignment by +11.0-14.2 points, with gains on all 15 task families; on real-robot tasks, it further raises average success from 47.5% to 65.0% (+17.5 points). Representation and action-probe analyses show that fine-grained alignment better preserves specification geometry and action-relevant variation, linking local behavior grounding to downstream control.
Problem

Research questions and friction points this paper is trying to address.

Text-to-3D Policy
Unseen Specification Generalization
Fine-grained Language-Behavior Alignment
Robotic Manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-grained language-behavior alignment
Token-level correspondence
Unseen specification generalization
3D diffusion policy
Text-to-3D