Task-Adaptive Grounded 3D-Programmers Using 2D VLMs

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottlenecks of data scarcity and insufficient reasoning in 2D vision-language models for 3D understanding by proposing the 3D-Prog framework. This framework introduces canonical coordinate anchoring and a task-adaptive dynamic feedback mechanism, integrating Euclidean space mapping with iterative closed-loop reasoning to transform 2D large models into geometry-aware 3D programmers without retraining. By transcending the limitations of existing paradigms, this approach enables open-vocabulary 3D understanding, manipulation, and generation. The resulting outputs exhibit strong consistency, interpretability, and high quality, establishing a novel paradigm for cross-modal 3D intelligence.
📝 Abstract
Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operate reliably in 3D by introducing 3D grounding and iterative feedback loops with two novel concepts: Canonical Coordinate Framing (CCF) and Task-Adaptive Feedback (TAF). CCF serves as a unified visual representation that anchors both inputs and outputs to a shared Euclidean coordinate system, solving common challenges in 3D grounding such as axis ambiguity, inconsistent metric scale, and floating references. Complementary to this structured framing of the 3D inputs, TAF closes the reasoning loop with task-adaptive dynamic feedback that enables 2D VLMs to perform varied open-vocabulary tasks within their native visual context. Building on this foundation, we introduce 3D-Prog, a 3D understanding, reasoning, and generation framework that jointly employs the capabilities of CCF and TAF together with powerful VLMs. Without requiring any retraining, 3D-Prog performs open-vocabulary 3D understanding, manipulation, and generation across both object-level and scene-level tasks. Our experiments show that the joint use of CCF and TAF transforms 2D VLMs into geometry-aware 3D programmers, achieving consistent, interpretable, and high-quality results across diverse 3D tasks.
Problem

Research questions and friction points this paper is trying to address.

3D understanding
vision-language models
3D grounding
open-vocabulary tasks
2D VLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Canonical Coordinate Framing
Task-Adaptive Feedback
3D grounding
Vision-Language Models
Open-vocabulary 3D tasks
🔎 Similar Papers
No similar papers found.