EZBlender: Efficient 3D Editing with Plan-and-ReAct Agent

📅 2026-01-12
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inefficiency of traditional 3D scene editing and the challenge faced by existing vision-language agents in balancing precision with responsiveness. We propose EZBlender, a novel hybrid agent framework that integrates deliberative task decomposition with reactive local autonomy through a pioneering Plan-and-ReAct architecture. This design ensures semantic consistency and high editing quality while substantially reducing latency and computational overhead. To rigorously evaluate our approach, we introduce the first systematic benchmark for multi-task 3D editing, on which EZBlender demonstrates significant advantages over prior methods in terms of language model preference, response speed, and computational cost.

Technology Category

Computer Vision: 3D Computer VisionHumans and AI: Human-Aware Planning and Behavior PredictionPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
As a cornerstone of the modern digital economy, 3D modeling and rendering demand substantial resources and manual effort when scene editing is performed in the traditional manner. Despite recent progress in VLM-based agents for 3D editing, the fundamental trade-off between editing precision and agent responsiveness remains unresolved. To overcome these limitations, we present EZBlender, a Blender agent with a hybrid framework that combines planning-based task decomposition and reactive local autonomy for efficient human AI collaboration and semantically faithful 3D editing. Specifically, this unexplored Plan-and-ReAct design not only preserves editing quality but also significantly reduces latency and computational cost. To further validate the efficiency and effectiveness of the proposed edge-autonomy architecture, we construct a dedicated multi-tasking benchmark that has not been systematically investigated in prior research. In addition, we provide a comprehensive analysis of language model preference, system responsiveness, and economic efficiency.
Problem

Research questions and friction points this paper is trying to address.

3D editing
editing precision
agent responsiveness
human-AI collaboration
computational efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Plan-and-ReAct
3D editing
Blender agent
edge-autonomy
task decomposition
🔎 Similar Papers
No similar papers found.