🤖 AI Summary
This study addresses the issue of unreliable editing in text-driven 3D Gaussian editing caused by inconsistent multi-view supervision quality. To this end, we propose a view-importance-aware framework that estimates keyframe significance by integrating geometric visibility analysis with semantic discriminability computation. The method guides the editing process around reliable keyframes while suppressing noisy back-propagated interference through asymmetric signal propagation and importance-aware optimization mechanisms. Experimental results demonstrate that our approach achieves the highest CLIP similarity across 23 scenes, requires only four minutes per scene for editing, and effectively preserves cross-view consistency.
📝 Abstract
Text-driven 3D Gaussian editing commonly does not distinguish the editing reliability of rendered views, although different viewpoints provide supervision of substantially different quality. Views that clearly show the scene and match the edit instruction provide reliable guidance, while less informative views may weaken the edit when all views are treated equally. We present View Matters, a view-importance-aware framework that conducts editing around reliable keyframes. Keyframe Importance Estimation (KIE) identifies reliable views using geometric visibility, semantic distinctiveness, and edit relevance. Keyframe-Guided Editing (KGE) then propagates their editing signals asymmetrically to non-keyframes without noisy reverse influence, while Importance-Aware Optimization (IAO) preserves this reliability preference during 3DGS optimization. Across 23 scene-prompt pairs, View Matters achieves the highest average CLIP text-image similarity of 0.2822 and directional similarity of 0.2564 among the evaluated methods, with a four-minute editing time. Additional adjacent-view analysis indicates that the fidelity-oriented editing process maintains cross-view coherence.