Score
Design and implement reconstruction and rendering systems that represent 3D scene geometry and appearance as collections of anisotropic Gaussian primitives and render them using splatting-based differentiable or feedforward pipelines. This includes methods for initializing and optimizing Gaussian primitives, Bayesian uncertainty estimation, mirror/reflection-aware modeling, inpainting and compositional reconstruction, refinement and shell supervision, multi-GPU and real-time/ feedforward implementations, and tooling to export or integrate refined Gaussian-based assets.
This work presents a systematic survey of 3D Gaussian Splatting (3D GS), clarifying its theoretical foundations as an explicit radiance field representation paradigm, its key technical advancements, and its practical applicability boundaries. We introduce the first 3D GS knowledge graph, elucidating its intrinsic mechanisms—explicit spatial parameterization, differentiable rasterization, and strong editability—and rigorously distinguishing it from implicit NeRF-based approaches. Methodologically, we integrate adaptive Gaussian optimization, density-aware regularization, and multi-view geometric constraints to enable end-to-end training and real-time rendering (>100 FPS). Comprehensive experiments benchmark state-of-the-art models across reconstruction accuracy, inference speed, and editing flexibility, revealing critical limitations—including low geometric fidelity and weak support for dynamic scenes. Finally, we identify promising future directions: scalability to large-scale scenes, physical plausibility enforcement, and cross-modal integration with vision-language or sensor-fusion frameworks.
To address key limitations of 3D Gaussian Splatting (3DGS)—including high memory consumption, poor generalization due to light baking, and weak support for secondary-light effects—this paper proposes an explicit, differentiable, and decoupled 3D Gaussian lattice representation. Our method introduces illumination-decoupled modeling and a differentiable voxel-based splatting mechanism, explicitly separating geometry, albedo, and illumination components while enhancing modeling fidelity for shadows, specular reflections, and other secondary-light phenomena. The representation is fully compatible with standard graphics pipelines and enables feedforward real-time novel-view synthesis (>30 FPS). Evaluated on virtual human reconstruction, dynamic scene modeling, and generative tasks, our approach achieves significant gains: 37% lower memory footprint and 1.8× faster training compared to vanilla 3DGS, alongside improved visual fidelity. This work establishes a new paradigm bridging implicit-explicit hybrid representations and physics-informed rendering.
Gaussian splatting (GS) surface reconstruction suffers from limited representational capacity when using a single primitive type (e.g., ellipses or ellipsoids), hindering accurate modeling of complex geometries. To address this, we propose a hybrid-primitive Gaussian point latticization framework. Our method introduces, for the first time, differentiable geometric primitives—specifically ellipses and ellipsoids—into the GS representation. We further design a composite lattice construction strategy, a hybrid-primitive co-initialization mechanism, and an adaptive vertex pruning algorithm to enhance modeling flexibility and geometric fidelity. Experiments demonstrate that our approach significantly outperforms single-primitive baselines across diverse complex shapes, achieving consistent improvements in PSNR, SSIM, and surface geometric error metrics. This work establishes a more expressive and generalizable paradigm for GS-based surface reconstruction.
To address the challenge of jointly rendering RGB images, depth maps, surface normals, and semantic logits while preserving cross-scene geometric consistency in high-fidelity multimodal 3D reconstruction, this paper proposes a geometry-aware unified Gaussian rasterization framework. Our method introduces two key innovations: (1) a differentiable ray-ellipsoid intersection renderer that analytically derives gradients for depth and normal predictions, enabling joint optimization of Gaussian rotation and scale to enhance geometric fidelity; and (2) a CUDA-accelerated differentiable rasterization pipeline integrating learnable attributes and a differentiable pruning mechanism to balance efficiency and representational capacity. Evaluated on multiple benchmarks, our approach achieves state-of-the-art performance, significantly improving multimodal reconstruction quality and cross-view geometric consistency.
This work addresses the unclear trade-offs among reconstruction quality, model compactness, and rendering speed in dynamic 3D scene reconstruction. It presents the first systematic taxonomy and empirical evaluation of dynamic 3D Gaussian splatting methods, categorizing them into two paradigms: structure-guided approaches (e.g., deformation fields, canonical spaces, and meshes) and Gaussian-centric approaches (e.g., continuous functions and 4D representations). A comprehensive benchmarking study on the D-NeRF dataset reveals that structure-guided methods achieve superior reconstruction fidelity and model compactness, whereas Gaussian-centric methods enable faster rendering—often reaching real-time performance—at the cost of reduced quality stability and higher storage overhead. This study elucidates the fundamental trade-offs inherent in dynamic Gaussian splatting techniques and establishes a clear benchmark for future research.
Existing 3D reconstruction methods suffer from low efficiency and limited quality due to indirect geometric learning and coupled geometry-appearance modeling. To address this, we propose the first end-to-end jointly optimized framework integrating explicit triangular meshes with 3D Gaussian points: differentiable 3D Gaussians are rigidly bound to mesh faces and jointly optimized under photometric supervision to reconstruct both geometry and surface appearance. Our approach breaks the conventional paradigm of decoupled geometry and appearance modeling, enabling efficient, high-fidelity reconstruction and real-time rendering. Quantitatively, it achieves a +1.8 dB PSNR improvement over prior methods on the DTU and BlendedMVS benchmarks. Moreover, the framework supports interactive mesh editing and incremental updates for dynamic scenes, significantly enhancing reconstruction efficiency and editing flexibility.
This study investigates the structural properties and predictability limits of convergence solutions in multi-view optimization with 3D Gaussian Splatting (3DGS), focusing on how density distribution influences the coupling between geometry and appearance parameters. Through statistical analysis, learnability probing, and variance decomposition, the work reveals for the first time that rendering-optimal references (RORs) exhibit bimodal scale and radiance characteristics. A density-stratification mechanism is introduced to distinguish geometry-dominated regions from view-synthesis-dominated ones. Building on this insight, a rendering-free, density-aware predictor is developed, demonstrating that parameters in high-density regions can be accurately predicted from point clouds, whereas sparse regions suffer prediction failure due to strong parameter coupling. The proposed strategy substantially enhances training robustness.
This work addresses the inherent conflict in 3D Gaussian Splatting (3DGS) between jointly modeling appearance and geometry, where direct geometric extraction often degrades rendering quality. To resolve this, the authors propose a lightweight decoupling mechanism that assigns each Gaussian an independent geometric opacity parameter and employs an opacity-guided optimization strategy. Leveraging geometric priors from vision foundation models, this approach enables disentangled modeling of appearance and geometry. The method significantly improves both novel-view synthesis quality and geometric reconstruction accuracy across multiple datasets, demonstrating particularly strong performance in complex scenes containing transparent objects.
This work addresses the limitations of existing Gaussian splatting methods, which often suffer from multi-view inconsistencies and floating-point artifacts that hinder high-quality geometric reconstruction. To overcome these issues, the paper introduces a novel approach that rigorously models Gaussian primitives as stochastic entities and leverages their volumetric properties to construct an explicit geometric representation. This formulation enables high-fidelity depth map rendering and accurate extraction of fine-scale geometry. By establishing a geometrically interpretable theoretical foundation for Gaussian splatting and integrating multi-view geometric optimization, the proposed method achieves state-of-the-art performance in shape reconstruction accuracy and consistency, significantly outperforming existing approaches on public benchmarks.
This work addresses floating artifacts, flickering, and blurriness in 3D Gaussian splatting reconstructions of wild scenes, which arise from camera pose errors, insufficient coverage, and noisy geometric initialization. To resolve these issues, the authors propose a geometry-guided video-to-video generation approach that refines rendered outputs with temporal consistency. Their method introduces, for the first time, a geometry-aware video generation framework that constructs a Gaussian primitive video buffer using depth, normals, opacity, and covariance. Combined with a synthetic data training strategy capable of simulating diverse degradation patterns, this approach significantly enhances generalization. The method achieves state-of-the-art performance on novel view synthesis benchmarks, with an efficient variant running at 21 FPS, enabling interactive applications.
本文针对3D高斯点云方法在反射区域的表面塌陷问题,提出了一种基于物理的延迟渲染框架RGS,通过几何连续性学习和反射感知密集化策略来改善新视角合成的质量。