gaussian primitive encoding

Design and implement encoders, predictors, and refinement networks that represent spatial scene elements as parameterized Gaussian primitives—learning shared Gaussian banks, predicting 3D Gaussian parameters, and refining primitive location, shape, and feature embeddings. Aggregate multi‑sensor or cross‑modal features into these Gaussians using attention and adaptive capacity allocation, and project/splat refined primitives into bird’s‑eye‑view representations to enable compact, vectorized map decoding and downstream reasoning.

gaussianprimitiveencoding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of existing online high-definition map construction methods, which rely on fixed-resolution bird’s-eye-view (BEV) grids and struggle to efficiently model sparse yet geometrically precise map elements. To overcome this, the authors propose an online mapping framework based on adaptive Gaussian representations, replacing conventional dense BEV grids with learnable Gaussian primitives that enable dynamic allocation of modeling resources and high-fidelity geometric representation in critical regions. The approach leverages Gaussian interaction modeling and multi-sensor feature fusion to generate structured BEV features through a feedforward encoder, which are subsequently decoded into vectorized maps. Evaluated on the nuScenes and Argoverse 2 datasets, the method achieves state-of-the-art performance under both pure vision and camera–LiDAR fusion settings.

bird's-eye-view representationgeometric localizationHD map construction

C3G: Learning Compact 3D Representations with 2K Gaussians

Dec 03, 2025
HA
Honggyu An
🏛️ KAIST AI | ETH AI Center | ETH Zürich | SONY AI | Sony Group Corporation

Existing pixel-wise 3D Gaussian splatting methods for feedforward 3D scene reconstruction and understanding from pose-free sparse views suffer from high redundancy and inefficient multi-view feature aggregation. Method: We propose Token-GS, a token-guided Gaussian splatting framework that learns compact, geometry-aware Gaussians at salient spatial locations via learnable tokens, coupled with attention-driven Gaussian decoding and multi-view self-attention feature aggregation. It unifies 2D-to-3D feature enhancement within a single architecture. Contribution/Results: Token-GS achieves strong geometric awareness and significantly reduced memory footprint. It substantially improves feature fidelity and generalization in pose-free novel view synthesis and 3D open-vocabulary segmentation. Remarkably, it attains state-of-the-art performance using only ~2K Gaussians—yielding substantial memory efficiency gains over prior approaches.

Improving multi-view feature aggregation for scene understandingReconstructing 3D scenes from unposed sparse viewsReducing redundant Gaussians to lower memory overhead

Existing feedforward multi-view Gaussian generation methods naively concatenate Gaussians across views, leading to geometric inconsistencies, rendering artifacts, and representational redundancy. To address this, we propose the Gaussian Graph Neural Network (G-GNN), the first approach to model multi-view Gaussians as a geometry-aware graph structure. G-GNN introduces a Gaussian-level message-passing mechanism to explicitly capture inter-view geometric relationships, incorporates differentiable Gaussian pooling for compact representation learning, and jointly enforces multi-view geometric constraints within an end-to-end trainable differentiable rendering framework. Evaluated on RealEstate10K and ACID, our method achieves superior PSNR and SSIM with significantly fewer Gaussians, faster rendering speed, and strong cross-scene generalization—outperforming state-of-the-art methods across all metrics.

Enhances rendering speed and image quality efficientlyImproves 3D Gaussian representation from multi-view imagesReduces artifacts and memory cost in scene synthesis

Existing general-purpose novel view synthesis methods employ fixed allocation strategies for Gaussian primitives, which struggle to adapt to spatial complexity variations across scenes, resulting in redundant resources in smooth regions and insufficient representation in detailed areas. This work proposes SplatWeaver, a framework that introduces dynamic primitive allocation within a feed-forward 3D Gaussian splatting architecture for the first time. By integrating a mixture-of-Gaussians expert model with a pixel-level routing mechanism—guided by high-frequency structural priors and enhanced through routing regularization—the method adaptively allocates Gaussian primitives according to local geometric complexity. Experiments demonstrate that SplatWeaver significantly outperforms state-of-the-art approaches using fewer Gaussians, consistently achieving higher rendering fidelity and improved detail reproduction across diverse scenes.

3D Gaussian Splattinggeneralizable novel view synthesishigh-frequency details

ProtoGS: Efficient and High-Quality Rendering with 3D Gaussian Prototypes

Mar 21, 2025
ZG
Zhengqing Gao
🏛️ MBZUAI | University of Melbourne | Alibaba Group

To address the deployment challenge of 3D Gaussian Splatting (3DGS) on edge devices—stemming from its prohibitively large number of Gaussian primitives—this paper proposes a prototype-based lightweight representation. It replaces the massive set of original Gaussians with a small, learnable set of 3D Gaussian prototypes. Prototypes are constructed via SfM anchor-guided grouping and structured K-means clustering, and jointly optimized with anchors to ensure both geometric consistency and rendering fidelity. This work introduces the first Gaussian prototype learning paradigm. Evaluated on both real-world and synthetic datasets, it achieves over 90% reduction in Gaussian count and a 2.1× speedup in rendering, while surpassing existing compression methods in PSNR and SSIM. Crucially, it enables real-time novel-view synthesis on resource-constrained edge devices for the first time.

Maintain rendering quality while improving efficiencyOptimize memory usage during prototype trainingReduce Gaussian primitives for lightweight device deployment

Latest Papers

What's happening recently
View more

This work addresses the redundancy and inefficiency inherent in feed-forward 3D Gaussian splatting methods, which generate primitives on a per-pixel basis. To overcome this limitation, the authors propose a structure-aware primitive merging approach that leverages saliency-guided adaptive superpixel segmentation to achieve spatially and appearance-consistent clustering. The method is integrated into a learnable encoder-merger-multi-resolution decoder architecture, enabling flexible trade-offs between rendering quality and computational efficiency. The proposed framework can compress Gaussian primitives produced by any feed-forward method to as few as 1/20 of their original count while preserving high-fidelity rendering. Extensive evaluations demonstrate that this approach significantly outperforms existing compression techniques in terms of reconstruction accuracy and robustness, while also enabling highly efficient rendering.

3D Gaussian splattingcompact representationfeed-forward reconstruction

Existing feedforward Gaussian reconstruction methods struggle to accumulate gradient-based optimization of scene density during training and lack explicit modeling of cross-temporal observations. To address these limitations, this work proposes the Learning Gaussian Structure (LGS) framework, which dynamically adjusts Gaussian structures through an intervention-guided density control mechanism and introduces a cross-time query module to explicitly aggregate multi-temporal features, thereby enhancing the reliability of attribute prediction. LGS is the first to incorporate an intervention-response-based strategy for adding or removing Gaussians, enabling structure adaptation during inference and overcoming the temporal modeling constraints inherent in feedforward approaches. Experiments on the Waymo Open Dataset and PandaSet demonstrate that the proposed method significantly outperforms existing techniques, achieving more accurate reconstruction of dynamic driving scenes.

cross-time aggregationdensity controldriving scene reconstruction

The field of 3D vision suffers from fragmented data representations, learning paradigms, and benchmarking protocols, leading to a lack of unified understanding regarding efficiency, fidelity, and scalability. This work proposes the first cohesive conceptual framework that integrates geometric representations—such as point clouds, meshes, voxels, and 3D Gaussians—with diverse learning paradigms—including 2D-supervised learning, implicit neural representations, and 4D modeling—and connects them to real-world application scenarios. By constructing a structured knowledge graph of 3D vision, the study systematically relates dataset design, supervision mechanisms, and task requirements, clarifying the trade-offs between efficiency and fidelity and charting pathways for multimodal geometric grounding. This framework offers systematic guidance for reconstruction, generation, and dynamic scene modeling, advancing the field toward a unified and efficient paradigm.

3D visionbenchmark fragmentationdata representation

Hot Scholars

SP

Sida Peng

Zhejiang University
Computer VisionComputer Graphics
TD

Tianchen Deng

Shanghai Jiao Tong University
RoboticsComputer Vision
YD

Yuan Dong

Fudan University; Alibaba
Computer VisionMedical Image ComputingMachine Learning
YG

Yu-Gang Jiang

Professor, Fudan University. IEEE & IAPR Fellow
Video AnalysisEmbodied AITrustworthy AI