Score
Design and implement algorithms that choose or sample camera viewpoints or projection parameters to map spatial or high‑dimensional structures into 2D views. These methods build and evaluate single‑ or multi‑view projections that optimize view quality metrics such as readability, occlusion and crossing reduction, informativeness, diversity, and coverage.
This work proposes an efficient active view selection method to enhance 3D scene reconstruction quality by prioritizing camera viewpoints that maximize information gain. The approach introduces the COVER metric, which optimizes geometric coverage while minimizing a computationally tractable approximation of Fisher information gain, thereby favoring observations of under-explored regions. Unlike existing methods, COVER eliminates the need for costly transmittance estimation, offering an interpretable and lightweight view planning strategy that maintains robustness and computational efficiency. Experimental results demonstrate that the proposed method significantly outperforms current active view selection techniques across multiple real-world datasets and NeRF baselines integrated within Nerfstudio.
To address the challenge of modeling strong view-dependent phenomena—such as metallic specular highlights and fine-scale textures—in novel view synthesis, this paper introduces Camera Splatting. It is the first method to model cameras as differentiable 3D Gaussian distributions (“camera splats”) and deploy virtual point cameras near scene surfaces for continuous view sampling and optimization. Integrated within the 3D Gaussian Splatting framework, our approach jointly optimizes both the camera distribution parameters and the scene representation via differentiable rendering, enabling explicit, continuous modeling of view-dependent appearance. Experiments demonstrate that, compared to discrete sampling strategies like Farthest View Sampling, Camera Splatting achieves significantly improved reconstruction quality for complex view-dependent details—including metal reflections and textural patterns—while maintaining higher fidelity and consistency under extreme viewing angles.
This work addresses the generative problem of multi-view-consistent 3D reconstruction from a single 2D image. We propose a hierarchical probabilistic diffusion framework: first, a point-map-driven geometric prior module predicts a structured 3D point cloud; second, high-fidelity, geometrically consistent novel views are decoded conditioned on this point cloud. Our key contributions are (i) the first introduction of point-map representations as a multi-view geometric prior, enabling generalization to arbitrary input images; and (ii) a modular design that decouples geometric modeling from appearance synthesis, significantly improving cross-object transferability. Evaluated on real-world datasets—including ObjaverseXL and Google Scanned Objects—our method surpasses state-of-the-art approaches such as CAT3D and Free3D in both novel-view synthesis quality (FID, LPIPS) and 3D consistency (Chamfer distance), achieving new performance benchmarks.
This work addresses the limitations of existing feedforward view synthesis methods, which rely on Plücker ray representations that are highly sensitive to camera coordinate systems, resulting in poor cross-view geometric consistency. To overcome this, the authors propose a projection-conditioning strategy that replaces raw ray inputs with 2D projection cues from the target view, effectively reformulating the task as a stable image-to-image translation problem. A tailored masked autoencoder pretraining mechanism is introduced to leverage large-scale uncalibrated data under this new conditioning paradigm. The proposed approach significantly enhances model robustness and view consistency, achieving state-of-the-art performance across multiple novel view synthesis benchmarks. Notably, it outperforms ray-based baselines by a clear margin on geometric consistency metrics, demonstrating the effectiveness of decoupling geometry representation from explicit ray parameterization.
To address inefficient viewpoint selection and image redundancy in autonomous mobile robot inspection under occluded environments, this paper proposes a sampling-driven Next-Best-View (NBV) planning framework. Methodologically, it introduces a novel information reward modeling mechanism that jointly leverages ray tracing and Gaussian process interpolation to precisely quantify observation uncertainty. Furthermore, it replaces conventional grid search and gradient-based optimization with derivative-free optimization, significantly enhancing both the efficiency and robustness of candidate viewpoint search. Evaluated in simulation and real-world experiments across multiple robotic platforms, the proposed approach reduces image acquisition volume by 37% on average while improving coverage completeness of critical regions by a factor of 2.1—outperforming state-of-the-art NBV methods.
This study addresses the issue that redundant views increase computational overhead and degrade reconstruction quality in feed-forward novel view synthesis. We propose a lightweight, rendering-free keyframe selector that integrates geometric and image features to construct a multi-criteria scoring system based on coverage, redundancy, and clarity. By employing a genetic algorithm for offline optimal subset search and knowledge distillation to train a compact network, our method breaks away from the conventional selection paradigm that relies on reconstruction feedback. Experiments across six datasets demonstrate that the proposed approach outperforms existing baselines while significantly reducing selection costs. Notably, the curated subsets achieve superior performance compared to full-sequence inputs and exhibit strong cross-paradigm generalization capabilities.
To address the labor-intensive parameter tuning and poor generalizability of conventional stereo matching algorithms (SGBM+WLS) in UAV-based forestry applications, this paper proposes the first genetic algorithm (GA) framework for automated joint optimization of SGBM and WLS parameters. The method eliminates manual intervention by introducing a multi-objective evaluation function integrating Mean Squared Error (MSE), Peak Signal-to-Noise Ratio (PSNR), and Structural Similarity Index (SSIM), thereby significantly enhancing disparity map quality and cross-illumination-condition generalization. Evaluated on radiata pine branch imagery, the optimized configuration achieves a 42.86% reduction in MSE, an 8.47% improvement in PSNR, and a 28.52% gain in SSIM over the baseline—while maintaining real-time processing capability suitable for resource-constrained onboard UAV systems.
This work addresses the challenge of balancing rendering speed, model size, and performance under sparse input views—a key limitation for deploying novel view synthesis methods on resource-constrained devices. The authors propose a new approach based on Multi-Plane Image (MPI) representation, leveraging depth maps predicted by vision foundation models for geometric initialization. A one-step diffusion mechanism is introduced to jointly optimize the MPI representation in a differentiable manner and enhance the final rendered output. This design significantly improves scene completeness and visual fidelity from sparse views. Compared to representative 3D Gaussian splatting methods, the proposed method achieves a 30.7% faster inference speed and reduces model size to only 14.8% of the baseline, while delivering competitive synthesis quality in forward-facing scenes.
Traditional two-dimensional graph visualizations struggle to simultaneously preserve high-dimensional structural fidelity, aesthetic quality, and readability. This work proposes a novel approach that integrates high-dimensional graph embeddings with differentiable optimization. By introducing a differentiable proxy for edge crossings and jointly optimizing visual criteria such as angular resolution, the method searches for informative two-dimensional views within a multi-view projection space. The accompanying interactive system, DataFly, enables users to explore latent structures effectively. Experimental results demonstrate that the generated layouts surpass standard graph drawing techniques in both aesthetics and readability, and even outperform methods specifically tailored to optimize individual visual metrics. A user study further confirms the approach’s efficacy in revealing complex graph structures.
This work addresses the challenge of insufficient geometric consistency and reconstruction fidelity in novel view synthesis under sparse observations and without camera pose information. The authors propose ReNoV, a framework that, for the first time, systematically leverages the geometric and semantic correspondences embedded in the spatial attention of external visual representations. By designing a dedicated representation projection module, ReNoV injects these correspondences as conditioning signals into a diffusion-based generative process, enabling high-quality view synthesis without explicit pose supervision. Evaluated on standard benchmarks, ReNoV significantly outperforms existing diffusion-based methods, achieving notable improvements in reconstruction fidelity, image inpainting quality, and geometric consistency.