Score
Designs and implements a two-stage U-Net pipeline for tiled spatial fields where the first stage generates predictions on overlapping patches and overlap-aware blending strategies merge those tile outputs. The second-stage U-Net serves as postprocessing to refine the merged field, restoring sharp features and preserving peak intensities while reducing smoothing.
NeRF struggles to scale to global-scale Earth observation due to GPU memory constraints, limiting existing methods to small local scenes. To address this, we propose Snake-NeRF—a novel framework incorporating a 2×2 3D tile progressive partitioning strategy. Our approach leverages overlapping tiling, segment-wise ray sampling, and an out-of-core training mechanism, enabling end-to-end NeRF reconstruction of large-scale satellite imagery on a single GPU for the first time—eliminating tile-boundary artifacts entirely. Memory consumption during training scales linearly with scene size, while rendering quality remains stable. Experiments demonstrate efficient reconstruction of areas exceeding 100 km² and confirm scalability toward planetary-scale modeling. Snake-NeRF thus establishes the first practical, large-scene NeRF solution for remote sensing 3D reconstruction.
This work proposes a lightweight screen-space neural post-processing method to address the high computational cost of traditional geometric subdivision in real-time rendering of large-scale low-polygon models, which struggles to efficiently produce smooth silhouettes. By reformulating subdivision as an image-domain task—an approach novel to this domain—the method progressively deforms object contours through multi-scale convolutional operations. It further preserves visual consistency by analyzing discrepancies between geometric and shading normals and applying implicit texture remapping. Crucially, the technique is fully decoupled from the original mesh complexity, enabling it to generate smooth, coherent silhouettes comparable to those achieved by geometric subdivision at a constant per-frame computational cost, thereby significantly enhancing real-time rendering efficiency for large-scale scenes.
Traditional image stitching methods suffer from poor scalability, and existing automated approaches are largely confined to single-image texture generation, failing to support cross-domain, multi-image collaborative seamless tiling. To address this, we propose Tiled Diffusion—the first diffusion-based framework natively extended to general-purpose image tiling generation. It supports arbitrary topological connections (e.g., self-tiling, many-to-many stitching) and diverse application domains including textures, 360° panoramas, and game assets. Methodologically, we introduce latent-space tiling constraints, inter-tile feature consistency regularization, and a multi-scale boundary fusion mechanism to ensure both visual seamlessness and semantic coherence. Experiments demonstrate substantial improvements in automation and creative flexibility for texture synthesis, image extrapolation, and 360° content generation, achieving high-quality, spatially consistent outputs across multiple tiling tasks.
Neural radiance field (NeRF) modeling for unconstrained real-world scenes—such as tourist photo collections—remains challenging due to the violation of the closed-world assumption and lack of semantic priors. Method: This work breaks this assumption by integrating pre-trained CNN/Vision Transformer semantic priors into the K-Planes planar scene representation framework. We propose a prior-guided alternating optimization strategy: building upon voxel-based representations, we jointly optimize geometry and appearance through multi-stage feature distillation and rendering loss. Contribution/Results: Our method achieves significant improvements in novel-view synthesis quality on both synthetic and real-world outdoor imagery, yielding richer geometric and textural details. Quantitatively, it consistently outperforms state-of-the-art baselines across PSNR and SSIM metrics. This establishes a new paradigm for efficient, high-fidelity NeRF modeling in open-domain scenes.
To address the demand for high-fidelity, diverse 3D asset generation and flexible editing, this paper introduces SLAT—a structured 3D implicit representation that jointly encodes sparse 3D mesh topology and multi-view visual foundation model features, enabling unified decoding into multiple 3D formats (e.g., radiance fields, 3D Gaussians, explicit meshes). Methodologically, SLAT pioneers a 2B-parameter Transformer architecture based on Rectified Flow for large-scale 3D latent-space modeling—the first of its kind. We curate a high-quality dataset of 500K 3D assets and perform end-to-end training. SLAT supports text- and image-conditioned generation, achieving state-of-the-art performance in fidelity, diversity, and editability. It enables real-time local 3D editing and dynamic output format switching. All code, models, and data are publicly released.
This work addresses the challenge of generating thermal imagery from RGB inputs in aerial scenes, where paired RGB–thermal datasets are scarce. To this end, the authors propose a conditional U-Net architecture that incorporates weather condition metadata embedded into the bottleneck layer. The method further enhances image fidelity by integrating saturation and contrast adjustments as preprocessing steps and applying Gaussian blur as postprocessing within a Pix2Pix GAN framework. Systematic experiments demonstrate the critical role of auxiliary environmental information and tailored image processing in improving generation quality. Evaluated via five-fold cross-validation on a dataset of 612 image pairs, the proposed model significantly outperforms the ThermalGen baseline, achieving a PSNR of 14.55, an SSIM of 0.8095, and an LPIPS score as low as 0.1666.
This work addresses the challenge of fragmented or erroneously fused geometry commonly produced by existing methods when reconstructing large-scale real-world scenes from unstructured, non-overlapping in-the-wild images. To achieve global consistency, the authors propose a novel modeling framework that leverages semantic alignment to jointly optimize the 6DoF pose and scale between local reconstructions and a geographically accurate pseudo-synthetic reference model generated via Google Earth Studio. The reference model is represented using 3D Gaussian Splatting enriched with semantic features, enabling robust registration through inverse feature optimization. To support this task, the authors also introduce the WikiEarth dataset. Experiments demonstrate that the proposed approach significantly improves global alignment accuracy across both classical and learning-based reconstruction pipelines, effectively mitigating failure modes prevalent in end-to-end models.
While Gaussian splatting achieves strong performance in novel view synthesis, its requirement of millions of primitives for highly textured scenes incurs prohibitive storage and computational overhead. This work addresses real-time novel view synthesis for sparse-geometry scenes by proposing *Nexels*: a hybrid representation that decouples geometry (surfels) from appearance—modeled jointly via a global NeRF-style neural field and per-surfel color parameters. To our knowledge, this is the first approach to co-model neural-textured surfels with a fixed-sampling neural field. Rendering employs pixel-level sparse texture sampling, enabling efficient representation without compromising visual fidelity. Experiments demonstrate significant gains: for outdoor scenes, voxel count reduces by 9.7× and memory usage by 5.5×; for indoor scenes, voxel count drops by 31× and memory by 3.7×; rendering speed improves by 2×, while visual quality surpasses existing textured primitive methods.