Score
Designs and implements algorithms, data structures, and systems to model, represent, and render visual scenes and imagery—covering geometry and material representations, shading programs, illumination models, rasterization and ray-tracing renderers, and animation/simulation components. Builds and analyzes rendering pipelines, graphics hardware/software interfaces, and evaluation methods for performance, visual fidelity, numerical robustness, and perceptual quality.
This paper addresses the challenge that ray-based 3D Gaussian splatting (RayGS) struggles to balance rendering speed and quality under hardware rasterization, while suffering from aliasing due to MIP-level mismatch. We propose the first RayGS rendering framework fully compatible with standard GPU rasterization pipelines. Methodologically: (1) we introduce a geometry-driven ray–Gaussian integration model enabling efficient rasterization; (2) we design a MIP-aware anti-aliasing mechanism that unifies multi-scale sampling during both training and inference, eliminating aliasing artifacts. Our contributions are threefold: (i) the first real-time, high-fidelity novel-view synthesis solution based on RayGS; (ii) consistent superiority over existing RayGS methods in visual quality and rendering throughput across multiple benchmarks; and (iii) a practical pathway toward low-latency, high-fidelity applications such as VR and MR.
Remote sensing simulation faces challenges including heavy reliance on LiDAR data, extensive manual intervention, and difficulty in cross-spectral band modeling. Method: This paper proposes an end-to-end, physically grounded simulation framework driven solely by commercial satellite imagery. It automatically constructs 3D geometry from digital surface models (DSMs), infers material properties by fusing multi-source satellite imagery, and integrates physics-based rendering with radiative transfer modeling across a broad spectral range (200 nm–20 μm) to achieve fully automated, high-fidelity reconstruction of terrain, buildings, vegetation, and dynamic vehicles. Contribution/Results: To our knowledge, this is the first framework to eliminate LiDAR dependence and generate novel regional scenes without any manual intervention. Experiments demonstrate substantial reduction in modeling cost, full-spectrum support—from ultraviolet (UV) to long-wave infrared (LWIR)—for algorithm development and processing pipeline validation, and significant improvements in realism, generalizability, and scalability of geospatial scene simulation.
Traditional rendering approaches face significant challenges when handling complex non-fractal geometries, including high memory overhead, limited topological expressiveness, and inflexible animation control. This work proposes a general-purpose rendering framework based on composable function systems, leveraging GPU-accelerated mesh-free representations to enable efficient generation and manipulation of intricate objects. The framework introduces Quibble, a metaprogramming system that facilitates dynamic composition of functional components. By extending function systems beyond conventional rendering into broader visualization and simulation domains, the method supports modeling of topologically non-trivial structures, controllable in-between frame animation, and point-cloud deformation, while maintaining strong interoperability. Experimental results demonstrate that the framework achieves both high performance and expressive artistic control across diverse tasks, including image synthesis, animation authoring, and geometric simulation.
This work addresses the formal verification of vision-based autonomous systems, tackling the rigorous propagation of uncertainty under continuous camera pose variations during scene rendering—particularly challenging for Gaussian splatting due to non-differentiable operations such as matrix inversion and depth sorting, which hinder precision-efficiency trade-offs. We propose the first linear-relational abstraction method tailored for Gaussian splatting, integrating piecewise-linear bounding, tile- and batch-level parallelism, and memory-aware scheduling to construct a scalable, block-wise abstract rendering framework. Our approach enables sound uncertainty quantification over scenes with up to 750K Gaussians under continuous translation, rotation, and scene perturbations. Compared to state-of-the-art mesh-based abstraction, it achieves 2–14× speedup with zero precision loss, marking the first solution capable of real-time, reliable abstract rendering for large-scale Gaussian scenes.
This study addresses the challenges of low accuracy, high computational cost, and the difficulty of unifying real-time rendering with differentiable optimization in vector graphics rasterization. To this end, this work proposes Windfoil, an algorithm that for the first time treats both tasks as a unified problem. By deriving a closed-form analytical solution for the box-filtered winding number of quadratic Bézier contours, it eliminates the need for approximate sampling and enables efficient, GPU-friendly computation. Experimental results demonstrate that Windfoil surpasses Skia and Slug in rendering fidelity while maintaining comparable performance. Furthermore, in differentiable optimization tasks, it achieves equivalent or superior reconstruction quality with substantially reduced per-step computational costs, thereby supporting interactive rendering of scenes comprising tens of thousands of shapes.
该研究通过工作负载中心框架连接表示、算法和硬件架构,以优化3D高斯点渲染效率,识别出影响系统增益的关键因素。
This study addresses the limitation of existing video benchmarks in evaluating models' capacity to follow fine-grained procedural world events. To this end, it constructs a novel benchmark grounded in replayable world records, generating videos through synchronized multi-view rendering and agent representations. Furthermore, the work proposes a vision-language model-based logic-rendering alignment metric that enables fine-grained consistency verification from procedural states to visual outputs. This approach effectively quantifies entity control, long-term memory, and interaction success rates, thereby establishing a rigorous evaluation standard for the visual fidelity of programmable world models.
MeshSplatBench通过标准化评估和部署流程,解决了基于三角形的神经渲染在生产引擎中实际应用的问题。
This study addresses the lack of paired data and geometric evaluation benchmarks for video rendering-to-real translation by proposing the first end-to-end evaluation framework based on reconstructed digital twins. The framework establishes a benchmark comprising twelve scenes, providing editable 3D replicas, proxy renderings, and real target videos. By integrating neural reconstruction with visual geometry techniques, it automates data processing through static-dynamic decoupling and camera registration. Furthermore, combining multi-view verification with scene decomposition enables automated assessment of appearance fidelity and geometric-dynamic consistency. The project releases 1,496 frames of paired data with an 85.7% retention rate, achieving a depth si-RMSE of 0.2170 and an instance mIoU of 0.3673. These contributions significantly reduce manual modeling costs while effectively identifying causes of generation failures.
This study addresses the program-to-visual inconsistency arising when multimodal large language models (MLLMs) generate code and corresponding visual outputs. To mitigate this issue, it proposes a unified framework incorporating three core mechanisms—Persistent State (PEG), Traceable Generation Process (TGP), and Revision Awareness (REV)—which collectively establish a shared revision reference system to coordinate planning, execution, and feedback. By integrating MLLM-driven executable program generation, rendering backends, and closed-loop verification techniques, the framework enables controllable image and video synthesis. Extensive benchmark evaluations demonstrate that GPT-6-Astra achieves a 100% generation success rate within this framework. Furthermore, the experimental results reveal a significant discrepancy between general capability scores and actual visual generation performance, highlighting critical limitations in current evaluation paradigms for assessing the visual fidelity of MLLMs.