OneFixer: High-Quality and Consistent One-Step Autoregressive 3DGS Refinement for Driving Scenes

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of single-step autoregressive 3D Gaussian Splatting (3DGS) generation for autonomous driving simulation, including degraded visual quality, temporal inconsistency, and high deployment latency. To this end, we propose OneFixer, a novel framework that introduces a deployment-matched shared rolling mechanism, integrating flow matching with pixel-level perceptual supervision conditioned on lane geometry and dynamic agent states. This design enables efficient single-stage training and high-quality real-time rendering without requiring multi-stage distillation or bidirectional translation. Experimental results demonstrate that OneFixer achieves state-of-the-art performance on metrics such as FVD, halves GPU inference time, reduces closed-loop simulation collision rates by one-third, and exhibits significantly superior temporal consistency compared to existing baselines.
📝 Abstract
Autoregressive video diffusion is a promising render-time fixer for 3D Gaussian Splatting (3DGS) in autonomous-driving simulation, but deployment demands high visual quality and temporal consistency at low latency. This is especially hard for one-step causal generation, where each imperfect prediction immediately becomes context for subsequent frames. Existing approaches stabilize rollouts through staged training with multiple modules and rollout-aware regularization, yet one-step quality still falls short of what deployment requires. We introduce OneFixer, a one-step autoregressive video-diffusion fixer trained in a single task-specific adaptation stage. Our key idea is a deployment-matched shared rollout: the model's own one-step predictions serve as the causal context for flow matching, exposing training to deployment-time errors, while the same rollout receives direct pixel-space perceptual supervision to preserve fine detail. Because the predictions optimized for current-frame quality are exactly those reused as future context, fidelity and autoregressive robustness are learned jointly, without bidirectional-to-causal conversion or teacher-student distillation. OneFixer further exploits cues that driving simulation readily provides, lane geometry and dynamic-agent states, to improve geometric fidelity. On Waymo and proprietary driving scenes with 900-frame rollouts, OneFixer achieves the lowest FVD, LPIPS, and DISTS among all baselines at one step, with temporal consistency matching or exceeding multi-stage DMD pipelines. Under identical backbone and conditioning, it matches a multi-stage DMD-with-Self-Forcing pipeline in under half the GPU-hours and keeps improving beyond its plateau. In closed-loop simulation with a driving policy, OneFixer reduces the collision rate by a third relative to raw 3DGS rendering. Project page: https://onefixer-web.vercel.app/
Problem

Research questions and friction points this paper is trying to address.

3D Gaussian Splatting
Autoregressive video diffusion
One-step generation
Temporal consistency
Driving simulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

One-Step Autoregressive Video Diffusion
3D Gaussian Splatting Refinement
Shared Rollout Flow Matching
Driving Scene Simulation
Temporal Consistency
🔎 Similar Papers
2024-08-29arXiv.orgCitations: 30