Institution profile

42dot Inc.

Industry researchasia · kr
Official website
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

OneFixer: High-Quality and Consistent One-Step Autoregressive 3DGS Refinement for Driving Scenes

Sep 26, 2026

This study addresses the limitations of single-step autoregressive 3D Gaussian Splatting (3DGS) generation for autonomous driving simulation, including degraded visual quality, temporal inconsistency, and high deployment latency. To this end, we propose OneFixer, a novel framework that introduces a deployment-matched shared rolling mechanism, integrating flow matching with pixel-level perceptual supervision conditioned on lane geometry and dynamic agent states. This design enables efficient single-stage training and high-quality real-time rendering without requiring multi-stage distillation or bidirectional translation. Experimental results demonstrate that OneFixer achieves state-of-the-art performance on metrics such as FVD, halves GPU inference time, reduces closed-loop simulation collision rates by one-third, and exhibits significantly superior temporal consistency compared to existing baselines.

0 citationsRead paper

DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models

Jun 22, 2026

This work addresses the challenge that existing hybrid reasoning models struggle to dynamically allocate reasoning budgets, often leading to over-reasoning on simple problems and under-reasoning on complex ones. The authors propose a training-free adaptive routing mechanism that samples two zero-thought drafts and uses their consistency to decide whether to answer directly; if inconsistent, it predicts the required reasoning budget based on draft entropy. This approach achieves dynamic budget allocation for the first time without labeled data or gradient updates, leveraging the model’s own generated signals for decision-making and demonstrating compatibility across diverse model scales and architectures. Evaluated on mathematical and code reasoning tasks, the method improves accuracy by up to 9.0 and 22.5 percentage points while reducing reasoning tokens by 15–69% and 51–63%, respectively, confirming its efficiency and generality.

0 citationsRead paper

VR-Drive: Viewpoint-Robust End-to-End Driving with Feed-Forward 3D Gaussian Splatting

Oct 27, 2025

To address the degradation of end-to-end autonomous driving robustness caused by inter-camera viewpoint discrepancies, this paper proposes VR-Drive—a multi-view robust end-to-end framework. Methodologically, it introduces feedforward 3D Gaussian splatting for unsupervised novel-view synthesis, jointly optimizing 3D scene reconstruction and trajectory planning; it further designs a cross-view hybrid memory bank and a consistency distillation mechanism to enable online augmentation under sparse-view conditions and cross-view temporal modeling. Experiments on a custom-built multi-view benchmark demonstrate that VR-Drive significantly suppresses synthesis artifacts, markedly improving planning generalization and stability under unseen viewpoints. The framework establishes a novel paradigm for scalable, real-world deployment of end-to-end autonomous driving systems.

0 citationsRead paper

CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection based View Transformation

Sep 06, 2025

To address the degradation of camera-radar fusion performance in back-projection-based BEV transformation caused by image depth ambiguity, this paper proposes CRAB: a novel framework that (1) explicitly constrains image depth distributions using high-precision sparse depth priors from radar, thereby mitigating depth ambiguity during inverse projection; and (2) introduces a radar-context-enhanced cross-attention mechanism to achieve fine-grained alignment and fusion of image features with radar occupancy information directly in BEV space. CRAB jointly integrates inverse projection, view-specific feature aggregation, and spatially adaptive radar fusion into a single end-to-end trainable architecture for high-fidelity BEV representation learning. Evaluated on nuScenes, CRAB achieves 62.4% NDS and 54.0% mAP—setting the new state of the art among back-projection-based camera-radar fusion methods for 3D detection and semantic segmentation.

0 citationsRead paper
Recent publications

Latest Papers

OneFixer: High-Quality and Consistent One-Step Autoregressive 3DGS Refinement for Driving Scenes

Sep 26, 2026

This study addresses the limitations of single-step autoregressive 3D Gaussian Splatting (3DGS) generation for autonomous driving simulation, including degraded visual quality, temporal inconsistency, and high deployment latency. To this end, we propose OneFixer, a novel framework that introduces a deployment-matched shared rolling mechanism, integrating flow matching with pixel-level perceptual supervision conditioned on lane geometry and dynamic agent states. This design enables efficient single-stage training and high-quality real-time rendering without requiring multi-stage distillation or bidirectional translation. Experimental results demonstrate that OneFixer achieves state-of-the-art performance on metrics such as FVD, halves GPU inference time, reduces closed-loop simulation collision rates by one-third, and exhibits significantly superior temporal consistency compared to existing baselines.

0 citationsRead paper

DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models

Jun 22, 2026

This work addresses the challenge that existing hybrid reasoning models struggle to dynamically allocate reasoning budgets, often leading to over-reasoning on simple problems and under-reasoning on complex ones. The authors propose a training-free adaptive routing mechanism that samples two zero-thought drafts and uses their consistency to decide whether to answer directly; if inconsistent, it predicts the required reasoning budget based on draft entropy. This approach achieves dynamic budget allocation for the first time without labeled data or gradient updates, leveraging the model’s own generated signals for decision-making and demonstrating compatibility across diverse model scales and architectures. Evaluated on mathematical and code reasoning tasks, the method improves accuracy by up to 9.0 and 22.5 percentage points while reducing reasoning tokens by 15–69% and 51–63%, respectively, confirming its efficiency and generality.

0 citationsRead paper

VR-Drive: Viewpoint-Robust End-to-End Driving with Feed-Forward 3D Gaussian Splatting

Oct 27, 2025

To address the degradation of end-to-end autonomous driving robustness caused by inter-camera viewpoint discrepancies, this paper proposes VR-Drive—a multi-view robust end-to-end framework. Methodologically, it introduces feedforward 3D Gaussian splatting for unsupervised novel-view synthesis, jointly optimizing 3D scene reconstruction and trajectory planning; it further designs a cross-view hybrid memory bank and a consistency distillation mechanism to enable online augmentation under sparse-view conditions and cross-view temporal modeling. Experiments on a custom-built multi-view benchmark demonstrate that VR-Drive significantly suppresses synthesis artifacts, markedly improving planning generalization and stability under unseen viewpoints. The framework establishes a novel paradigm for scalable, real-world deployment of end-to-end autonomous driving systems.

0 citationsRead paper

CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection based View Transformation

Sep 06, 2025

To address the degradation of camera-radar fusion performance in back-projection-based BEV transformation caused by image depth ambiguity, this paper proposes CRAB: a novel framework that (1) explicitly constrains image depth distributions using high-precision sparse depth priors from radar, thereby mitigating depth ambiguity during inverse projection; and (2) introduces a radar-context-enhanced cross-attention mechanism to achieve fine-grained alignment and fusion of image features with radar occupancy information directly in BEV space. CRAB jointly integrates inverse projection, view-specific feature aggregation, and spatially adaptive radar fusion into a single end-to-end trainable architecture for high-fidelity BEV representation learning. Evaluated on nuScenes, CRAB achieves 62.4% NDS and 54.0% mAP—setting the new state of the art among back-projection-based camera-radar fusion methods for 3D detection and semantic segmentation.

0 citationsRead paper