Institution profile

NIO

Industry researchasia · cn
Official website
Research library24linked papers
Opportunities0open roles
Selected work

Representative Papers

ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory

Aug 29, 2025arXiv.org

Robot manipulation video generation suffers from data scarcity and 3D spatial ambiguity arising from 2D trajectory representations. To address these challenges, we propose the first diffusion-based framework integrating 3D occupancy-aware modeling and trajectory optimization. First, we construct a scene-level 3D occupancy map to ensure geometrically consistent scene understanding. Second, we optimize physically feasible end-effector trajectories in 3D space—replacing ambiguous 2D paths with explicit, collision-aware 3D motion priors. Third, we design a trajectory-conditioned latent diffusion model that synthesizes coherent, obstacle-avoiding manipulation videos in third-person view, end-to-end. Our approach eliminates reliance on error-prone 2D trajectory supervision and explicitly grounds video generation in 3D dynamics. Experiments demonstrate significant improvements over state-of-the-art methods in visual fidelity and action plausibility. Notably, our method autonomously generates realistic pick-and-place videos with minimal human annotation, substantially reducing dependence on costly labeled data.

6 citationsRead paper

PlanWAM: Planning-Shaped Future Representations for End-to-End Autonomous Driving

Oct 08, 2026

This work addresses the disconnect between prediction and planning in existing end-to-end autonomous driving world models, which neglect the practical utility of future representations for planning. We propose PlanWAM, the first framework to directly shape future latent representations via planning objectives. Specifically, it employs a temporal register pyramid to efficiently compress historical information and introduces a privileged future branch coupled with a hindsight-to-foresight distillation mechanism, enabling precise predictions tailored for trajectory planning. Evaluated on the NAVSIM benchmark, PlanWAM achieves 93.8 PDMS and 90.9 EPDMS, along with an HD-Score of 38.7 in closed-loop simulation, demonstrating state-of-the-art zero-shot planning performance.

0 citationsRead paper

Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR

Aug 04, 2026

This work addresses the challenge of cross-lingual optimization conflicts in multilingual large language model-based speech recognition, where joint training struggles to preserve language-specific characteristics. The authors propose a language-specific multi-teacher online policy distillation framework that integrates language routing with token-level knowledge fusion. To decouple language-specialized capabilities from general multilingual modeling, they introduce both static and dynamic acoustic prefix designs. Evaluated on a mixed benchmark comprising Mandarin, Chinese dialects, Cantonese, and English, the proposed method significantly outperforms reinforcement learning baselines and consistently surpasses all monolingual teacher models, demonstrating superior generalization performance.

0 citationsRead paper
Recent publications

Latest Papers

PlanWAM: Planning-Shaped Future Representations for End-to-End Autonomous Driving

Oct 08, 2026

This work addresses the disconnect between prediction and planning in existing end-to-end autonomous driving world models, which neglect the practical utility of future representations for planning. We propose PlanWAM, the first framework to directly shape future latent representations via planning objectives. Specifically, it employs a temporal register pyramid to efficiently compress historical information and introduces a privileged future branch coupled with a hindsight-to-foresight distillation mechanism, enabling precise predictions tailored for trajectory planning. Evaluated on the NAVSIM benchmark, PlanWAM achieves 93.8 PDMS and 90.9 EPDMS, along with an HD-Score of 38.7 in closed-loop simulation, demonstrating state-of-the-art zero-shot planning performance.

0 citationsRead paper

Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR

Aug 04, 2026

This work addresses the challenge of cross-lingual optimization conflicts in multilingual large language model-based speech recognition, where joint training struggles to preserve language-specific characteristics. The authors propose a language-specific multi-teacher online policy distillation framework that integrates language routing with token-level knowledge fusion. To decouple language-specialized capabilities from general multilingual modeling, they introduce both static and dynamic acoustic prefix designs. Evaluated on a mixed benchmark comprising Mandarin, Chinese dialects, Cantonese, and English, the proposed method significantly outperforms reinforcement learning baselines and consistently surpasses all monolingual teacher models, demonstrating superior generalization performance.

0 citationsRead paper

SparseStreet: Sparse Gaussian Splatting for Real-Time Street Scene Simulation

Jun 02, 2026

This work addresses the high storage overhead and low rendering efficiency of existing 3D Gaussian splatting methods in street-view reconstruction, which stem from their reliance on a large number of Gaussian primitives. To tackle this issue, the authors propose a general-purpose compression framework tailored for street scenes, introducing a hierarchical strategy that exploits the distinct characteristics of dynamic objects and static backgrounds. The framework employs a learnable node pruning mechanism to eliminate low-contribution primitives and applies secondary compression specifically to static regions. Evaluated on the Waymo and nuScenes datasets, the method achieves up to 80% reduction in the number of Gaussians while preserving high-fidelity reconstruction quality, thereby significantly improving both storage efficiency and rendering performance.

0 citationsRead paper