Institution profile

BCAI

Research institution
Research library22linked papers
Opportunities0open roles
Selected work

Representative Papers

SkillWeave: Weaving Heterogeneous Demonstrations into Long-Horizon Manipulation Skills

Oct 08, 2026

This study addresses the challenge of collecting demonstration data for long-horizon dexterous manipulation that simultaneously captures macro-level task progression and micro-level contact interactions. To this end, it proposes a novel heterogeneous demonstration framework that matches interaction modalities by integrating teleoperation with kinesthetic teaching. Furthermore, a mask-conditioned diffusion policy supervised via offline segmentation is designed to resolve visual mismatches, while a successor-aware steering algorithm is introduced to enable smooth policy transitions and mitigate distribution shift. Experimental results demonstrate an end-to-end success rate of 27%, with dexterous subtask success rates improving to 65% and the average policy composition efficiency reaching 87%.

0 citationsRead paper

Refine Connections, Close the Gap: A Reliable Enhancement Framework for Driving Scene Topology

Oct 07, 2026

This study addresses the theoretical performance gap in topological connectivity reasoning for autonomous driving, where existing threshold-based methods yield unreliable and logically inconsistent decision graphs. To overcome this, we propose TopoEnhance, a framework that formulates topological enhancement as a denoising process. By integrating denoising diffusion models, stochastic corruption reconstruction, and discrete topological optimization, the framework restores structural consistency and effectively resolves logical conflicts through denoised reconstruction. Functioning as a source-agnostic, plug-and-play module, TopoEnhance significantly boosts the performance of multiple state-of-the-art baselines without requiring retraining. It achieves substantial improvements in both the continuous TOP score and the discrete TJS metric, pushing topological reliability toward its theoretical upper bound.

0 citationsRead paper

Prediction-powered Neural Architecture Search

Oct 01, 2026

This study addresses the challenges of scarce ground-truth performance labels and noisy zero-cost proxies in neural architecture search (NAS) by proposing PPNAS. This method pioneers the integration of predictive performance inference (PPI) into label-efficient NAS, effectively synergizing multi-source heterogeneous signals by fusing a limited number of ground-truth labels with abundant zero-cost proxy data. Specifically, it leverages ordinal information to construct debiased pairwise ranking supervision. Experimental results demonstrate that PPNAS achieves state-of-the-art end-to-end NAS performance under constrained evaluation budgets, establishing a new paradigm for low-cost architecture search.

0 citationsRead paper

From Fragments to Global Maps: Learning Vectorized Map Aggregation with Large Language Models

Oct 01, 2026

This study addresses the limitations of traditional handcrafted rules in handling road structure variations and prediction noise during high-definition map aggregation for autonomous driving by proposing the MapMergeLLM framework. This method formulates vector map aggregation as a conditional sequence generation task, leveraging large language models to directly predict global map polylines. Furthermore, it introduces a coordinate tokenizer integrated with geometry-aware pretraining, supplemented by a line-level association loss, which effectively decouples the model from specific upstream detectors through synthetic data training. Experimental results demonstrate that the proposed framework significantly outperforms heuristic baselines on the Argoverse2 and nuScenes datasets, achieving efficient map aggregation without requiring retraining for different detectors.

0 citationsRead paper

SymRegFlow: Symmetry-Regularized Flow Matching for Video World Models

Oct 01, 2026

This study addresses the reliance of video world models on expensive paired data and their lack of novel-view supervision under continuous camera poses. To this end, we propose a symmetric regularized flow matching framework. The method generates noisy anchors by geometrically warping source views, introducing masked dual-anchor supervision and cross-anchor denoising consistency. By integrating affine Gaussian surrogate modeling with geometric reprojection, it achieves multi-view consistent video generation without ground-truth novel-view supervision. We further provide theoretical proof that this regularization recovers the clean-reference optimal solution. Experiments on the nuScenes dataset demonstrate that our approach reduces Fréchet Video Distance (FVD) by over 31%, achieving state-of-the-art performance in FVD, FVMD, and FID metrics while exhibiting superior instance preservation capabilities.

0 citationsRead paper
Recent publications

Latest Papers

SkillWeave: Weaving Heterogeneous Demonstrations into Long-Horizon Manipulation Skills

Oct 08, 2026

This study addresses the challenge of collecting demonstration data for long-horizon dexterous manipulation that simultaneously captures macro-level task progression and micro-level contact interactions. To this end, it proposes a novel heterogeneous demonstration framework that matches interaction modalities by integrating teleoperation with kinesthetic teaching. Furthermore, a mask-conditioned diffusion policy supervised via offline segmentation is designed to resolve visual mismatches, while a successor-aware steering algorithm is introduced to enable smooth policy transitions and mitigate distribution shift. Experimental results demonstrate an end-to-end success rate of 27%, with dexterous subtask success rates improving to 65% and the average policy composition efficiency reaching 87%.

0 citationsRead paper

Refine Connections, Close the Gap: A Reliable Enhancement Framework for Driving Scene Topology

Oct 07, 2026

This study addresses the theoretical performance gap in topological connectivity reasoning for autonomous driving, where existing threshold-based methods yield unreliable and logically inconsistent decision graphs. To overcome this, we propose TopoEnhance, a framework that formulates topological enhancement as a denoising process. By integrating denoising diffusion models, stochastic corruption reconstruction, and discrete topological optimization, the framework restores structural consistency and effectively resolves logical conflicts through denoised reconstruction. Functioning as a source-agnostic, plug-and-play module, TopoEnhance significantly boosts the performance of multiple state-of-the-art baselines without requiring retraining. It achieves substantial improvements in both the continuous TOP score and the discrete TJS metric, pushing topological reliability toward its theoretical upper bound.

0 citationsRead paper

Prediction-powered Neural Architecture Search

Oct 01, 2026

This study addresses the challenges of scarce ground-truth performance labels and noisy zero-cost proxies in neural architecture search (NAS) by proposing PPNAS. This method pioneers the integration of predictive performance inference (PPI) into label-efficient NAS, effectively synergizing multi-source heterogeneous signals by fusing a limited number of ground-truth labels with abundant zero-cost proxy data. Specifically, it leverages ordinal information to construct debiased pairwise ranking supervision. Experimental results demonstrate that PPNAS achieves state-of-the-art end-to-end NAS performance under constrained evaluation budgets, establishing a new paradigm for low-cost architecture search.

0 citationsRead paper

From Fragments to Global Maps: Learning Vectorized Map Aggregation with Large Language Models

Oct 01, 2026

This study addresses the limitations of traditional handcrafted rules in handling road structure variations and prediction noise during high-definition map aggregation for autonomous driving by proposing the MapMergeLLM framework. This method formulates vector map aggregation as a conditional sequence generation task, leveraging large language models to directly predict global map polylines. Furthermore, it introduces a coordinate tokenizer integrated with geometry-aware pretraining, supplemented by a line-level association loss, which effectively decouples the model from specific upstream detectors through synthetic data training. Experimental results demonstrate that the proposed framework significantly outperforms heuristic baselines on the Argoverse2 and nuScenes datasets, achieving efficient map aggregation without requiring retraining for different detectors.

0 citationsRead paper

SymRegFlow: Symmetry-Regularized Flow Matching for Video World Models

Oct 01, 2026

This study addresses the reliance of video world models on expensive paired data and their lack of novel-view supervision under continuous camera poses. To this end, we propose a symmetric regularized flow matching framework. The method generates noisy anchors by geometrically warping source views, introducing masked dual-anchor supervision and cross-anchor denoising consistency. By integrating affine Gaussian surrogate modeling with geometric reprojection, it achieves multi-view consistent video generation without ground-truth novel-view supervision. We further provide theoretical proof that this regularization recovers the clean-reference optimal solution. Experiments on the nuScenes dataset demonstrate that our approach reduces Fréchet Video Distance (FVD) by over 31%, achieving state-of-the-art performance in FVD, FVMD, and FID metrics while exhibiting superior instance preservation capabilities.

0 citationsRead paper