Institution profile

Institute of AI for Industries

Academic institution
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation

Sep 27, 2026

This study addresses the scarcity of training data and the inherent trade-off between correctness and performance when large language models generate GPU kernels. To tackle these challenges, we propose a co-evolutionary framework comprising two synergistic models: a Proposer and a Coder. Methodologically, we introduce a frontier-driven module to construct an automatic curriculum learning system that continuously generates high-quality training data. Furthermore, we design a Correctness-Aware Group Relative Policy Optimization (CA-GRPO) algorithm to facilitate alternating reinforcement learning between the two models. Experimental results demonstrate that our approach significantly outperforms advanced models such as Claude-4.5 on the KernelBench benchmark, achieving pass@1 scores of 75.8% for CUDA and 77.2% for Triton. These findings confirm that the proposed framework effectively achieves simultaneous improvements in both kernel generation correctness and execution performance.

0 citationsRead paper

Efficient Diffusion Planning with Temporal Diffusion

Nov 25, 2025

Diffusion-based planning methods suffer from high computational overhead, low decision frequency, and susceptibility to plan–reality discrepancies due to frequent full replanning. To address these issues, we propose the Temporal Diffusion Planner (TDP). TDP distributes the denoising process across the temporal dimension, enabling progressive, stepwise refinement of a blurred long-horizon plan—eliminating the need for per-step full replanning. It further introduces a state-consistency-driven automatic replanning strategy that enhances real-world alignment while preserving planning continuity. By integrating offline reinforcement learning with dynamic temporal denoising, TDP significantly reduces computational burden. On the D4RL benchmark, TDP achieves an 11–24.8× increase in decision frequency over baseline diffusion planners, while matching or exceeding their task performance.

0 citationsRead paper
Recent publications

Latest Papers

KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation

Sep 27, 2026

This study addresses the scarcity of training data and the inherent trade-off between correctness and performance when large language models generate GPU kernels. To tackle these challenges, we propose a co-evolutionary framework comprising two synergistic models: a Proposer and a Coder. Methodologically, we introduce a frontier-driven module to construct an automatic curriculum learning system that continuously generates high-quality training data. Furthermore, we design a Correctness-Aware Group Relative Policy Optimization (CA-GRPO) algorithm to facilitate alternating reinforcement learning between the two models. Experimental results demonstrate that our approach significantly outperforms advanced models such as Claude-4.5 on the KernelBench benchmark, achieving pass@1 scores of 75.8% for CUDA and 77.2% for Triton. These findings confirm that the proposed framework effectively achieves simultaneous improvements in both kernel generation correctness and execution performance.

0 citationsRead paper

Efficient Diffusion Planning with Temporal Diffusion

Nov 25, 2025

Diffusion-based planning methods suffer from high computational overhead, low decision frequency, and susceptibility to plan–reality discrepancies due to frequent full replanning. To address these issues, we propose the Temporal Diffusion Planner (TDP). TDP distributes the denoising process across the temporal dimension, enabling progressive, stepwise refinement of a blurred long-horizon plan—eliminating the need for per-step full replanning. It further introduces a state-consistency-driven automatic replanning strategy that enhances real-world alignment while preserving planning continuity. By integrating offline reinforcement learning with dynamic temporal denoising, TDP significantly reduces computational burden. On the D4RL benchmark, TDP achieves an 11–24.8× increase in decision frequency over baseline diffusion planners, while matching or exceeding their task performance.

0 citationsRead paper