Institution profile

Shanghai Qizhi Institute

Academic institutionasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Forking: Sudden Overfitting Under Replay

Sep 30, 2026

This study addresses the abrupt divergence between training and validation losses at epoch boundaries observed in NanoGPT during data replay. We reveal that this "bifurcation" phenomenon stems from an n-gram memory module that, through repeated updates, amplifies context-specific subspaces while suppressing the probabilities of unseen sequences, thereby triggering sudden overfitting. Through controlled experiments utilizing NanoGPT and DeepSeek-style Engram models, we systematically dissect the underlying n-gram encoding mechanisms and the influence of low-frequency contexts. Our contributions include successfully reproducing and confirming the prevalence of this bifurcation phenomenon across short-budget repetitive scenarios in both supervised fine-tuning and reinforcement learning. This work provides a novel perspective for understanding model generalization failure and highlights potential adverse side effects arising from techniques generated by autonomous AI research agents.

0 citationsRead paper

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Sep 30, 2026

This study addresses the high cost of acquiring end-to-end demonstration data and the difficulty of focusing supervision when optimizing long-horizon robotic skill compositions. To this end, it proposes a world-model-guided coaching framework centered on the RIDI loop. By leveraging an action-conditioned world model to simulate failure trajectories, the method precisely identifies weak subtasks and expert adapters, thereby guiding targeted demonstration requests and modular skill updates. This paradigm shifts the role of the world model from passive prediction to active coaching. Integrated with a progress judge and an aggregated record selection mechanism, the proposed approach improves task success rates from 13.3% to 75.0% on real robots using only limited data. The method significantly outperforms baselines while demonstrating strong generalization capabilities.

0 citationsRead paper

Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching

Sep 28, 2026

This study addresses the semantic misalignment between action representations and autoregressive backbones caused by existing action tokenizers, which limits robot learning. We propose CATok, which reformulates action tokenization as a causal generative process, achieving coarse-to-fine semantic alignment via conditional annealed flow matching. This approach constructs a causal token space that inherently decouples high-level reasoning from low-level execution without requiring explicit attention masks, while integrating a multimodal diffusion Transformer (MMDiT) with a discrete bottleneck architecture to enhance modeling capacity. Experiments demonstrate that CATok significantly improves reconstruction fidelity, inference efficiency, and Vision-Language-Action (VLA) success rates across both simulated and real-world tasks, establishing a foundation for high-performance autoregressive robot learning.

0 citationsRead paper
Recent publications

Latest Papers

Forking: Sudden Overfitting Under Replay

Sep 30, 2026

This study addresses the abrupt divergence between training and validation losses at epoch boundaries observed in NanoGPT during data replay. We reveal that this "bifurcation" phenomenon stems from an n-gram memory module that, through repeated updates, amplifies context-specific subspaces while suppressing the probabilities of unseen sequences, thereby triggering sudden overfitting. Through controlled experiments utilizing NanoGPT and DeepSeek-style Engram models, we systematically dissect the underlying n-gram encoding mechanisms and the influence of low-frequency contexts. Our contributions include successfully reproducing and confirming the prevalence of this bifurcation phenomenon across short-budget repetitive scenarios in both supervised fine-tuning and reinforcement learning. This work provides a novel perspective for understanding model generalization failure and highlights potential adverse side effects arising from techniques generated by autonomous AI research agents.

0 citationsRead paper

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Sep 30, 2026

This study addresses the high cost of acquiring end-to-end demonstration data and the difficulty of focusing supervision when optimizing long-horizon robotic skill compositions. To this end, it proposes a world-model-guided coaching framework centered on the RIDI loop. By leveraging an action-conditioned world model to simulate failure trajectories, the method precisely identifies weak subtasks and expert adapters, thereby guiding targeted demonstration requests and modular skill updates. This paradigm shifts the role of the world model from passive prediction to active coaching. Integrated with a progress judge and an aggregated record selection mechanism, the proposed approach improves task success rates from 13.3% to 75.0% on real robots using only limited data. The method significantly outperforms baselines while demonstrating strong generalization capabilities.

0 citationsRead paper

Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching

Sep 28, 2026

This study addresses the semantic misalignment between action representations and autoregressive backbones caused by existing action tokenizers, which limits robot learning. We propose CATok, which reformulates action tokenization as a causal generative process, achieving coarse-to-fine semantic alignment via conditional annealed flow matching. This approach constructs a causal token space that inherently decouples high-level reasoning from low-level execution without requiring explicit attention masks, while integrating a multimodal diffusion Transformer (MMDiT) with a discrete bottleneck architecture to enhance modeling capacity. Experiments demonstrate that CATok significantly improves reconstruction fidelity, inference efficiency, and Vision-Language-Action (VLA) success rates across both simulated and real-world tasks, establishing a foundation for high-performance autoregressive robot learning.

0 citationsRead paper