Noisy Test-Time Reinforcement Learning for Code LLMs

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited robustness of large code models under real-world noisy instructions and the prohibitive cost of acquiring paired data. To this end, it proposes NTRL-Code, a framework enabling test-time noise-robust self-evolution without requiring paired supervision. Methodologically, NTRL-Code employs conservative self-denoising to establish semantic anchors, leverages abstract syntax tree (AST) structural aggregation to estimate surrogate objectives, and introduces a multi-signal hybrid reward strategy to drive reinforcement learning optimization. This paradigm uniquely endows models with autonomous noise resilience during the inference phase. Experiments across three benchmarks featuring multi-level perturbations demonstrate that NTRL-Code significantly enhances model robustness, effectively stabilizes predictive outputs, and improves overall performance.
📝 Abstract
Large language models (LLMs) have demonstrated remarkable performance across various code-related tasks. However, unlike carefully curated datasets that are typically high-quality and error-free, real-world user instructions are often vague and error-prone, posing significant challenges to the robustness of code LLMs. Furthermore, robustness-oriented fine-tuning relies on paired clean-noisy samples, which are costly to curate and require sophisticated noisy simulation techniques. To address these challenges, we propose the Noisy Test-time Reinforcement Learning framework (NTRL-Code), which enables robust self-evolution of code LLMs using only unlabeled noisy data during the testing stage. Specifically, NTRL-Code uses conservative self-denoising to obtain a cleaner semantic anchor for target estimation, and employs an abstract-syntax-tree (AST)-based structural aggregation mechanism to estimate a proxy target from multiple candidate programs. The policy is then optimized on the original noisy prompts with a hybrid reward that combines format validity, code similarity, and anti-repetition signals. Extensive experiments on three benchmarks, each incorporating character-level, word-level, and paragraph-level perturbations, demonstrate that NTRL-Code yields robust and consistent improvements, stabilizing the predictions of various base models. Our code is available at https://github.com/Xikai97/NTRL-Code.
Problem

Research questions and friction points this paper is trying to address.

Code LLMs
Robustness
Noisy Instructions
Test-Time Reinforcement Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-time Reinforcement Learning
Noisy Prompts
Self-denoising
AST-based Aggregation
Code LLMs
🔎 Similar Papers
No similar papers found.