Score
Designs and implements training pipelines that generate candidate solutions from a base model, automatically construct weakly labeled positive and negative critic examples from model trajectories, train a learned critic without human annotations to score candidates, and iteratively select high‑scoring outputs to fine‑tune the model.
This work addresses the high cost, poor scalability, and diminishing effectiveness of human-supervised approaches for improving large language models, especially as model capabilities approach human-level performance. To overcome these limitations, the paper proposes a closed-loop self-improvement framework that structures the self-enhancement process into four tightly coupled stages: data acquisition, selection, model optimization, and inference refinement. A key innovation is the introduction of an autonomous evaluation layer that coordinates and guides transitions across these stages. This framework offers the first systematic, lifecycle-oriented modeling of self-improvement, unifying critical components such as self-generated data, automated evaluation, iterative training, and inference-time optimization. By comprehensively mapping existing technical pathways and their limitations, the study lays the groundwork for realizing fully autonomous, self-evolving language models.
Existing LLM-driven automated visual modeling approaches rely on global, one-shot optimization, resulting in poor attribution, slow convergence, low stability, and limited accessibility for non-experts. Method: We propose an “iterative single-component fine-tuning” strategy, inspired by expert human practice, wherein only one module in the pipeline is optimized per iteration. This is integrated with training-feedback-guided modular updates, zero-shot prompt engineering, and a multi-domain evaluation protocol to construct an end-to-end LLM agent framework. Contribution/Results: Our approach significantly enhances interpretability, stability, and convergence efficiency of optimization. Evaluated across multiple standard benchmarks and Kaggle datasets, it consistently outperforms state-of-the-art zero-shot LLM methods, achieving superior classification accuracy and generalization capability.
Traditional neural network training relies on fixed optimization pipelines, rendering it inflexible in dynamically addressing training instability and anomalies. To address this limitation, we propose the first interactive training framework enabling real-time human–AI collaborative intervention. Our method employs a lightweight control server that integrates expert human directives with AI agent feedback to dynamically adjust hyperparameters, data sampling strategies, and model checkpoints during training. This framework introduces, for the first time, a closed-loop interactive paradigm into neural network training, establishing a scalable human–machine collaboration interface coupled with automated response mechanisms. Experimental results demonstrate significant improvements in training stability, reduced sensitivity to initial hyperparameter configurations, and enhanced real-time responsiveness to user-specified customization requirements. The effectiveness is validated across multiple benchmark tasks.
This work addresses black-box program behavior modeling by proposing a reversible, differentiable, and constraint-aware, grammar-driven neural modeling framework. Methodologically, it generates input-output (I/O) pairs from formal grammars of input and output languages, and employs a lightweight (<6.3M-parameter) sequence-to-sequence model to cast program I/O mapping as a bidirectional neural machine translation task—enabling both forward prediction and backward inference—while supporting fine-grained behavioral constraints and fault- or coverage-guided input synthesis. Its key contribution is the first end-to-end joint modeling of reversibility, differentiability, and syntactic consistency in program behavior models. Evaluated on structured tasks such as Markdown and HTML generation, the framework achieves 95.4% accuracy and a BLEU score of 0.98±0.04, significantly outperforming existing irreversible or syntax-agnostic approaches.
This work addresses the challenge that large language model–based code agents struggle to efficiently acquire strategic reasoning capabilities through end-to-end training, a process that is both ineffective and computationally expensive. To overcome this limitation, the authors propose freezing the primary agent and introducing a lightweight critic model that delivers real-time, fine-grained supervisory feedback during trajectory execution, thereby guiding the agent to refine its decision-making rather than directly producing final answers. The critic is trained via supervised fine-tuning and demonstrates strong cross-model transferability—evidenced by successful deployment with both CWM-32B and Qwen-family models. Evaluated on SWE-bench Verified, the approach improves accuracy by 3.0–5.2 percentage points over baseline methods, achieving 25.2% accuracy compared to Qwen3-Next-80B-A3B alone while reducing inference cost to $0.04.
This work addresses fundamental bottlenecks in current AI systems—namely, inefficient knowledge acquisition, heavy reliance on human-annotated data, and rigid, manually designed training paradigms. To overcome these limitations, the authors propose a synthetic data–driven self-improvement framework that enhances small-scale corpora with synthetically generated data to accelerate knowledge updating. The approach employs distillation-free self-guided pretraining, replacing human-labeled data with model-generated content, and leverages algorithmic space search at test time to automatically discover learning strategies superior to handcrafted ones. This methodology substantially improves knowledge acquisition under data-scarce conditions, reduces dependence on human-provided data, and expands the frontier of autonomous learning through algorithmic innovation.
This work systematically investigates three critical yet often overlooked design factors that profoundly influence the effectiveness of iterative generative optimization with large language models: the choice of initial artifacts, the scope of credit assignment in execution trajectories, and the batching strategy for trial-and-error samples. Through extensive experiments across diverse benchmarks—including MLAgentBench, Atari, and BigBench Hard—combined with execution feedback and iterative editing mechanisms, the study empirically demonstrates that these “hidden” choices decisively determine optimization success or failure. Specifically, different initial artifacts significantly affect the reachability of the solution space, truncated trajectories can still enhance performance on Atari tasks, and increasing batch size does not necessarily improve generalization. These findings provide both theoretical grounding and practical guidance for building robust iterative self-improvement systems.
Existing LLM tool-use methods rely on static data pipelines, decoupling data generation from model training—hindering adaptive focus on model weaknesses and effective removal of noisy labels, thus impairing training efficiency. This paper introduces the first open-source, model-aware data evolution framework, establishing a closed-loop training paradigm comprising three tightly integrated modules: *capability diagnosis*, *label verification*, and *error-driven expansion*. It jointly optimizes data and model through iterative refinement: greedy capability probing identifies model deficiencies; discriminator-guided label verification purifies training data; and error feedback steers targeted data augmentation. The resulting 8B model achieves state-of-the-art performance on BFCL-v3 and ACEBench—surpassing same-scale SOTA models and even outperforming its 32B data generator—marking the first demonstration of data–model co-evolution within an open-source ecosystem.
This study investigates the evolutionary dynamics of generative models trained iteratively on synthetic data contaminated with real data, aiming to mitigate model collapse induced by data pollution. Through statistical modeling, mixture distribution analysis, and theoretical analysis of iterative training dynamics—complemented by theoretical derivations and simulations based on next-token prediction language models—the work demonstrates that model collapse can be effectively avoided and the true data distribution even recovered, provided the mixture weight of real data remains non-zero over time and is paired with sufficient sample sizes. This mechanism consistently enhances performance across diverse model classes, offering both theoretical guarantees and practical guidance for sustainable iterative training.