Score
Design and build methods and pipelines that automatically construct training tasks and their supervisory signals for task-driven learning, including creating synthetic tasks, sampling diverse meta-training tasks, and producing pseudo- or soft-label assignments from models or unlabeled data. Analyze and evaluate the quality, diversity, and utility of these generated tasks and labels for downstream adaptation and meta-learning.
This work addresses the high cost, poor scalability, and diminishing effectiveness of human-supervised approaches for improving large language models, especially as model capabilities approach human-level performance. To overcome these limitations, the paper proposes a closed-loop self-improvement framework that structures the self-enhancement process into four tightly coupled stages: data acquisition, selection, model optimization, and inference refinement. A key innovation is the introduction of an autonomous evaluation layer that coordinates and guides transitions across these stages. This framework offers the first systematic, lifecycle-oriented modeling of self-improvement, unifying critical components such as self-generated data, automated evaluation, iterative training, and inference-time optimization. By comprehensively mapping existing technical pathways and their limitations, the study lays the groundwork for realizing fully autonomous, self-evolving language models.
Existing AI research agents often produce seemingly plausible but ineffective machine learning solutions due to a lack of systematic training. To address this, this work proposes the first scalable synthetic task generation framework that automatically constructs high-quality, executable research tasks through topic sampling, proposal generation grounded in real-world Hugging Face datasets, and self-debugging validation. The framework further leverages trajectory distillation—transferring effective research behaviors from GPT-5 to Qwen3—to guide student models in learning valid scientific reasoning paths. Evaluated on the MLGym benchmark, Qwen3-4B and Qwen3-8B models trained with this approach achieve 9% and 12% relative improvements in Area Under the Performance curve (AUP), respectively, substantially outperforming baseline methods.
In the era of large language models, selecting and fine-tuning pre-trained models remains challenging due to the absence of efficient adaptation mechanisms across vast model zoos and severe limitations in labeled data, hindering few-shot generalization. Method: (1) We systematically integrate meta-learning into the deep learning pipeline, constructing a task-prior-driven pipeline ranking surrogate model; (2) we quantitatively characterize the critical role of data augmentation in self-supervised learning; (3) we propose a differentiable neural synthetic data generator, replacing conventional reinforcement learning–based approaches. Contribution/Results: Our framework significantly outperforms human-crafted fine-tuning baselines on standard CV and NLP benchmarks. It achieves substantial gains in few-shot settings and enables zero-shot cross-environment generalization of the synthetic data generator, demonstrating robust adaptability without domain-specific retraining.
This work addresses the limitation of existing large language models in synthetic data generation, which typically treat tasks as isolated events and thus fail to accumulate or transfer synthesis experience across tasks. To overcome this, the authors propose StreamSynth, a novel paradigm that formulates synthetic data generation as an experience-driven continual learning process. By incorporating streaming task inputs and a feedback mechanism, StreamSynth enables the model to continuously learn from and reuse effective synthesis strategies across a sequence of tasks. The proposed SynLearner framework integrates diverse exploration, feedback-based learning, and a balanced optimization of quality and diversity. Experimental results demonstrate that the approach effectively leverages early-task experience to enhance performance on subsequent tasks, exhibiting robust cross-task transfer and cumulative learning capabilities across multiple benchmarks.
Instruction fine-tuning (IFT) suffers from heavy reliance on large-scale annotated examples and poor few-shot cross-task generalization. To address this, we propose an instruction-driven zero-shot adapter generation framework. Our method introduces three key innovations: (1) the first end-to-end paradigm mapping natural-language instructions directly to adapter parameters; (2) a two-stage hypernetwork training scheme that decouples instruction understanding from parameter generation; and (3) the first integration of knowledge distillation into instruction learning to align instruction-level and instance-level training signals. Evaluated on Super-Natural Instructions and P3 benchmarks, our approach matches or surpasses state-of-the-art meta-trained and hypernetwork-based models in task performance, while significantly reducing inference computational overhead. This work establishes a new paradigm for efficient, low-resource generalization of large language models.
This paper addresses the challenges of structural design and continual optimization for scaffolded language models (LMs) in multi-step tasks. We propose a novel paradigm—*language-supervised training*—that enables non-parametric optimization via natural-language instructions, tool-call trajectories, and human-readable/editable linguistic feedback. This framework dynamically adapts external variables—including prompts, toolchains, and scaffolding code—without modifying model parameters. It is compatible with closed-source LMs, mitigates catastrophic forgetting, and supports human-in-the-loop streaming learning. We introduce the first taxonomy of non-parametric variables tailored to language supervision, unifying prompt engineering, multi-step reasoning orchestration, and feedback modeling. Our approach provides both theoretical foundations and a systematic implementation pathway for deploying hybrid autonomous agents in real-world settings.
Existing query-first data synthesis approaches struggle to generate valid and executable tool-use sequences. This work proposes SyntheticAgentTraceQA, a novel framework that introduces an "execution-first" paradigm: it first constructs high-level workflows, maps and validates feasible tool trajectories, and then synthesizes corresponding natural language tasks and reference answers. The method integrates dependency-aware tool assignment, trajectory validation in a controlled environment, and reasoning-augmented annotation generation, followed by fine-tuning and evaluation using the Qwen model. Experimental results demonstrate that this framework substantially improves large language model (LLM) agents’ tool execution accuracy, trajectory consistency, and answer quality. Furthermore, the study reveals that masked supervision outperforms full supervision for models at the 9B scale.
This study addresses the high cost of manually constructing task configurations for scientific information extraction and the difficulty of automatically generating complete pipelines from brief objectives alone. To overcome these challenges, this work proposes an end-to-end framework that formulates task construction as a series of optimizable components. Methodologically, schemas, instructions, and evaluation criteria are automatically generated from weak specifications, while textual gradient feedback and failure-focused update mechanisms are introduced to enable joint optimization, ensuring the entire pipeline remains fully editable. Experiments conducted on a heterogeneous corpus of catalysis literature demonstrate that the joint optimization of schemas and instructions consistently yields superior performance across all settings, a finding further corroborated by blind human evaluations.
为解决工具使用语言模型训练数据昂贵问题,提出技能预训练(SPT)方法,利用公开技能包作为训练数据,提高模型性能。
研究提出任务模型归纳法,从自然计算机使用痕迹中提取多线程工作流程,解决现有方法仅能处理单一工作流的问题。
This study addresses the high cost of expert data and the lack of factual grounding and internal consistency in synthetic tasks for training LLM agents. We propose an evidence-based method for occupational scenario construction and execution-guided consistency verification. By retrieving authentic occupational knowledge from O*NET to synthesize contexts and render reference deliverables, combined with execution feedback for defect attribution and iterative repair, our approach automatically generates high-quality, executable training data. Fine-tuned on merely 20,000 samples, the resulting Fx-Work-35B model surpasses peer models across multiple benchmarks and even outperforms frontier large language models on certain metrics. This work demonstrates a low-cost, high-fidelity paradigm for the automated generation of agent training data.