Score
Design and implement procedures, adapter layers, or transform pipelines that adapt pre-trained embedding representations to a target dataset or downstream predictor by updating embedding weights or learning lightweight modifications to improve downstream predictive accuracy. Build training, evaluation, and deployment practices that address limited-data regimes and that ensure deterministic, reproducible embedding extraction (fixed checkpoints, seed control, or deterministic transforms) for governance and downstream model reliability.
To address severe pipeline bubbles and limited throughput in heterogeneous large-model training, this paper proposes AdaPtis, an adaptive pipeline parallelism system. Methodologically, AdaPtis introduces (1) a generalizable pipeline performance model; (2) the first joint optimization of model partitioning, device placement, and micro-batch scheduling; and (3) a unified pipeline executor supporting diverse parallelism strategies. Experiments on representative heterogeneous hardware configurations demonstrate that AdaPtis achieves an average 1.42× speedup over Megatron-LM’s I-1F1B baseline, with peak improvements reaching 2.14×. These gains translate into significantly enhanced training efficiency and improved hardware resource utilization, without compromising model accuracy or training stability.
论文提出了一种保真度感知的训练框架和C-DPPO方法,解决了编码和终端代理后训练中的令牌与控制保真度错误问题。
Existing theoretical frameworks struggle to explain why larger-scale pre-trained models substantially reduce sample complexity on downstream tasks. This work proposes a novel theoretical framework—termed “caulking”—inspired by parameter-efficient fine-tuning methods such as adapters, low-rank adaptation, and partial fine-tuning. It establishes, for the first time, a provable relationship between the scale of pre-trained models and the sample complexity of downstream tasks. By rigorously linking stronger pre-training capabilities to reduced data requirements in transfer learning, this study not only addresses a critical gap in current theoretical understanding but also provides a solid foundation for empirically observed scaling laws, demonstrating that enhanced pre-training capacity can significantly decrease the number of samples needed for effective downstream adaptation.
本文提出UpgradeBench基准,通过四个连续模型版本、六项任务和两种模型大小,评估在升级过程中保留、迁移或重新训练特定任务适配器的有效性。
This study addresses the over-reliance of language agents on external frameworks for state tracking and control decision-making by proposing Harness Annealing Training, a method designed to internalize control capabilities. Introducing the novel concept of harness annealing, this approach integrates explicit control supervision with a progressive attenuation strategy. Through curriculum learning on teacher trajectories, it distills successful trajectories into the model’s intrinsic decision-making capacity while dynamically adjusting the allocation of control authority between human and machine. Experimental results demonstrate that 9B- and 35B-parameter models, when relying solely on tool use, achieve performance comparable to full framework deployment, thereby significantly reducing runtime control overhead.
This work addresses the fragmented landscape of post-training adaptation techniques, which suffer from inconsistent terminology and a lack of unified comparative or governance frameworks. To resolve this, the paper introduces the first six-dimensional taxonomy—spanning mechanism, objective, data requirements, persistence, structural scope, and model type—that systematically integrates mainstream approaches such as fine-tuning, retrieval augmentation, prompt engineering, model editing, and machine unlearning. This framework clarifies conceptual boundaries and reveals evolutionary and compositional relationships among methods. Beyond standardizing terminology, it enables standardized technical documentation, model change tracking, and AI governance analysis. The study further identifies critical challenges, including evaluation rigor, reproducibility, continual adaptation, multimodal alignment, and governance-aware workflows.
This work addresses the vulnerability of open-source large language models (LLMs) to malicious prompts—even on harmless inputs—caused by the degradation of safety alignment during downstream fine-tuning. To mitigate this, the authors propose SafeGene, a reusable safety adapter module that, for the first time, explicitly models safety alignment as a task-agnostic, standalone representation. SafeGene extracts safety vectors by contrasting aligned and degraded models, and integrates a data-aware layer selection strategy with few-shot, layer-wise coefficient calibration. Evaluated across multiple model families, downstream tasks, and safety benchmarks, SafeGene significantly reduces harmful response rates while preserving task performance, outperforming existing safety adaptation methods.
This study addresses the absence of supervision signals caused by the unavailability of specialized testing frameworks during deployment. To overcome this limitation, we propose a recursive self-rewriting framework built upon the Qwen-3.8-27B model. By orchestrating the collaboration of a planner, a critic, and an executor, our approach reconstructs problem-solving experiences from multiple frameworks into generalized training trajectories. Furthermore, reliability is ensured through runbook extraction, leakage screening, and sandboxed execution techniques. Experimental results demonstrate that the proposed method improves the task resolution rate by 34.3% and significantly outperforms baselines on the Pass@3 metric across benchmarks such as Terminal-Bench. These findings confirm that our framework effectively enables the reuse and generalization of cross-framework experiences in practical deployment scenarios.
This study addresses the challenges of missing source data and catastrophic forgetting during fine-tuning when transferring pretrained chip placement models across designs. We propose CARVE, a continual adaptation framework that enables secure knowledge reuse through frozen base policies, immutable expert modules, and task credential mechanisms. Furthermore, we introduce a novel local-validation-based expert selection strategy, establishing theoretical bounds for expected performance guarantees under fixed distributions and worst-case sample complexity with bounded loss. Experimental results demonstrate that our method reduces training time by 58.5% while preserving nearly lossless HPWL gains. Notably, on IBM circuits, CARVE achieves an average 5.76% HPWL improvement without requiring any receiver-side training.
This work proposes DataOrchestra, a novel framework that addresses the limitations of conventional pretraining data processing, which typically applies uniform strategies and fails to account for sample-level heterogeneity. DataOrchestra introduces the first automated, per-sample orchestration mechanism, featuring a hierarchical decision architecture: a high-level orchestrator dynamically selects among discarding, retaining, or combining multiple cleaning operations—such as programmatic editing and large language model–based rewriting—while low-level tool models execute the chosen operations accordingly. Evaluated across 11 benchmark tasks on models ranging from 0.5B to 7B parameters, this approach consistently outperforms any single static processing strategy, demonstrating particularly strong gains in mathematical continual pretraining while substantially reducing computational overhead.