Score
Designs and implements training procedures that optimize memory-related components one layer or stage at a time, tuning submodules sequentially to assemble an integrated, retrieval-capable memory system. Uses layer-wise or stage-wise optimization to stabilize learning of multi-component memory architectures and to enable inspection and analysis of each submodule’s behavior.
To address catastrophic forgetting (CF) and instruction overfitting in continual instruction tuning (CIT) of large language models (LLMs)—which degrade task generalization and instruction-following capability—this paper proposes a task-aware dynamic tuning framework. Methodologically, it (1) identifies semantically decisive instruction fragments via Key Part Information Gain (KPIG), enabling fine-grained task awareness; (2) introduces a dynamic data replay and target refinement mechanism that preserves task essence rather than superficial patterns; and (3) establishes the first dual-metric evaluation system—P-score (measuring generalization) and V-score (assessing instruction adherence). Experiments demonstrate that our approach significantly outperforms existing baselines on both seen and unseen tasks, effectively mitigating CF and overfitting while improving instruction-following accuracy and cross-task generalization performance.
This work addresses the significant yet underexplored impact of parameter configuration on performance in memory tiering systems, where efficient automated tuning mechanisms are lacking. The authors propose PTMT, a lightweight framework that systematically categorizes memory tiering parameters and reveals their performance sensitivity for the first time. PTMT introduces a hybrid adaptive tuning mechanism that combines offline performance profiling with online reinforcement learning to achieve high-efficacy optimization at low overhead. By co-designing memory access profiling and page migration, PTMT improves performance by 30%, 26%, 21%, and 14% over TPP, UPM, Colloid, and AutoNUMA, respectively, and outperforms the best existing approaches by 32% on average.
To address the prominent GPU memory bottleneck in large language model (LLM) pretraining, this paper proposes the Staged Parameter-Efficient Training (SPET) framework. SPET is the first to deeply integrate parameter-efficient fine-tuning techniques—such as LoRA—into the *entire* pretraining pipeline, synergistically combining gradient checkpointing with staged architectural expansion to enable dynamic model growth and on-demand memory optimization. Implemented in PyTorch, SPET introduces a memory-aware training scheduler that reduces peak GPU memory consumption by up to 53.9% versus full-parameter baselines, while preserving pretraining performance. Downstream task performance after instruction tuning remains unchanged. The core contribution lies in bridging the paradigmatic divide between standard pretraining and parameter-efficient adaptation, establishing a scalable, memory-efficient, and unified pretraining paradigm.
Sequential supervised fine-tuning (SFT) followed by preference learning (e.g., DPO or RLHF) in large language model post-training induces catastrophic forgetting, degrading SFT task performance while optimizing for preference alignment. Method: This work theoretically establishes the suboptimality of sequential training and proposes the first jointly optimized framework with provable convergence guarantees. It unifies SFT and DPO objectives via a scalable multi-objective loss function and performs joint gradient updates—without additional computational overhead. Contribution/Results: The method enables effective knowledge fusion across both stages. Experiments demonstrate a 23% improvement in SFT task retention and a 9.7% gain in preference alignment accuracy over sequential baselines, while maintaining comparable computational cost.
This study investigates the intrinsic mechanisms underlying zero-shot generalization in instruction tuning, revealing that it fundamentally arises from instance-level input similarity—termed “early generalization”—and is highly sensitive to training data ordering. To address this, we propose the first test-centric multi-turn arrangement framework, which integrates dynamic loss analysis, instance similarity modeling, granularity-aware sorting, and progressive training scheduling. Our method significantly enhances zero-shot generalization on unseen tasks, accelerates convergence, reduces training loss, and improves generalization stability. The core contribution is the first empirical identification of instance-level similarity—not task-level structural alignment—as the primary driver of zero-shot generalization; further, we establish data ordering as a novel, controllable lever for optimizing generalization behavior, enabling precise, target-driven adaptation without architectural or objective modifications.
This work addresses the limitations of existing language models, whose memory mechanisms typically rely on static storage and struggle to effectively manage information retention, forgetting, and user trust in memory reliability. To overcome these challenges, the authors propose StageMem, a novel framework that models memory as a dynamic process with a well-defined lifecycle, comprising transient, working, and persistent stages. Each memory item is explicitly assigned a confidence score and strength, enabling fine-grained control. Through stage-specific policies for admission, promotion, updating, and eviction, StageMem decouples shallow writing from long-term commitment, substantially enhancing the flexibility and reliability of memory management. Experimental results demonstrate that StageMem effectively preserves critical late-stage information under controlled stress, reduces contamination and overall load in deep memory, and remains compatible with sophisticated retrieval systems.
Existing large language model compression methods are constrained by whole-layer granularity and continuous selection, making it difficult to precisely align with the non-uniform redundancy inherent in Transformer architectures. This work proposes SubFit, the first approach to enable discontinuous, differentiated compression at the sub-module level for both attention and feed-forward networks, complemented by lightweight residual fitting bypasses tailored to each sub-module type. Operating within a post-training framework, SubFit efficiently adapts using only calibration data. Evaluated across ten large language models, SubFit consistently achieves the best trade-off between perplexity and downstream accuracy under various sparsity levels: at 25% sparsity, it retains 84.6% of downstream task accuracy with only a 2.42× increase in perplexity, while significantly accelerating inference and reducing KV cache overhead.
This work addresses the challenge that current large language models struggle to effectively accumulate and reuse transferable experience across problems during inference, as existing memory mechanisms either exhibit limited generalization or are not optimized for answer correctness. To overcome this, the paper proposes the MILES framework, which dynamically constructs asymmetric modular memory units composed of subgoal embeddings and sub-instructions, and introduces a learnable memory selection head to enable a coarse-to-fine two-stage retrieval and composition mechanism. Designed for incremental test-time scenarios, MILES leverages confidence-based supervision signals and a fine-grained reranking strategy, achieving strong performance across multiple tasks—either significantly outperforming or matching state-of-the-art methods—while striking a favorable balance among accuracy, efficiency, robustness, and transferability.
This work addresses the challenge of effectively preserving and leveraging memory states across context boundaries. It proposes a learnable memory consolidation mechanism implemented via a lightweight Consolidator module, which dynamically transforms and accumulates routed short-term memory into long-term memory without erasing existing long-term content. The consolidated long-term memory serves dual roles: as a retrievable knowledge store and as a guiding signal for selecting subsequent memory slots, thereby significantly enhancing cross-context memory retention. Built upon the Phasor memory network architecture and incorporating KV cache eviction with hierarchical routing, the approach trains only the Consolidator—comprising 12.35K parameters (0.041% of the total model)—while keeping the backbone frozen. On a two-phase modulo-10 mapping task, this method improves recall accuracy for updated mappings from 44.38% to 87.02%.