layer-wise memory training

Designs and implements training procedures that optimize memory-related components one layer or stage at a time, tuning submodules sequentially to assemble an integrated, retrieval-capable memory system. Uses layer-wise or stage-wise optimization to stabilize learning of multi-component memory architectures and to enable inspection and analysis of each submodule’s behavior.

layer-wisememorytraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.06
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Don't Half-listen: Capturing Key-part Information in Continual Instruction Tuning

Mar 15, 2024
YH
Yongquan He
🏛️ Meituan | Institute of Information Engineering, Chinese Academy of Sciences | National University of Defense Technology

To address catastrophic forgetting (CF) and instruction overfitting in continual instruction tuning (CIT) of large language models (LLMs)—which degrade task generalization and instruction-following capability—this paper proposes a task-aware dynamic tuning framework. Methodologically, it (1) identifies semantically decisive instruction fragments via Key Part Information Gain (KPIG), enabling fine-grained task awareness; (2) introduces a dynamic data replay and target refinement mechanism that preserves task essence rather than superficial patterns; and (3) establishes the first dual-metric evaluation system—P-score (measuring generalization) and V-score (assessing instruction adherence). Experiments demonstrate that our approach significantly outperforms existing baselines on both seen and unseen tasks, effectively mitigating CF and overfitting while improving instruction-following accuracy and cross-task generalization performance.

Address catastrophic forgetting in continual instruction tuningImprove task-aware information capture in LLMsMeasure generalization and instruction-following abilities

This work addresses the significant yet underexplored impact of parameter configuration on performance in memory tiering systems, where efficient automated tuning mechanisms are lacking. The authors propose PTMT, a lightweight framework that systematically categorizes memory tiering parameters and reveals their performance sensitivity for the first time. PTMT introduces a hybrid adaptive tuning mechanism that combines offline performance profiling with online reinforcement learning to achieve high-efficacy optimization at low overhead. By co-designing memory access profiling and page migration, PTMT improves performance by 30%, 26%, 21%, and 14% over TPP, UPM, Colloid, and AutoNUMA, respectively, and outperforms the best existing approaches by 32% on average.

configuration sensitivitymemory tieringparameter tuning

STEP: Staged Parameter-Efficient Pre-training for Large Language Models

Apr 05, 2025
KY
Kazuki Yano
🏛️ Tohoku University | Langsmith Inc. | RIKEN | NII LLMC

To address the prominent GPU memory bottleneck in large language model (LLM) pretraining, this paper proposes the Staged Parameter-Efficient Training (SPET) framework. SPET is the first to deeply integrate parameter-efficient fine-tuning techniques—such as LoRA—into the *entire* pretraining pipeline, synergistically combining gradient checkpointing with staged architectural expansion to enable dynamic model growth and on-demand memory optimization. Implemented in PyTorch, SPET introduces a memory-aware training scheduler that reduces peak GPU memory consumption by up to 53.9% versus full-parameter baselines, while preserving pretraining performance. Downstream task performance after instruction tuning remains unchanged. The core contribution lies in bridging the paradigmatic divide between standard pretraining and parameter-efficient adaptation, establishing a scalable, memory-efficient, and unified pretraining paradigm.

Integrates efficient tuning with model growthMaintains performance with less memoryReduces memory use in LLM pre-training

Mitigating Forgetting in LLM Supervised Fine-Tuning and Preference Learning

Oct 20, 2024
HF
Heshan Fernando
🏛️ Rensselaer Polytechnic Institute | IBM Research

Sequential supervised fine-tuning (SFT) followed by preference learning (e.g., DPO or RLHF) in large language model post-training induces catastrophic forgetting, degrading SFT task performance while optimizing for preference alignment. Method: This work theoretically establishes the suboptimality of sequential training and proposes the first jointly optimized framework with provable convergence guarantees. It unifies SFT and DPO objectives via a scalable multi-objective loss function and performs joint gradient updates—without additional computational overhead. Contribution/Results: The method enables effective knowledge fusion across both stages. Experiments demonstrate a 23% improvement in SFT task retention and a 9.7% gain in preference alignment accuracy over sequential baselines, while maintaining comparable computational cost.

Mitigate forgetting in LLM post-trainingOptimize SFT and RLHF/DPO trade-offPropose joint post-training framework

The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning

Jun 17, 2024
BH
Bingxiang He
🏛️ Tsinghua University | University of Illinois Urbana-Champaign | Renmin University of China

This study investigates the intrinsic mechanisms underlying zero-shot generalization in instruction tuning, revealing that it fundamentally arises from instance-level input similarity—termed “early generalization”—and is highly sensitive to training data ordering. To address this, we propose the first test-centric multi-turn arrangement framework, which integrates dynamic loss analysis, instance similarity modeling, granularity-aware sorting, and progressive training scheduling. Our method significantly enhances zero-shot generalization on unseen tasks, accelerates convergence, reduces training loss, and improves generalization stability. The core contribution is the first empirical identification of instance-level similarity—not task-level structural alignment—as the primary driver of zero-shot generalization; further, we establish data ordering as a novel, controllable lever for optimizing generalization behavior, enabling precise, target-driven adaptation without architectural or objective modifications.

Explores instance-level similarity between training and test data for generalizationInvestigates how data arrangement affects zero-shot generalization in instruction tuningProposes a test-centric framework to improve continual learning and loss reduction

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing language models, whose memory mechanisms typically rely on static storage and struggle to effectively manage information retention, forgetting, and user trust in memory reliability. To overcome these challenges, the authors propose StageMem, a novel framework that models memory as a dynamic process with a well-defined lifecycle, comprising transient, working, and persistent stages. Each memory item is explicitly assigned a confidence score and strength, enabling fine-grained control. Through stage-specific policies for admission, promotion, updating, and eviction, StageMem decouples shallow writing from long-term commitment, substantially enhancing the flexibility and reliability of memory management. Experimental results demonstrate that StageMem effectively preserves critical late-stage information under controlled stress, reduces contamination and overall load in deep memory, and remains compatible with sophisticated retrieval systems.

information retentionlanguage modelsmemory lifecycle

Existing large language model compression methods are constrained by whole-layer granularity and continuous selection, making it difficult to precisely align with the non-uniform redundancy inherent in Transformer architectures. This work proposes SubFit, the first approach to enable discontinuous, differentiated compression at the sub-module level for both attention and feed-forward networks, complemented by lightweight residual fitting bypasses tailored to each sub-module type. Operating within a post-training framework, SubFit efficiently adapts using only calibration data. Evaluated across ten large language models, SubFit consistently achieves the best trade-off between perplexity and downstream accuracy under various sparsity levels: at 25% sparsity, it retains 84.6% of downstream task accuracy with only a 2.42× increase in perplexity, while significantly accelerating inference and reducing KV cache overhead.

granularityLLM compressionpost-training compression

This work addresses the challenge that current large language models struggle to effectively accumulate and reuse transferable experience across problems during inference, as existing memory mechanisms either exhibit limited generalization or are not optimized for answer correctness. To overcome this, the paper proposes the MILES framework, which dynamically constructs asymmetric modular memory units composed of subgoal embeddings and sub-instructions, and introduces a learnable memory selection head to enable a coarse-to-fine two-stage retrieval and composition mechanism. Designed for incremental test-time scenarios, MILES leverages confidence-based supervision signals and a fine-grained reranking strategy, achieving strong performance across multiple tasks—either significantly outperforming or matching state-of-the-art methods—while striking a favorable balance among accuracy, efficiency, robustness, and transferability.

learnable selectionmemory-based methodsmodular memory

This work addresses the challenge of effectively preserving and leveraging memory states across context boundaries. It proposes a learnable memory consolidation mechanism implemented via a lightweight Consolidator module, which dynamically transforms and accumulates routed short-term memory into long-term memory without erasing existing long-term content. The consolidated long-term memory serves dual roles: as a retrievable knowledge store and as a guiding signal for selecting subsequent memory slots, thereby significantly enhancing cross-context memory retention. Built upon the Phasor memory network architecture and incorporating KV cache eviction with hierarchical routing, the approach trains only the Consolidator—comprising 12.35K parameters (0.041% of the total model)—while keeping the backbone frozen. On a two-phase modulo-10 mapping task, this method improves recall accuracy for updated mappings from 44.38% to 87.02%.

context boundarylong-term memorymemory consolidation

Hot Scholars

YD

Yatin Dandi

EPFL, IIT Kanpur
Deep Learning TheoryStatistical PhysicsOptimization
FK

Florent Krzakala

École polytechnique fédérale de Lausanne
Statistical MechanicsStatisticsMachine LearningInformation theory
MP

Miao Pan

Professor, Electrical and Computer Engineering, University of Houston
Wireless for AICybersecurity for AIMobile/Edge AI SystemsUnderwater IoT Nets
KC

Koki Chinzei

Fujitsu Research
Quantum computingQuantum machine learningQuantum many-body physics