Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the incompatibility between conventional weight tying strategies in decoder-based large language models and efficient ghost clipping techniques during differentially private (DP) training. Focusing on Transformer architectures such as GPT-2, this work evaluates architectural choices under DP-SGD and proposes an embedding untieing strategy to accommodate the ghost clipping algorithm. Experimental results demonstrate that untied embeddings significantly outperform weight tying in privacy-preserving settings, improving model accuracy by 4.74% while reducing memory consumption by over 60%. This research elucidates the failure mechanism of traditional weight tying in DP training and confirms that untied embeddings constitute a superior paradigm for privacy-preserving training, effectively balancing predictive performance with memory efficiency.
📝 Abstract
Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and improved language modeling performance in the non-private setting. However, the impact of weight tying under differentially private training remains largely unexplored. In this work, we investigate the role of weight tying in the DP setting using GPT2 and DistilGPT2 as representative decoder-only architectures. Interestingly, we find that untied embeddings consistently outperform weight-tied models under DP-SGD, achieving gains of up to 4.74% points in accuracy on SST-2, QNLI, and QQP. Beyond improved utility, untying embeddings enables the use of memory-efficient ghost clipping for DP-SGD. By contrast, weight tying introduces shared-parameter interactions that complicate standard ghost norm computation and largely negate its computational advantages. As a result, untied models achieve over 60% lower memory usage while preserving the benefits of ghost clipping. Our results indicate that untied embeddings provide a more effective and scalable design for differentially private training of decoder-only LLMs and highlight the need to revisit standard LLM architectural choices in the privacy-preserving setting.
Problem

Research questions and friction points this paper is trying to address.

Weight Tying
Differential Privacy
DP-SGD
Decoder-Only LLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Differential Privacy
Weight Tying
DP-SGD
Ghost Clipping
Decoder-Only LLMs
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Razan El Mais
Department of Electrical and Computer Engineering, American University of Beirut, Beirut, Lebanon
Ali Chehab
Ali Chehab
Professor & Chair of ECE Department, American University of Beirut
CryptographyAI for CybersecurityAI for Medicine
Ibrahim Issa
Ibrahim Issa
Assistant Professor, American University of Beirut
Privacy and SecurityInformation TheoryStatistical LearningQuantum Information Theory
R
Razane Tajeddine
Department of Electrical and Computer Engineering, American University of Beirut, Beirut, Lebanon