Score
Designs and implements algorithms and system components that adapt models and pipelines to individual users, including per-user personalization strategies, policy optimization for engagement, and methods to integrate personalization into end-to-end workflows. Builds lightweight on-device adaptation techniques (e.g., local tokenizer or small-module fine-tuning), local model updates and retrieval-augmented inference, and privacy-preserving/local-only personalization approaches that can keep core pretrained models frozen while adapting behavior to user signals.
This paper addresses the longstanding fragmentation between “personalized text generation” and “personalized downstream applications” (e.g., recommendation systems) in large language model (LLM) personalization research. To bridge this gap, we propose the first unified, multidimensional taxonomy and formal framework for personalized LLMs. Our framework systematically integrates diverse techniques—including parameter-efficient fine-tuning, prompt engineering, memory augmentation, user modeling, and preference alignment—across five dimensions: granularity, technical methodology, data paradigm, evaluation criteria, and application scenarios. We comprehensively survey existing benchmarks, metrics, and open challenges. Crucially, we formally define the novel paradigm, usage patterns, and ideal properties of personalized LLMs, and construct a full-stack, structured knowledge map. This work unifies disparate strands of research, resolves conceptual ambiguities, and establishes a rigorous theoretical foundation and practical roadmap for future investigation and deployment.
Large language models (LLMs) excel at general-knowledge tasks but struggle to capture user-specific characteristics—such as affective tendencies, writing style, and preferences. This paper presents a systematic survey of personalized large language models (PLLMs), proposing the first three-dimensional taxonomy spanning the input layer (e.g., context-aware prompt customization), model layer (e.g., parameter-efficient fine-tuning methods like LoRA and Adapter), and objective layer (e.g., human preference alignment techniques such as RLHF and DPO). Synthesizing over one hundred seminal works, we identify critical bottlenecks—including privacy-utility trade-offs and long-term memory modeling—and delineate six key frontiers: scalable personalization, safety-aligned adaptation, cross-domain transfer, continual learning, efficient inference, and interpretable personalization. Our framework provides a comprehensive foundation for both theoretical advancement and practical deployment of PLLMs.
This work addresses the challenge of efficiently constructing and managing massive numbers of persistent personalized models atop trillion-parameter foundation models. It proposes leveraging parameter-efficient fine-tuning (PEFT) as a lightweight and reliable personalization substrate, combining a shared large model with small, trainable adapters to encode user preferences, skills, and memory. The authors introduce MinT, an infrastructure that integrates adapter identity management, version control, provenance tracking, evaluation, and serving mechanisms, and define three scaling dimensions: Scale Up, Scale Down, and Scale Out. Experimental results demonstrate that, even under strong shared priors, compact adapters can stably capture personalized behaviors, offering a viable pathway toward large-scale deployment of millions of persistent personal models.
To address catastrophic forgetting caused by adapter interference in continual personalization of diffusion models—under the stringent constraint of inaccessible historical task adapters—this work proposes a three-stage strategy for multi-task sequential customization. First, orthogonal initialization decouples newly introduced adapters from prior knowledge subspaces. Second, selective weight freezing, guided by task relevance, suppresses updates to irrelevant parameters. Third, dynamic adapter sequence fusion enables adaptive integration of historical functionalities during inference. Built upon the LoRA architecture, the method requires neither storage nor reloading of past adapter parameters. Experiments demonstrate substantial mitigation of forgetting, with stable preservation of prior concept generation quality across multiple sequential personalization rounds. The implementation is publicly available.
Existing parameter-efficient fine-tuning (PEFT) methods require per-user adapter training, incurring high computational overhead and hindering real-time personalization. To address this, we propose Profile-to-PEFT: an end-to-end trainable hypernetwork that directly maps user profiles to full LoRA adapter parameters—enabling zero-shot, on-the-fly personalization without user-side training. Our framework supports localized deployment and privacy-preserving inference, while demonstrating strong generalization to unseen users and diverse behavioral patterns. Experiments show that Profile-to-PEFT outperforms prompt-based learning and single-user fine-tuning in personalization quality, maintaining robustness while drastically reducing deployment computational cost. This advances the efficiency, scalability, and practicality of large language model personalization at scale.
Current large language models (LLMs) suffer from insufficient personalization and heavy reliance on cloud infrastructure. To address these limitations, we propose a cloud-edge collaborative personalization framework: the cloud leverages LLMs to generate high-quality synthetic data and performs parameter-efficient fine-tuning (PEFT), while the edge integrates real and synthetic data for lightweight personalized training and standalone inference. This approach is the first to systematically synergize cloud-side generalization capability with edge-side personalization demands, alleviating user data sparsity through collaborative data augmentation. Evaluated across six downstream tasks, our method significantly improves personalization performance while eliminating network dependency—ensuring low-latency response and preserving local data privacy. Our core contribution is the establishment of the first cloud-edge joint fine-tuning paradigm that simultaneously achieves personalization, computational efficiency, and privacy preservation.
To address the challenge of fine-tuning large language models (LLMs) on memory- and compute-constrained edge devices, this paper proposes an efficient on-device fine-tuning method that requires no modifications to the inference engine. Our approach comprises three key contributions: (1) Parallelized Random Gradient Estimation (P-RGE), a low-overhead gradient approximation technique operating within a zeroth-order optimization framework; (2) a lightweight LoRA-FA module fully compatible with the ExecuTorch runtime, requiring no intrusive changes to the execution stack; and (3) the synergistic integration of LoRA-based parameter-efficient fine-tuning with P-RGE, achieving up to 68% reduction in GPU memory consumption and significantly lower computational overhead. Experiments demonstrate that our method maintains fine-tuning accuracy while accelerating training by 3.2×, enabling real-time, personalized LLM deployment on edge devices. This work provides a practical pathway for continual learning of LLMs in resource-constrained environments.
This work proposes an AI-driven approach to dynamic front-end personalization that overcomes the limitations of traditional static designs or rule-based systems, which often fail to accurately respond to user behavior. By leveraging a user behavior prediction model, the system dynamically adjusts interface layout and content in real time, while incorporating reinforcement learning to optimize the prioritization of functional elements. The architecture is designed to be scalable and adaptive, enabling efficient deployment and continuous iteration. Experimental results demonstrate that the proposed method significantly outperforms conventional rule-driven approaches in terms of user experience and interaction efficiency, thereby validating the effectiveness and practical value of integrating artificial intelligence into front-end personalization.
This study addresses the trade-off users face in large language model services between personalization methods—supervised fine-tuning (SFT) and in-context learning (ICL)—and congestion due to shared computational resources. The authors develop an analytically tractable framework integrating statistical and economic factors, combining game-theoretic equilibrium analysis, theoretical modeling, and experiments on GPT-2, complemented by an empirical survey of 21 leading AI platforms. Their findings reveal how pretraining coverage, signal-to-noise ratio, and system congestion jointly determine the relative performance of SFT versus ICL. The work demonstrates that offering both personalization strategies simultaneously maximizes platform profit—a prediction corroborated by real-world adoption, as the share of platforms supporting dual-mode personalization rose from 9.5% in 2021 to 71.4% in 2025, underscoring the practical relevance and foresight of the proposed design.
This work addresses the limitations of traditional recommender systems, which rely on platform-owned implicit user profiles that lack transparency, editability, and cross-platform portability, thereby limiting users’ understanding and control over their personalized experiences. The paper introduces a novel paradigm—“governable personalization”—and presents the first systematic architecture centered on large language model (LLM) agents for recommendation. By integrating cross-domain memory, privacy-preserving modeling, intent alignment, and trustworthy monetization mechanisms, the proposed framework empowers users to inspect, revise, transfer, and exert influence over their profiles across services. This study establishes foundational design principles and a research agenda for recommender systems in the LLM era, offering both theoretical grounding and technical pathways toward user-driven personalization.
This work addresses the challenge in federated learning where task heterogeneity makes it difficult to balance model generalization and personalization. To this end, the authors propose Potara, a novel framework that provides the first theoretical foundation for personalized model fusion in federated settings. Leveraging linear mode connectivity theory, Potara derives a closed-form solution for optimal mixing weights that efficiently combine a global model with locally fine-tuned models—such as those adapted via LoRA—yielding a merged model whose loss upper bound is provably tighter than either constituent model alone. Extensive experiments demonstrate that Potara significantly enhances personalized performance across vision and language benchmarks while substantially reducing communication overhead, achieving an excellent trade-off between accuracy and communication efficiency.