personalization-generalization analysis

Designs and executes evaluation protocols, metrics, and analyses that quantify both per-client personalized performance and aggregate/global model generalization, and that compare personalized approaches to global or centralized models. Builds experiments and diagnostic studies to measure the personalization-vs-generalization tradeoff, report per-client and overall accuracy, and identify factors or conditions that degrade one objective relative to the other.

personalization-generalizationanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address the client drift and generalization imbalance in federated learning under non-independent and identically distributed (Non-IID) data, this paper identifies a critical limitation of existing personalization methods: their excessive focus on local accuracy while neglecting out-of-distribution (OOD) generalization—a fundamental pillar of FedAvg’s robustness. We propose a unified evaluation paradigm that jointly optimizes local accuracy and OOD generalization, and design FLIU, an adaptive personalization update mechanism. Within the FedAvg framework, FLIU introduces learnable, client-specific scaling factors to dynamically balance global consistency and local adaptability. Extensive experiments across MNIST and CIFAR-10 under IID, pathological Non-IID, and Dirichlet Non-IID settings demonstrate that FLIU achieves high local accuracy while significantly improving OOD generalization—outperforming state-of-the-art personalized federated learning methods.

Addressing client drift in federated learning with heterogeneous data distributionsEvaluating local performance versus out-of-distribution generalization in personalized FLProposing individualized updates to improve both local and global model performance

Personalized Generation In Large Model Era: A Survey

Mar 04, 2025
YX
Yiyan Xu
🏛️ University of Science and Technology of China | University of Chinese Academy of Sciences | University of Massachusetts at Amherst | Nanyang Technological University | National University of Singapore

This survey addresses personalized generation (PGen)—the multimodal content creation tailored to user preferences and requirements—in the era of large language and foundation models. We propose the first unified conceptual framework, formally defining its core components, objectives, and abstract workflow. A hierarchical taxonomy is introduced, spanning text, image, audio, and other modalities while jointly considering personalization contexts and task types. We systematically review technical advances, benchmark datasets, and evaluation metrics. Our analysis identifies critical challenges, including cross-modal coordination and dynamic preference modeling, and highlight key future directions: interpretability, privacy-preserving personalization, and standardized evaluation protocols. As the inaugural comprehensive, structured, and extensible reference for PGen, this work bridges academic research and industrial practice across disciplines, enabling rigorous, reproducible, and ethically grounded development of personalized generative systems.

Formalizing key components and objectives of personalized generation.Proposing taxonomy and reviewing advancements in personalized generation.Surveying personalized content generation in large models era.

Communication barriers between data scientists and domain experts arise from oversimplified, accuracy-centric model performance reporting, hindering shared understanding of model limitations and contextual applicability. Method: We propose a visualization-mediated model explanation framework grounded in human-computer interaction principles, participatory design, and visual narrative techniques. This yields the first domain-expert-oriented model communication guideline—emphasizing risk, trade-offs, and situational appropriateness rather than isolated metrics like accuracy. An iterative empirical study was conducted using regression models, incorporating structured expert feedback for evaluation. Contribution/Results: The framework significantly improves domain experts’ ability to identify model limitations, recognize inherent trade-offs, and proactively make context-driven adoption decisions. Its core innovation lies in repositioning visualization as an interdisciplinary consensus-building medium—shifting the paradigm from “metric reporting” to “collaborative understanding.”

Communication gaps between data scientists and subject matter experts hinder model understanding.Traditional metrics fail to convey model risks, strengths, and limitations effectively.Visualization guidelines improve model performance communication and decision-making confidence.

From User Surveys to Telemetry-Driven Agents: Exploring the Potential of Personalized Productivity Solutions

Jan 17, 2024
SN
Subigya Nepal
🏛️ Microsoft Research | Dartmouth College | Brown University | Microsoft

Information workers often struggle to translate enterprise-provided productivity metrics into actionable behavioral improvements. To address this, we designed and evaluated a privacy-aware, personalized AI productivity agent powered by GPT-4, grounded in a mixed-methods approach: a survey of 363 knowledge workers and telemetry data from Microsoft Viva Insights. Our method introduces a two-stage “survey-driven + telemetry-informed” paradigm, integrating personified interaction, fine-grained behavioral modeling, and user-controllable privacy mechanisms. In a 40-participant A/B controlled experiment, the agent significantly outperformed conventional dashboards and narrative-based tools—increasing task completion efficiency by 27% and achieving a user satisfaction rating of 4.6/5.0. This work presents the first empirical validation of a human-centered, dual-loop (data + insight) AI agent design for enhancing knowledge worker effectiveness, demonstrating both its feasibility and efficacy in real-world organizational settings.

Addressing productivity challenges in modern workplaces using AI agentsBalancing personalization and privacy in AI-assisted productivity toolsTranslating productivity metrics into actionable insights for workers

MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences

Oct 03, 2024
GI
Genta Indra Winata
🏛️ Capital One | University of Toronto | Monash University Indonesia | Boston University

To address the misalignment between automatic evaluation metrics and human preferences in generative tasks, this paper proposes MetaMetrics—a calibratable meta-metric that supervisely weights and fuses existing metrics to model fine-grained human preferences across multimodal (language/vision), multilingual, and multi-domain settings. Methodologically, it introduces the first preference-dimension-aware metric calibration framework, enabling cross-modal unified evaluation and plug-and-play integration. The approach combines supervised meta-learning, multi-task joint optimization, and explicit modeling of human preference annotations. Experiments demonstrate that MetaMetrics significantly improves correlation with human judgments across multilingual text and vision generation tasks (average Kendall’s τ increase of +18.7%). Moreover, it exhibits strong generalization to unseen domains and models, maintaining robust alignment with human preferences without task-specific retraining.

Calibrate metrics to align with human preferences.Evaluate generation tasks across different modalities.Optimize existing metrics for multilingual and multi-domain scenarios.

Latest Papers

What's happening recently
View more

This study addresses the challenge of reliably evaluating business ideas generated by large language models, where expert judgments often exhibit structural disagreement due to multidimensional evaluation criteria. To investigate this issue, the authors construct PBIG-DATA, a dataset comprising 300 patent-derived business ideas and 3,000 expert ratings, and systematically compare aggregate versus personalized automated evaluators. Results demonstrate that expert disagreement is structural rather than random, and that personalized evaluators—trained on a target expert’s historical ratings—significantly outperform aggregate approaches, achieving higher fidelity to individual expert judgments across all evaluation dimensions. Notably, model reasoning similarity correlates significantly with human agreement only under personalized settings. These findings reveal that evaluation models trained on unified labels are fragile in diverse assessment contexts and underscore the necessity of evaluator-conditioned designs for robust automated assessment.

aggregate judgebusiness idea evaluationexpert disagreement

This study addresses the fundamental question of whether personalized interventions yield significantly greater benefits than a uniform optimal intervention. To this end, the authors propose a statistical hypothesis testing framework based on historical observational data, integrating nonparametric inference, causal inference, and asymptotic theory. The resulting test statistic is rigorously controlled for Type I error, asymptotically normal, and achieves minimal variance, thereby offering the first reliable tool for quantifying the incremental value of personalization. Extensive experiments across diverse real-world datasets—including job training programs, depression treatment trials, educational interventions, and recommendation systems—demonstrate the method’s broad applicability and superior performance.

hypothesis testinginterventionpersonalization

This work addresses the lack of a precise definition of “personalization” in existing algorithmic recourse methods, which hinders systematic evaluation of its impact on effectiveness, cost, and reasonableness. The paper formalizes personalization as individualized actionability by incorporating hard constraints—restricting the set of actionable features—and soft constraints—modeling users’ preferences over the value and cost of recommended actions—within a causal recourse framework. It further introduces a pre-recourse user prompting mechanism to enable personalized recommendations. Experimental results demonstrate that hard constraints substantially reduce both the effectiveness and reasonableness of recourse suggestions. Moreover, significant disparities emerge across social groups in terms of recourse cost and reasonableness, revealing a complex trade-off between personalized design and fairness.

algorithmic recoursefairnessindividual actionability

LLM4Perf: Large Language Models Are Effective Samplers for Multi-Objective Performance Modeling (Copy)

Dec 17, 2025
XW
Xin Wang
🏛️ The Hong Kong University of Science and Technology (Guangzhou) | York University

Software systems face challenges in multi-objective performance modeling due to vast configuration spaces and low sampling efficiency. Method: This paper proposes LLM4Perf—the first large language model (LLM)-based feedback-driven collaborative sampling framework. It innovatively integrates semantic information from configuration documentation with runtime performance feedback to enable dynamic configuration space pruning and online optimization of sampling strategies. Contribution/Results: Unlike conventional approaches, LLM4Perf empirically demonstrates, for the first time, the LLM’s generalizable pruning capability in performance modeling—significantly enhancing multiple baseline methods. Across 112 evaluation scenarios, it achieves optimal performance in 68.8%; across 448 baseline experiments, 91.5% show performance improvement attributable to its pruning mechanism. This work establishes a reproducible framework and robust empirical foundation for LLM-enabled performance engineering.

LLM4Perf framework outperforms traditional sampling methodsLLMs prune configuration space and refine strategies via feedbackLLMs sample configurations for multi-objective performance modeling

This work addresses the limitation of current large language models, which rely on single-pass generation during inference and struggle to improve personalized output quality even with increased computation. To overcome this, we propose a Test-Time Personalization (TTP) framework that samples multiple candidates from a personalized policy model and selects the best via a personalized reward model, enabling scalable optimization at inference time. We establish the first unified scaling law for Best-of-N performance of personalized reward models, revealing two failure modes—user-level collapse and query-level reward gaming—and introduce a probabilistic reward model with learnable variance to mitigate them. Experiments demonstrate that TTP consistently yields scaling gains across diverse policy models and personalized generation tasks, and the proposed scaling law accurately predicts empirical performance curves.

Best-of-NPersonalized Text GenerationReward Model

Hot Scholars

MN

Mohammad Naghizadeh

Associate Professor of Technology Management, Allameh Tabataba'i University
Innovation NetworkArtificial Intelligence and Text AnalyticsTechnology CollaborationSustainability
NT

Nathan Tsoi

Postdoctoral Researcher, The University of Texas at Austin
Robot LearningSystemsHuman-Robot Interaction
MS

Micol Spitale

Politecnico di Milano, Department of Electronics, Information, and Bioengineering
Human-Robot InteractionSocially Assistive RoboticsSocial Artificial Intelligence
MF

Michael Franke

University of Tübingen
Pragmatics (FormalExperimental & Computational)Probabilistic ModelingLanguage Evolution
MS

Muhammad Shafique

Professor, ECE, New York University (AD-UAE, Tandon-USA), Director eBRAIN Lab
Embedded Machine LearningBrain-Inspired ComputingRobust & Energy-Efficient System DesignSmart