Score
Designs and executes evaluation protocols, metrics, and analyses that quantify both per-client personalized performance and aggregate/global model generalization, and that compare personalized approaches to global or centralized models. Builds experiments and diagnostic studies to measure the personalization-vs-generalization tradeoff, report per-client and overall accuracy, and identify factors or conditions that degrade one objective relative to the other.
To address the client drift and generalization imbalance in federated learning under non-independent and identically distributed (Non-IID) data, this paper identifies a critical limitation of existing personalization methods: their excessive focus on local accuracy while neglecting out-of-distribution (OOD) generalization—a fundamental pillar of FedAvg’s robustness. We propose a unified evaluation paradigm that jointly optimizes local accuracy and OOD generalization, and design FLIU, an adaptive personalization update mechanism. Within the FedAvg framework, FLIU introduces learnable, client-specific scaling factors to dynamically balance global consistency and local adaptability. Extensive experiments across MNIST and CIFAR-10 under IID, pathological Non-IID, and Dirichlet Non-IID settings demonstrate that FLIU achieves high local accuracy while significantly improving OOD generalization—outperforming state-of-the-art personalized federated learning methods.
This survey addresses personalized generation (PGen)—the multimodal content creation tailored to user preferences and requirements—in the era of large language and foundation models. We propose the first unified conceptual framework, formally defining its core components, objectives, and abstract workflow. A hierarchical taxonomy is introduced, spanning text, image, audio, and other modalities while jointly considering personalization contexts and task types. We systematically review technical advances, benchmark datasets, and evaluation metrics. Our analysis identifies critical challenges, including cross-modal coordination and dynamic preference modeling, and highlight key future directions: interpretability, privacy-preserving personalization, and standardized evaluation protocols. As the inaugural comprehensive, structured, and extensible reference for PGen, this work bridges academic research and industrial practice across disciplines, enabling rigorous, reproducible, and ethically grounded development of personalized generative systems.
Communication barriers between data scientists and domain experts arise from oversimplified, accuracy-centric model performance reporting, hindering shared understanding of model limitations and contextual applicability. Method: We propose a visualization-mediated model explanation framework grounded in human-computer interaction principles, participatory design, and visual narrative techniques. This yields the first domain-expert-oriented model communication guideline—emphasizing risk, trade-offs, and situational appropriateness rather than isolated metrics like accuracy. An iterative empirical study was conducted using regression models, incorporating structured expert feedback for evaluation. Contribution/Results: The framework significantly improves domain experts’ ability to identify model limitations, recognize inherent trade-offs, and proactively make context-driven adoption decisions. Its core innovation lies in repositioning visualization as an interdisciplinary consensus-building medium—shifting the paradigm from “metric reporting” to “collaborative understanding.”
Information workers often struggle to translate enterprise-provided productivity metrics into actionable behavioral improvements. To address this, we designed and evaluated a privacy-aware, personalized AI productivity agent powered by GPT-4, grounded in a mixed-methods approach: a survey of 363 knowledge workers and telemetry data from Microsoft Viva Insights. Our method introduces a two-stage “survey-driven + telemetry-informed” paradigm, integrating personified interaction, fine-grained behavioral modeling, and user-controllable privacy mechanisms. In a 40-participant A/B controlled experiment, the agent significantly outperformed conventional dashboards and narrative-based tools—increasing task completion efficiency by 27% and achieving a user satisfaction rating of 4.6/5.0. This work presents the first empirical validation of a human-centered, dual-loop (data + insight) AI agent design for enhancing knowledge worker effectiveness, demonstrating both its feasibility and efficacy in real-world organizational settings.
To address the misalignment between automatic evaluation metrics and human preferences in generative tasks, this paper proposes MetaMetrics—a calibratable meta-metric that supervisely weights and fuses existing metrics to model fine-grained human preferences across multimodal (language/vision), multilingual, and multi-domain settings. Methodologically, it introduces the first preference-dimension-aware metric calibration framework, enabling cross-modal unified evaluation and plug-and-play integration. The approach combines supervised meta-learning, multi-task joint optimization, and explicit modeling of human preference annotations. Experiments demonstrate that MetaMetrics significantly improves correlation with human judgments across multilingual text and vision generation tasks (average Kendall’s τ increase of +18.7%). Moreover, it exhibits strong generalization to unseen domains and models, maintaining robust alignment with human preferences without task-specific retraining.
This study addresses the challenge of reliably evaluating business ideas generated by large language models, where expert judgments often exhibit structural disagreement due to multidimensional evaluation criteria. To investigate this issue, the authors construct PBIG-DATA, a dataset comprising 300 patent-derived business ideas and 3,000 expert ratings, and systematically compare aggregate versus personalized automated evaluators. Results demonstrate that expert disagreement is structural rather than random, and that personalized evaluators—trained on a target expert’s historical ratings—significantly outperform aggregate approaches, achieving higher fidelity to individual expert judgments across all evaluation dimensions. Notably, model reasoning similarity correlates significantly with human agreement only under personalized settings. These findings reveal that evaluation models trained on unified labels are fragile in diverse assessment contexts and underscore the necessity of evaluator-conditioned designs for robust automated assessment.
This study addresses the fundamental question of whether personalized interventions yield significantly greater benefits than a uniform optimal intervention. To this end, the authors propose a statistical hypothesis testing framework based on historical observational data, integrating nonparametric inference, causal inference, and asymptotic theory. The resulting test statistic is rigorously controlled for Type I error, asymptotically normal, and achieves minimal variance, thereby offering the first reliable tool for quantifying the incremental value of personalization. Extensive experiments across diverse real-world datasets—including job training programs, depression treatment trials, educational interventions, and recommendation systems—demonstrate the method’s broad applicability and superior performance.
This work addresses the lack of a precise definition of “personalization” in existing algorithmic recourse methods, which hinders systematic evaluation of its impact on effectiveness, cost, and reasonableness. The paper formalizes personalization as individualized actionability by incorporating hard constraints—restricting the set of actionable features—and soft constraints—modeling users’ preferences over the value and cost of recommended actions—within a causal recourse framework. It further introduces a pre-recourse user prompting mechanism to enable personalized recommendations. Experimental results demonstrate that hard constraints substantially reduce both the effectiveness and reasonableness of recourse suggestions. Moreover, significant disparities emerge across social groups in terms of recourse cost and reasonableness, revealing a complex trade-off between personalized design and fairness.
Software systems face challenges in multi-objective performance modeling due to vast configuration spaces and low sampling efficiency. Method: This paper proposes LLM4Perf—the first large language model (LLM)-based feedback-driven collaborative sampling framework. It innovatively integrates semantic information from configuration documentation with runtime performance feedback to enable dynamic configuration space pruning and online optimization of sampling strategies. Contribution/Results: Unlike conventional approaches, LLM4Perf empirically demonstrates, for the first time, the LLM’s generalizable pruning capability in performance modeling—significantly enhancing multiple baseline methods. Across 112 evaluation scenarios, it achieves optimal performance in 68.8%; across 448 baseline experiments, 91.5% show performance improvement attributable to its pruning mechanism. This work establishes a reproducible framework and robust empirical foundation for LLM-enabled performance engineering.
This work addresses the limitation of current large language models, which rely on single-pass generation during inference and struggle to improve personalized output quality even with increased computation. To overcome this, we propose a Test-Time Personalization (TTP) framework that samples multiple candidates from a personalized policy model and selects the best via a personalized reward model, enabling scalable optimization at inference time. We establish the first unified scaling law for Best-of-N performance of personalized reward models, revealing two failure modes—user-level collapse and query-level reward gaming—and introduce a probabilistic reward model with learnable variance to mitigate them. Experiments demonstrate that TTP consistently yields scaling gains across diverse policy models and personalized generation tasks, and the proposed scaling law accurately predicts empirical performance curves.