personalized image evaluation

Designs and implements evaluation protocols, metrics, and datasets that measure how well image outputs (including generated images) align with individual user preferences and profile attributes, including few-shot example scenarios. Builds profile-inclusive benchmarks and diagnostic analyses to compare methods across profiling dimensions and to identify and characterize personalization failure modes.

personalizedimageevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current text-to-image generation models struggle to align with users’ implicit visual preferences. To address this limitation, this work proposes the first user-profile-driven evaluation benchmark that integrates psychological and demographic dimensions. The authors establish a collaborative data generation pipeline involving real users and AI agents and introduce a multidimensional framework for evaluating personalized image generation. Leveraging this benchmark, they systematically assess state-of-the-art personalization methods, uncovering critical shortcomings in their ability to align with individual user preferences. The findings highlight key challenges and offer concrete directions for future research in personalized generative modeling.

aesthetic preferencesimage evaluationpersonalized image generation

Current image generation evaluation benchmarks lag behind rapid technical advances and inadequately cover complex, creative tasks arising in real-world applications. To address this, we propose ECHO: a novel, application-oriented evaluation framework grounded in 31,000 user prompts and associated feedback crawled from social media—enabling dynamic, usage-informed benchmark construction. ECHO identifies previously unaddressed scenarios, including cross-lingual image editing and receipt generation with specified monetary amounts. It further introduces new quality metrics targeting color fidelity, identity consistency, and structural controllability. Extensive experiments demonstrate that ECHO effectively discriminates performance differences among state-of-the-art models, uncovers behavioral biases under practical usage conditions, and advances the evaluation paradigm from static, task-agnostic benchmarks toward user-driven, scenario-aware assessment.

Current evaluations fail to capture emerging real-world use casesExisting benchmarks lag behind rapidly advancing image generation capabilitiesThere is a gap between community perceptions and formal evaluation

LAPIS: A novel dataset for personalized image aesthetic assessment

Apr 10, 2025
AM
Anne-Sofie Maerten
🏛️ KU Leuven

This work addresses the lack of tailored data and models for Personalized Image Aesthetic Assessment (PIAA) on art images. We introduce LAPIS, the first PIAA benchmark dataset specifically designed for artworks, comprising 11,723 high-quality art images annotated with fine-grained visual attributes and multidimensional annotator-specific traits (e.g., art training background, aesthetic preferences). Methodologically, we propose the first image–individual co-annotation paradigm and conduct systematic benchmarking and ablation studies using state-of-the-art models. Results show that both individual and image attributes significantly improve prediction accuracy (−8.2% MAE), with their interaction proving critical; moreover, current models exhibit systematic biases on art images—e.g., overestimating aesthetic scores for abstract works. This work establishes a new benchmark, introduces a novel annotation paradigm, and provides reproducible pathways for advancing personalized aesthetic modeling in computational aesthetics.

Evaluates existing PIAA models using art-specific attributesIdentifies limitations in current artistic aesthetic prediction modelsIntroduces LAPIS dataset for personalized image aesthetic assessment

AI-generated faces (AIGFs) suffer from artifacts, identity drift, and text–image inconsistency; existing image quality assessment (IQA) metrics fail to capture human fine-grained preferences accurately. Method: We introduce FaceQ—the first large-scale, human-annotated face generation quality benchmark (12,255 images, 32,742 MOS scores), covering generation, customization, and restoration tasks—and systematically model human multidimensional preferences across realism, identity fidelity, and text–image alignment. We further propose F-Bench, a statistically rigorous, significance-driven evaluation framework for metric benchmarking. Results: Experiments reveal that mainstream IQA, face-specific QA (FQA), and AIGC-oriented IQA metrics exhibit consistently low correlation (<0.3) with human judgments. FaceQ is publicly released, and F-Bench establishes the first end-to-end, human-preference-aligned benchmark for face generation evaluation, enabling paradigmatic advances in generative quality assessment.

Assessing authenticity, identity fidelity and text-image correspondenceDeveloping comprehensive benchmark for face generation and restoration modelsEvaluating AI-generated face quality against human preferences

MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences

Oct 03, 2024
GI
Genta Indra Winata
🏛️ Capital One | University of Toronto | Monash University Indonesia | Boston University

To address the misalignment between automatic evaluation metrics and human preferences in generative tasks, this paper proposes MetaMetrics—a calibratable meta-metric that supervisely weights and fuses existing metrics to model fine-grained human preferences across multimodal (language/vision), multilingual, and multi-domain settings. Methodologically, it introduces the first preference-dimension-aware metric calibration framework, enabling cross-modal unified evaluation and plug-and-play integration. The approach combines supervised meta-learning, multi-task joint optimization, and explicit modeling of human preference annotations. Experiments demonstrate that MetaMetrics significantly improves correlation with human judgments across multilingual text and vision generation tasks (average Kendall’s τ increase of +18.7%). Moreover, it exhibits strong generalization to unseen domains and models, maintaining robust alignment with human preferences without task-specific retraining.

Calibrate metrics to align with human preferences.Evaluate generation tasks across different modalities.Optimize existing metrics for multilingual and multi-domain scenarios.

Latest Papers

What's happening recently
View more

This work addresses the challenge of modeling individual aesthetic preferences in zero-shot image aesthetic assessment, where user historical ratings are unavailable. It introduces user profiles as contextual signals and proposes a profile-guided personalization paradigm: building upon a frozen multimodal large language model, the method incorporates a profile-aware selective fusion module to enable controllable integration of visual and textual information, followed by profile-conditioned inference for personalized prediction. Notably, the approach requires no fine-tuning and achieves competitive zero-shot performance across multiple PIAA benchmarks. Its robustness persists even with coarse-grained user profiles, demonstrating the efficacy and potential of leveraging user profiles for zero-shot personalized aesthetic evaluation.

Aesthetic PreferenceMultimodal LLMPersonalized Image Aesthetics Assessment

This work addresses the unclear commercial value of current image generation models in real-world design scenarios by proposing the first framework that directly links image quality assessment to human payment decisions. The authors introduce ServImageBench, a dataset comprising 1.07k commercial tasks and 2.05k deliverables, along with ServImageScore—a multidimensional evaluation metric encompassing functional requirements, visual quality, and business necessity. Leveraging 33k human-annotated images, they train ServImageModel, a payment prediction model that achieves 82.00% accuracy in forecasting whether users are willing to pay for a given image. The model further outputs calibrated payment probabilities, offering an effective quantitative measure of an image’s commercial viability.

commercial benchmarkeconomic valuehuman payment decisions

Existing text-to-image generation models typically optimize for average population-level aesthetics, failing to capture individual users’ subjective preferences. To address this limitation, this work presents the first systematic modeling of individual aesthetic variation by introducing PAMELA, a personalized image evaluation dataset comprising 70,000 user ratings. The authors propose a joint personalized reward model that integrates high-quality human annotations with existing aesthetic data, trained jointly with prompt optimization and multi-user rating signals. This approach achieves higher accuracy in predicting individual preferences than most current state-of-the-art models attain on population-level preference tasks, substantially improving the personalization fidelity of generated images. The dataset and model are publicly released to support further research in personalized generative modeling.

aesthetic judgmentpersonalizationsubjective preference

Current evaluation of long-form question answering systems predominantly relies on human pairwise preference judgments, which often fail to capture the nuanced, expert-level assessment of in-depth research report quality. This work systematically examines the applicability and limitations of such meta-evaluation approaches in scientific QA using the ScholarQA-CS2 benchmark. The study finds that pairwise preferences are suitable only for system-level comparisons, whereas metric-level evaluation requires explicit dimension-wise annotations combined with domain-expert review. It identifies subjectivity as a central challenge and proposes a set of meta-evaluation design guidelines aligned with expert expectations, offering practical recommendations for future evaluation frameworks, annotator expertise matching, and reporting practices in deep research-oriented QA systems.

deep-research systemsevaluation benchmarkhuman pairwise preference

This work proposes the first creator-centered, fine-grained evaluation framework for text-to-image (T2I) generation, addressing the limitations of existing benchmarks that primarily assess basic alignment while neglecting the realism and creative expressiveness required in authentic artistic practice. Integrating expert insights from professional artists, the framework establishes a hierarchical structure of 56 verifiable metrics and conducts comprehensive multidimensional evaluation using 1,000 stratified prompts. For the first time, professional art creation workflows are incorporated into T2I assessment through a unified judgment model, Q-Judger (based on Qwen3.6-27B), trained under the supervision of 80 experts from global art institutions and enhanced with blind labeling and triple-review mechanisms. Experiments demonstrate that the framework effectively discriminates between state-of-the-art models and significantly outperforms existing benchmarks in evaluating realism and creative generation, offering reliable and interpretable signals for industrial-scale model optimization.

Artistic WorkflowsBenchmarkingCreative Generation

Hot Scholars

YF

Yuming Fang

Jiangxi University of Finance and Economics
Image ProcessingVideo Processing3D Multimedia Processing
BS

Bernt Schiele

Professor, Max Planck Institute for Informatics, Saarland University, Saarland Informatics Campus
Computer VisionMachine LearningArtificial IntelligenceAutonomous Driving
MY

Mengping Yang

East China University of Science and Technology
Few-shot LearningGenerative Models
DV

Deepak Vasisht

University of Illinois at Urbana-Champaign
Wireless NetworksInternet of Things
ZZ

Zecheng Zhao

The University of Queensland
Video Retrieval