PortraitAes: Intent-Conditioned Structured Portrait Aesthetics Assessment

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inconsistent evaluation criteria in existing portrait aesthetics assessment methods caused by neglecting photographic intent. To this end, it proposes an intent-conditioned assessment framework that constructs an intent decomposition structure alongside an expert rule system. The method designs an end-to-end training pipeline incorporating multi-task learning, intent routing, and score fusion, while introducing Gaussian calibration and cross-image ranking mechanisms to achieve precise scoring. Experimental results demonstrate that the proposed approach attains a Pearson correlation coefficient of 0.924, significantly outperforming general-purpose multimodal large models and existing baselines. This work establishes a novel paradigm for intent-aware portrait aesthetics assessment.
📝 Abstract
Portrait aesthetic assessment assigns comparable scores according to how effectively human-centered images fulfill their photographic intent. These scores support data filtering, candidate selection, and preference modeling in image-generation pipelines. Existing methods typically predict a single aesthetic score or use general-purpose MLLMs without conditioning on photographic intent. This omission matters because the same blur, pose, lighting, or framing choice may serve one photographic intent but undermine another. These models thus learn context-agnostic aesthetic priors and yield inconsistent, inaccurate, misleading judgments for portraits with distinct photographic objectives. We introduce PortraitAes-Bench, an 11K-scale benchmark that decomposes this task into intent-conditioned subjudgments. Expert-authored rubrics define nine photographic intents, six first-level dimensions, and 22 secondary criteria. They support a structured pipeline for intent routing, specialist assessment, verification, and score fusion. Following this structure, we train PortraitAes with multi-task supervision. We then improve score comparability through Gaussian score calibration and within-dimension cross-image ranking. On the standard benchmark, PortraitAes achieves a Pearson correlation of 0.924 and a Spearman rank correlation of 0.934. On the hard-case set, its Pearson correlation is 0.829 and its Spearman rank correlation is 0.795. Across both sets, PortraitAes outperforms the evaluated general-purpose MLLMs and specialized aesthetic baselines.
Problem

Research questions and friction points this paper is trying to address.

Portrait Aesthetic Assessment
Photographic Intent
Aesthetics Evaluation
Multi-modal Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Intent-Conditioned Assessment
Structured Pipeline
Multi-task Supervision
Gaussian Score Calibration
Cross-Image Ranking
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Junzhou Xie
Southeast University
H
Haozhong Xiong
Qwen Business Unit of Alibaba
X
Xunyun Tian
Qwen Business Unit of Alibaba
Kaile Du
Kaile Du
Southeast University
Continual learningClass-incremental learningMulti-label learning
T
Tianchen Yu
Qwen Business Unit of Alibaba
Q
Qiang Li
Qwen Business Unit of Alibaba
W
Wei Liu
Qwen Business Unit of Alibaba
J
Jiaming Liu
Qwen Business Unit of Alibaba
R
Ruihua Huang
Qwen Business Unit of Alibaba
Y
Yang Shi
Qwen Business Unit of Alibaba
Guangcan Liu
Guangcan Liu
Professor of Southeast University, Nanjing, China.
machine learningcomputer visionimage processing