generation-based evaluation

Designs, implements, and analyzes evaluation protocols and quantitative metrics for open-ended generative systems, covering automatic generation-evaluation metrics, grounded and production-oriented assessments, and multimodal output quality measures. Conducts human-annotation studies and analytical experiments to validate outputs and decoding behavior (e.g., beam vs sampling), compare evaluation paradigms, identify human-preferred generations, and summarize strengths and failure modes.

generation-basedevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.68
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$209K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences

Oct 03, 2024
GI
Genta Indra Winata
🏛️ Capital One | University of Toronto | Monash University Indonesia | Boston University

To address the misalignment between automatic evaluation metrics and human preferences in generative tasks, this paper proposes MetaMetrics—a calibratable meta-metric that supervisely weights and fuses existing metrics to model fine-grained human preferences across multimodal (language/vision), multilingual, and multi-domain settings. Methodologically, it introduces the first preference-dimension-aware metric calibration framework, enabling cross-modal unified evaluation and plug-and-play integration. The approach combines supervised meta-learning, multi-task joint optimization, and explicit modeling of human preference annotations. Experiments demonstrate that MetaMetrics significantly improves correlation with human judgments across multilingual text and vision generation tasks (average Kendall’s τ increase of +18.7%). Moreover, it exhibits strong generalization to unseen domains and models, maintaining robust alignment with human preferences without task-specific retraining.

Calibrate metrics to align with human preferences.Evaluate generation tasks across different modalities.Optimize existing metrics for multilingual and multi-domain scenarios.

Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation

Sep 25, 2025
SA
Seyed Amir Kasaei
🏛️ Sharif University of Technology

Automated evaluation metrics for text-to-image generation are often adopted empirically, lacking systematic validation against human judgments—particularly for compositional alignment involving objects, attributes, and relational semantics. Method: We propose a multidimensional analytical framework and conduct the first unified benchmarking of three metric families—VQA-based, embedding-based, and image-only—using large-scale human annotations as the gold standard. Contribution/Results: No single metric family dominates across all dimensions; VQA-based metrics are not universally superior, and certain embedding-based metrics exhibit higher discriminative power for fine-grained relational alignment. Image-only metrics show limited capacity to model compositional alignment. Our findings reveal strong task dependency in metric behavior, providing empirical guidance and methodological insights for principled metric selection in text-to-image evaluation.

Analyzing metric performance across different types of compositional challengesAssessing automated metrics' alignment with human judgment in evaluationEvaluating how well text-to-image models capture compositional prompts

This work addresses the diversity calibration imbalance of large language models (LLMs) in open-ended generation: insufficient diversity in creative tasks and excessive, hallucination-prone diversity in factual tasks. We propose the “Generative Space Size” (GSS) theoretical framework to unify the analysis of both phenomena. By constructing GSSBench—a dedicated benchmark—and developing an evaluation suite based on unsupervised internal representation metrics (e.g., EigenScore), we empirically reveal that both imbalances stem from structural biases in the model’s latent space. Our method integrates hallucination detection with semantic diversity quantification, enabling interpretable diagnosis of ambiguous prompts and reasoning biases, and supporting controllable generation. Experiments demonstrate that the framework significantly enhances diversity in creative generation while effectively suppressing factual hallucinations, thereby improving output consistency and reliability.

Addressing model collapse and hallucination through generation spaceCalibrating LLM output diversity across different tasksDetecting prompt ambiguity and improving model grounding

Existing evaluation methods for generative AI struggle to simultaneously balance human-perceived quality, scalability, and consistency in open-ended, creative tasks. This work proposes the QQJ framework, which uniquely integrates structured qualitative judgments with large language model (LLM)-based assessment. By leveraging expert-defined, multidimensional scoring criteria to explicitly articulate quality constructs and calibrating LLM evaluators with few-shot, high-quality human annotations, QQJ achieves strong alignment with human judgments while maintaining automation. The framework decouples quality definition from evaluation execution, substantially enhancing interpretability, stability, and cross-task generalization. Empirical results demonstrate that QQJ outperforms conventional automatic metrics and unconstrained LLM evaluators across both text and image generation tasks, and effectively detects critical failure modes such as hallucination and intent misalignment.

evaluation biasgenerative AI evaluationhuman alignment

This study addresses the lack of quantifiable definitions and evaluation criteria for “creativity” in generative AI. We propose, for the first time, a behaviorally grounded framework for assessing practical creativity in image generation models. Methodologically, we introduce an interpretable, three-dimensional metric encompassing diversity, novelty, and appropriateness; integrate multi-model comparative experiments, human perceptual evaluation, and statistical significance testing to ensure reproducibility, comparability, and alignment with human intuition. Validation across mainstream image-to-image translation models demonstrates strong agreement between our metric rankings and human subjective scores (Spearman ρ > 0.85), significantly outperforming existing black-box evaluation approaches. The framework enables objective, cross-model comparison of creative performance and provides empirical guidance for users selecting optimal generative models according to task-specific requirements.

Defining creative behavior in AI image generatorsEvaluating measures against human intuitionQuantifying creativity for model selection

Latest Papers

What's happening recently
View more

This work addresses the high computational cost, rigidity, and lack of interpretability and user customization in existing evaluation methods for visual generative models. We propose the Evaluation Agent (EA) framework, which emulates human-like rapid judgment from few examples by decomposing natural language evaluation requests into subtasks through multi-turn dynamic interaction. The framework autonomously generates prompts, samples content, invokes tools, and iteratively refines its evaluation plan. Leveraging a locally deployed agent, EA-3B—fine-tuned from Qwen2.5-3B-Instruct with multi-turn reasoning planning, tool calling, and history-conditioned instructions—the method achieves comparable performance to standard T2I/T2V benchmarks while reducing evaluation time to just 10% of conventional approaches. We release the Open-EA framework and the EA-CoT-10K dataset containing 10K reasoning chains, and demonstrate partial cross-family transferability on video generation models.

computational costevaluation efficiencyexplainable assessment

This work addresses the lack of standardized, automated evaluation methods for design animation video generation, which hinders objective assessment of generation quality under structured constraints. To bridge this gap, we propose the first multidimensional automatic evaluation framework tailored specifically for design animations. Leveraging computer vision and video analysis techniques, the framework quantifies key generative attributes across four dimensions: layout fidelity, motion correctness, temporal consistency, and content fidelity. Operating without human intervention, it delivers an objective and reproducible benchmark that enables fair comparison among diverse generative models and supports sustained progress in the field.

compositional fidelitydesign video generationevaluation framework

Current evaluations of generative AI alignment predominantly rely on single benchmarks, which inadequately capture the diversity of human judgment across cultural, demographic, and contextual dimensions. This work proposes a personified evaluation framework grounded in state-space constraints, modeling alignment assessment for the first time as a structured dynamical system on a latent manifold. By synthesizing diverse cognitive personas, the approach enables perspective-dependent, pluralistic evaluation. It incorporates a dynamic, survival-driven regulatory mechanism to ensure cognitive fidelity and validates robustness through stochastic prompt perturbations and sequential reasoning stability analysis. Experiments demonstrate that generative models can consistently instantiate multifaceted evaluative personas, while also revealing state drift and semantic inconsistencies under static alignment constraints.

alignment stabilitycognitive emulationgenerative AI

This study addresses the limitations of existing AI evaluation methods, which often fail to align with real-world user needs, contextual nuances, and local policies, while manual assessment remains difficult to scale. To bridge this gap, the authors propose an auditable and iterative, context-aware evaluation framework that integrates persona-driven test case generation, domain-specific scoring rubrics, and a hybrid adjudication mechanism combining human reviewers and LLM-based judges. Automated scoring is activated only when sufficient agreement between LLM judgments and human annotations is achieved. A three-week pilot across four organizations involving 108 annotated question-answer pairs demonstrates that the approach effectively balances policy alignment with scalable automation, enabling reliable end-to-end evaluation of AI systems.

AI evaluationcontextual alignmenthuman-aligned scoring

Current evaluations of AI models lack standardized protocols, with institutions selectively employing benchmarks in ways that hinder cross-study comparability and raise concerns about scientific validity. This work introduces Benchmarking-Cultures-25, a dataset encompassing 231 benchmarks from 139 model releases, and combines qualitative content analysis with a unified categorization framework to systematically expose the fragmentation in benchmark selection: 63.2% of benchmarks are used by only a single institution, and 38.5% appear just once. Moreover, many benchmarks marketed as “general-purpose” disproportionately emphasize STEM—particularly mathematics—while often neglecting construct validity. The study further proposes a taxonomy aligning ostensibly disparate terminologies to their underlying measurement signals and develops an interactive tool revealing that benchmarks frequently serve marketing narratives rather than rigorous scientific assessment.

AI evaluationbenchmarkingconstruct validity

Hot Scholars

YL

Yanyu Li

PhD, Northeastern University
Machine Learning
JG

Jiuxiang Gu

Adobe Research
Computer VisionNatural Language ProcessingMachine Learning
ST

Sergey Tulyakov

Director of Research, Snap Inc.
computer visionmachine learning
WM

Willi Menapace

University of Trento
deep learningcomputer vision