Score
Designs, implements, and analyzes evaluation protocols and quantitative metrics for open-ended generative systems, covering automatic generation-evaluation metrics, grounded and production-oriented assessments, and multimodal output quality measures. Conducts human-annotation studies and analytical experiments to validate outputs and decoding behavior (e.g., beam vs sampling), compare evaluation paradigms, identify human-preferred generations, and summarize strengths and failure modes.
Existing automated evaluation methods for generative content lack a systematic, cross-modal framework. Method: This paper conducts a large-scale literature review and cross-modal comparative analysis to establish, for the first time, a unified evaluation taxonomy covering text, image, and speech modalities. It identifies five fundamental evaluation paradigms and empirically validates their consistent applicability across three representative generative tasks. Furthermore, it introduces a comparability analysis framework to construct a structured knowledge graph that clarifies capability boundaries and limitations of existing methods per modality. Contributions/Results: (1) The first cross-modal unified classification system for generative evaluation; (2) abstraction of generalizable, transferable evaluation paradigms; and (3) a theoretical foundation and practical methodology for cross-modal consistent evaluation and joint metric design. This work bridges critical gaps in evaluating multimodal generative models and enables principled, interoperable assessment across modalities.
To address the misalignment between automatic evaluation metrics and human preferences in generative tasks, this paper proposes MetaMetrics—a calibratable meta-metric that supervisely weights and fuses existing metrics to model fine-grained human preferences across multimodal (language/vision), multilingual, and multi-domain settings. Methodologically, it introduces the first preference-dimension-aware metric calibration framework, enabling cross-modal unified evaluation and plug-and-play integration. The approach combines supervised meta-learning, multi-task joint optimization, and explicit modeling of human preference annotations. Experiments demonstrate that MetaMetrics significantly improves correlation with human judgments across multilingual text and vision generation tasks (average Kendall’s τ increase of +18.7%). Moreover, it exhibits strong generalization to unseen domains and models, maintaining robust alignment with human preferences without task-specific retraining.
Automated evaluation metrics for text-to-image generation are often adopted empirically, lacking systematic validation against human judgments—particularly for compositional alignment involving objects, attributes, and relational semantics. Method: We propose a multidimensional analytical framework and conduct the first unified benchmarking of three metric families—VQA-based, embedding-based, and image-only—using large-scale human annotations as the gold standard. Contribution/Results: No single metric family dominates across all dimensions; VQA-based metrics are not universally superior, and certain embedding-based metrics exhibit higher discriminative power for fine-grained relational alignment. Image-only metrics show limited capacity to model compositional alignment. Our findings reveal strong task dependency in metric behavior, providing empirical guidance and methodological insights for principled metric selection in text-to-image evaluation.
This work addresses the diversity calibration imbalance of large language models (LLMs) in open-ended generation: insufficient diversity in creative tasks and excessive, hallucination-prone diversity in factual tasks. We propose the “Generative Space Size” (GSS) theoretical framework to unify the analysis of both phenomena. By constructing GSSBench—a dedicated benchmark—and developing an evaluation suite based on unsupervised internal representation metrics (e.g., EigenScore), we empirically reveal that both imbalances stem from structural biases in the model’s latent space. Our method integrates hallucination detection with semantic diversity quantification, enabling interpretable diagnosis of ambiguous prompts and reasoning biases, and supporting controllable generation. Experiments demonstrate that the framework significantly enhances diversity in creative generation while effectively suppressing factual hallucinations, thereby improving output consistency and reliability.
Existing evaluation methods for generative AI struggle to simultaneously balance human-perceived quality, scalability, and consistency in open-ended, creative tasks. This work proposes the QQJ framework, which uniquely integrates structured qualitative judgments with large language model (LLM)-based assessment. By leveraging expert-defined, multidimensional scoring criteria to explicitly articulate quality constructs and calibrating LLM evaluators with few-shot, high-quality human annotations, QQJ achieves strong alignment with human judgments while maintaining automation. The framework decouples quality definition from evaluation execution, substantially enhancing interpretability, stability, and cross-task generalization. Empirical results demonstrate that QQJ outperforms conventional automatic metrics and unconstrained LLM evaluators across both text and image generation tasks, and effectively detects critical failure modes such as hallucination and intent misalignment.
This study addresses the lack of quantifiable definitions and evaluation criteria for “creativity” in generative AI. We propose, for the first time, a behaviorally grounded framework for assessing practical creativity in image generation models. Methodologically, we introduce an interpretable, three-dimensional metric encompassing diversity, novelty, and appropriateness; integrate multi-model comparative experiments, human perceptual evaluation, and statistical significance testing to ensure reproducibility, comparability, and alignment with human intuition. Validation across mainstream image-to-image translation models demonstrates strong agreement between our metric rankings and human subjective scores (Spearman ρ > 0.85), significantly outperforming existing black-box evaluation approaches. The framework enables objective, cross-model comparison of creative performance and provides empirical guidance for users selecting optimal generative models according to task-specific requirements.
This work addresses the high computational cost, rigidity, and lack of interpretability and user customization in existing evaluation methods for visual generative models. We propose the Evaluation Agent (EA) framework, which emulates human-like rapid judgment from few examples by decomposing natural language evaluation requests into subtasks through multi-turn dynamic interaction. The framework autonomously generates prompts, samples content, invokes tools, and iteratively refines its evaluation plan. Leveraging a locally deployed agent, EA-3B—fine-tuned from Qwen2.5-3B-Instruct with multi-turn reasoning planning, tool calling, and history-conditioned instructions—the method achieves comparable performance to standard T2I/T2V benchmarks while reducing evaluation time to just 10% of conventional approaches. We release the Open-EA framework and the EA-CoT-10K dataset containing 10K reasoning chains, and demonstrate partial cross-family transferability on video generation models.
This work addresses the lack of standardized, automated evaluation methods for design animation video generation, which hinders objective assessment of generation quality under structured constraints. To bridge this gap, we propose the first multidimensional automatic evaluation framework tailored specifically for design animations. Leveraging computer vision and video analysis techniques, the framework quantifies key generative attributes across four dimensions: layout fidelity, motion correctness, temporal consistency, and content fidelity. Operating without human intervention, it delivers an objective and reproducible benchmark that enables fair comparison among diverse generative models and supports sustained progress in the field.
Current evaluations of generative AI alignment predominantly rely on single benchmarks, which inadequately capture the diversity of human judgment across cultural, demographic, and contextual dimensions. This work proposes a personified evaluation framework grounded in state-space constraints, modeling alignment assessment for the first time as a structured dynamical system on a latent manifold. By synthesizing diverse cognitive personas, the approach enables perspective-dependent, pluralistic evaluation. It incorporates a dynamic, survival-driven regulatory mechanism to ensure cognitive fidelity and validates robustness through stochastic prompt perturbations and sequential reasoning stability analysis. Experiments demonstrate that generative models can consistently instantiate multifaceted evaluative personas, while also revealing state drift and semantic inconsistencies under static alignment constraints.
This study addresses the limitations of existing AI evaluation methods, which often fail to align with real-world user needs, contextual nuances, and local policies, while manual assessment remains difficult to scale. To bridge this gap, the authors propose an auditable and iterative, context-aware evaluation framework that integrates persona-driven test case generation, domain-specific scoring rubrics, and a hybrid adjudication mechanism combining human reviewers and LLM-based judges. Automated scoring is activated only when sufficient agreement between LLM judgments and human annotations is achieved. A three-week pilot across four organizations involving 108 annotated question-answer pairs demonstrates that the approach effectively balances policy alignment with scalable automation, enabling reliable end-to-end evaluation of AI systems.
Current evaluations of AI models lack standardized protocols, with institutions selectively employing benchmarks in ways that hinder cross-study comparability and raise concerns about scientific validity. This work introduces Benchmarking-Cultures-25, a dataset encompassing 231 benchmarks from 139 model releases, and combines qualitative content analysis with a unified categorization framework to systematically expose the fragmentation in benchmark selection: 63.2% of benchmarks are used by only a single institution, and 38.5% appear just once. Moreover, many benchmarks marketed as “general-purpose” disproportionately emphasize STEM—particularly mathematics—while often neglecting construct validity. The study further proposes a taxonomy aligning ostensibly disparate terminologies to their underlying measurement signals and develops an interactive tool revealing that benchmarks frequently serve marketing narratives rather than rigorous scientific assessment.