QQJ: Quantifying Qualitative Judgment for Scalable and Human-Aligned Evaluation of Generative AI

📅 2026-05-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing evaluation methods for generative AI struggle to simultaneously balance human-perceived quality, scalability, and consistency in open-ended, creative tasks. This work proposes the QQJ framework, which uniquely integrates structured qualitative judgments with large language model (LLM)-based assessment. By leveraging expert-defined, multidimensional scoring criteria to explicitly articulate quality constructs and calibrating LLM evaluators with few-shot, high-quality human annotations, QQJ achieves strong alignment with human judgments while maintaining automation. The framework decouples quality definition from evaluation execution, substantially enhancing interpretability, stability, and cross-task generalization. Empirical results demonstrate that QQJ outperforms conventional automatic metrics and unconstrained LLM evaluators across both text and image generation tasks, and effectively detects critical failure modes such as hallucination and intent misalignment.
📝 Abstract
The rapid progress of generative artificial intelligence has exposed fundamental limitations in existing evaluation methodologies, particularly for open-ended, creative, and human-facing tasks. Traditional automatic metrics rely on surface-level statistical similarity and often fail to reflect human perceptions of quality, while purely human evaluation, although reliable, is costly, subjective, and difficult to scale. Recent approaches using large language models as evaluators offer improved scalability but frequently lack explicit grounding in human-defined evaluation principles, leading to bias and inconsistency. In this paper, we introduce Quantifying Qualitative Judgment (QQJ), a scalable and human-centric evaluation framework that explicitly bridges the gap between human judgment and automated assessment. QQJ separates the definition of quality from its execution by anchoring evaluation in expert-designed, multi-dimensional rubrics and calibrating large language model evaluators to align with expert reasoning using a small, high-quality annotation set. This design enables consistent, interpretable, and scalable evaluation across diverse generative tasks and modalities. Extensive experiments on text and image generation demonstrate that QQJ achieves substantially stronger alignment with human judgment than traditional automatic metrics and unconstrained LLM-based evaluators. Moreover, QQJ exhibits improved stability across repeated evaluations and superior diagnostic capability in identifying critical failure modes such as hallucination and intent mismatch. These results indicate that structured qualitative judgment can be operationalized at scale without sacrificing interpretability or human alignment, positioning QQJ as a practical foundation for reliable evaluation of modern generative AI systems.
Problem

Research questions and friction points this paper is trying to address.

generative AI evaluation
human alignment
scalability
qualitative judgment
evaluation bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantifying Qualitative Judgment
human-aligned evaluation
large language model calibration
multi-dimensional rubrics
generative AI evaluation
💼 Related Jobs
No related jobs found.
M
Marjan Veysi
AI Lab, Arioobarzan Engineering Team, Shiraz, Iran
P
Pirooz Shamsinejadbabaki
Department of Computer Engineering and Information Technology, Shiraz University of Technology, Shiraz, Iran
M
Mohammad Zare
AI Lab, Arioobarzan Engineering Team, Shiraz, Iran
M
Mohammad Sabouri
Department of Informatics, Bioengineering, Robotics and Systems Engineering, University of Genoa, Genoa, Italy