Mubric: Mutation Testing-Guided Rubric Generation for LLM Evaluation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant challenge of automatically generating evaluation rubrics that reliably capture task-specific quality requirements for large language models (LLMs). To this end, this work pioneers the introduction of mutation testing from software engineering into LLM evaluation by establishing a closed-loop "defect injection–scoring feedback" mechanism. Through defect mining grounded in real-world preference data, the abstraction of reusable mutation operators, and an iterative feedback algorithm, the proposed approach automates the generation of highly robust rubrics. Evaluated across 703 tasks spanning four domains, the method achieves state-of-the-art overall evaluation accuracy, outperforming the strongest baseline by 7.48 percentage points. These results demonstrate its effectiveness in enhancing the capacity of rubrics to discriminate among varying levels of response quality.
📝 Abstract
Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We introduce Mubric, a mutation testing-guided approach to rubric generation. Mutation testing, a classic software testing methodology, evaluates a test suite by injecting faults into programs and checking whether the tests detect them. We draw an analogy between test suites and rubrics: if a rubric captures an important quality requirement, introducing a corresponding defect into an otherwise high-quality response should reduce its score. Mubric first mines common defects from real pairs of preferred and dispreferred responses and abstracts these defects into reusable mutation operators, each specifying how to introduce a particular type of response defect. For a new task, it applies relevant operators to a reference response, checks whether the injected defects reduce response quality, and uses insufficiently penalized defects to refine the rubric. We evaluate Mubric on 703 tasks across four representative domains against six advanced rubric generation methods. Mubric achieves the highest overall evaluation accuracy, outperforming the strongest baseline by 7.48 percentage points.
Problem

Research questions and friction points this paper is trying to address.

Rubric Generation
LLM Evaluation
Mutation Testing
Response Quality Assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mutation Testing
Rubric Generation
LLM Evaluation
Mutation Operators
Automated Assessment
🔎 Similar Papers
No similar papers found.