MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决LLM评估者对响应、指令和评分标准变化敏感的问题,提出MAWILE工作台,通过控制扰动检验评估者的敏感性。
📝 Abstract
Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to incidental changes in the evaluated response, judge instructions, and scoring rubric. Existing systems examine important subsets of these failure modes, but auditing a configured judge requires testing both the judge instrument and the items it evaluates. We introduce MAWILE, a developer-facing workbench for auditing judge sensitivity across four surfaces: the judge prompt, judge rubric, target-system input, and target-system output. Given a user-supplied judge and representative evaluation items, MAWILE constructs and validates controlled perturbations, re-executes the judge, and localizes the resulting sensitivity. Each perturbation declares whether the verdict should remain invariant or change in a specified direction, allowing the same system to measure both robustness to irrelevant variations and sensitivity to meaningful changes. MAWILE audits binary, ordinal, and pairwise judges without requiring gold labels. The code for this tool is available at: github.com/megagonlabs/mawile-judge.
Problem

Research questions and friction points this paper is trying to address.

Large language model (LLM) judges
sensitivity to incidental changes
auditing a configured judge
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Axis Workbench
LLM Evaluators
Controlled Perturbations
Sensitivity Auditing
No Gold Labels Required
🔎 Similar Papers
No similar papers found.