Beyond Binary Correctness: Scaling Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks

📅 2026-03-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing evaluation methods struggle to effectively assess agent performance in subjective, context-dependent, long-horizon enterprise tasks due to their reliance on binary correctness judgments. To address this limitation, this work proposes LH-Bench, a three-pillar evaluation framework that integrates expert-designed rubrics, step-level ground-truth artifact annotations, and pairwise human preference comparisons to enable fine-grained and scalable quantitative assessment. Validation on two real-world scenarios—Figma-to-code translation and procedural content generation—demonstrates that expert-crafted rubrics achieve substantially higher inter-rater agreement than LLM-generated ones (Cohen’s Kappa: 0.60 vs. 0.46), and human preference data significantly align with the framework’s rankings (p < 0.05). The associated dataset has been publicly released.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageHumans and AI: Learning Human Values and PreferencesCognitive Modeling & Cognitive Systems: Agent Architectures

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Large language models excel on objectively verifiable tasks such as math and programming, where evaluation reduces to unit tests or a single correct answer. In contrast, real-world enterprise work is often subjective and context-dependent: success hinges on organizational goals, user intent, and the quality of intermediate artifacts produced across long, multi-tool workflows. We introduce LH-Bench, a three-pillar evaluation design that moves beyond binary correctness to score autonomous, long-horizon execution on subjective enterprise tasks. The pillars are: (i) expert-grounded rubrics that give LLM judges the domain context needed to score subjective work, (ii) curated ground-truth artifacts that enable stepwise reward signals (e.g., chapter-level annotation for content tasks), and (iii) pairwise human preference evaluation for convergent validation. We show that domain-authored rubrics provide substantially more reliable evaluation signals than LLM-authored rubrics (kappa = 0.60 vs. 0.46), and that human preference judgments confirm the same top-tier separation (p < 0.05), evidence that expert-grounded evaluation can scale without sacrificing reliability. We release public datasets and report results on two environments: Figma-to-code (33 real .fig tasks against the Figma API via MCP) and Programmatic content (41 courses comprising 183 individually-evaluated chapters on a course platform serving 30+ daily users).
Problem

Research questions and friction points this paper is trying to address.

subjective evaluation
long-horizon tasks
enterprise workflows
LLM assessment
context-dependent tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

long-horizon evaluation
subjective tasks
expert-grounded rubrics
stepwise reward signals
human preference validation
🔎 Similar Papers
2023-08-22Frontiers Comput. Sci.Citations: 866
A
Abhishek Chandwani
I
Ishan Gupta