DAYJOB: A Benchmark for Long-Horizon Professional Work

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear capacity of current AI systems to handle ambiguous requirements and verify premises in long-horizon professional tasks within domains such as healthcare and finance. To investigate this, we construct a benchmark comprising 130 real-world scenario tasks and deploy it within a Harbor-based containerized environment. Furthermore, we introduce the first fine-grained binary evaluation framework tailored for extended professional workflows, integrating expert rubrics with an agentic judge to enable rigorous automated assessment. Our findings reveal a critical vulnerability: models tend to accept flawed premises without adequate verification. Empirical results demonstrate that even the most capable model achieves a pass rate of only approximately 24%, while the median performance falls below 3%. These findings substantiate that contemporary AI systems remain insufficiently equipped to autonomously execute complex, long-horizon professional work.
📝 Abstract
Professional work often starts with a brief request that leaves the professional to work out what is needed, which documents matter, and whether the request's premise holds. We introduce DAYJOB, a benchmark of 130 tasks built by professionals in healthcare (50) and finance (80). The tasks are estimated to take a professional 13.6 hours on average in healthcare and 16.6 in finance. Each task is a containerized Harbor environment with an expert rubric of binary criteria (median 47.5 and 57.5 per task) that an agentic judge applies to the delivered files, and an attempt passes only if it meets every criterion. Across 30 model configurations from 13 developers, the strongest, Claude Opus 5.5, passes 24.7% of healthcare and 23.9% of finance attempts, and the median configuration passes 0.6% and 2.5%. In case studies, agents accept premises that the record contradicts and carry wrong inputs through otherwise consistent analyses. We release all healthcare tasks, 50 of the 80 finance tasks, the evaluation harness, and the leaderboard.
Problem

Research questions and friction points this paper is trying to address.

long-horizon tasks
professional work
AI agents
benchmark evaluation
complex reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-Horizon Benchmark
Agentic Judge
Containerized Environment
Professional Work Evaluation
Expert Rubric
S
Stephanie Finley
Liudas Panavas
Liudas Panavas
PhD Student at Northeastern University
data visualizationhuman computer interaction
T
Thomas Mikkelson
C
Cam Hinton
S
Stacey Ganss
B
Bradley Monton
E
Emily Kendall
M
Michelle Spradlin
L
Lydia Bye
M
Michael O'Brien
L
Lauren Ylvisaker
D
Derek Ray
S
Suhaas Garre
S
Sushant Mehta
E
Edwin Chen