EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of comprehensive evaluation benchmarks for LLM agents in enterprise email workflows by constructing a self-contained assessment environment comprising 206 scenarios. Methodologically, it proposes a hybrid evaluation protocol integrating static assertions with LLM-based scoring. Leveraging a deterministic synthetic corpus and typed API specifications, the framework systematically examines agent performance in retrieval, reasoning, and multi-step coordination. This work fills the gap in evaluating typed email APIs and reveals that successful tool invocation does not equate to task completion. Experimental results demonstrate that while the optimal configuration achieves a 99.7% error-free tool call rate, the scenario pass rate reaches only 33.5%, confirming substantial capability gaps in current agents.
📝 Abstract
Enterprise email agents must combine information retrieval, structured state changes, temporal reasoning, and multi-step coordination. Recent agent benchmarks include productivity tasks, but few center on typed email workflows in a self-contained environment. We introduce EmailBench, a benchmark of 206 email and productivity scenarios across 16 task categories. The benchmark couples a typed email API specification with provider-neutral naming, a deterministic synthetic Enron-inspired corpus, and a scenario suite whose topic selection was informed by aggregate task-intent telemetry from an interactive prototype. Its hybrid evaluation protocol combines 258 executable static assertions with 211 LLM rubrics. We evaluate eight LM configurations on a fixed single-user corpus. The best-performing configuration passes only 33.5% of scenarios despite 99.7% of its tool calls completing without an observed API failure, with pass rates varying substantially across task categories. This gap shows that valid tool execution is not equivalent to task completion. EmailBench provides a self-contained environment for end-to-end email-agent evaluation, with broader tool coverage, multi-persona testing, and repeated-run evaluation as future work areas.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
email benchmark
task completion
tool execution
enterprise productivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Agents
Benchmark
Email Workflows
Hybrid Evaluation
Tool Use
🔎 Similar Papers
No similar papers found.