EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of existing benchmarks for evaluating enterprise-level long-horizon, dynamic, and decision-making capabilities under uncertainty by constructing a comprehensive evaluation framework spanning from static question answering to dynamic decision-making. The benchmark incorporates three interactive scenarios: consulting, the Beer Game, and digital twins. Methodologically, it integrates large language model agents with classical supply chain simulations and enterprise project modeling techniques, utilizing uniformly restructured data alongside multi-turn information retrieval and delayed feedback mechanisms to simulate authentic decision environments. This work bridges the gap in evaluating interactive strategic reasoning, reveals the limitations of current agents in cross-task reliability, and establishes practical evaluation standards for enterprise-level decision-making.
📝 Abstract
LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as information extraction, numerical calculation, domain knowledge, and financial QA, leaving interactive and long-horizon decision-making underexplored. To bridge this gap, we introduce EnterpriseBench, a benchmark that evaluates LLM agents across this spectrum, from static question answering to dynamic decision-making. Specifically, EnterpriseBench reorganizes existing enterprise and financial QA datasets into a unified foundational suite annotated by capability and difficulty, and introduces three professional interactive settings: Consulting, based on management-consulting-style business cases for client problem diagnosis through multi-turn information seeking; the Beer Game, adapted from a classic supply-chain management simulation for inventory control under delayed feedback; and Enterprise Digital Twin, a project-based business simulator for workforce, risk, and project planning. Experiments with nine agent methods under four backbone models show that current agents have not yet achieved stable, comprehensive, and cross-task reliability in enterprise scenarios. These results show that EnterpriseBench provides a practical benchmark for evaluating LLM agents in realistic enterprise strategic reasoning and decision-making.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
enterprise benchmarking
strategic reasoning
decision-making
long-horizon planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Agents
Enterprise Benchmark
Strategic Reasoning
Interactive Decision-Making
Digital Twin
🔎 Similar Papers
No similar papers found.