BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

๐Ÿ“… 2026-06-01
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the lack of a unified, production-ready evaluation framework for enterprise AI systems tackling both natural language-to-SQL translation and multi-step agent reasoning, which often fails to jointly assess execution accuracy and behavioral plausibility. To bridge this gap, we propose BADGERโ€”a cohesive framework integrating text-to-SQL evaluation with agent behavior assessment, supporting deployment in customer-managed data environments and configurable LLM-based judging backends for continuous evaluation. Key innovations include LLM-assisted SQL component extraction, a hybrid execution accuracy metric (Hybrid-EX), and an agent evaluation suite combining RAGAS and G-Eval while introducing a novel metric, Excess Tool Usage. Evaluated on 150 industrial queries, Hybrid-EX achieves a Cohenโ€™s ฮบ of 0.717 and 87.3% balanced accuracy, significantly outperforming six baselines (p โ‰ค 0.001).
๐Ÿ“ Abstract
Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approaches fundamentally different from academic benchmarks. Spider and BIRD established execution-accuracy protocols; G-Eval and RAGAS advanced LLM-based assessment; and recent work such as Spider 2.0, BEAVER, and BIRD-Interact has begun to address enterprise and agentic dimensions. No single framework unifies text-to-SQL assessment with agentic behavior evaluation into a production-grade pipeline calibrated against human expert judgment. We present BADGER, developed at Merkle, a unified evaluation framework integrating text-to-SQL assessment with agentic behavior evaluation. BADGER offers three contributions. First, LLM-assisted SQL component extraction extending Spider methodology to handle CTE-heavy, dialect-specific SQL. Second, a hybrid execution accuracy metric (Hybrid-EX) resolving column-aliasing and numeric-tolerance brittleness by using an LLM to infer structural alignments before deterministic cell-level scoring. Validated on 150 human-annotated industry queries, Hybrid-EX achieves Cohen's kappa=0.717 [95% CI: 0.600-0.822] (Substantial agreement) and 87.3% balanced accuracy, outperforming all six competing frameworks (Delta-kappa: 0.322-0.502, all p<=0.001). Third, an enterprise agentic evaluation suite assembling RAGAS, G-Eval, and agent benchmark metrics into a unified pipeline; Excess Tool Usage is the sole novel element. BADGER runs entirely within the client's governed data environment, supports configurable LLM judge backends, and enables rapid prototyping of client-specific judges and metrics, serving as a continuous evaluation backbone rather than a one-time quality gate.
Problem

Research questions and friction points this paper is trying to address.

text-to-SQL
agentic reasoning
enterprise AI evaluation
execution accuracy
LLM-based assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid-EX
Agentic Evaluation
Text-to-SQL
Enterprise AI
LLM-assisted SQL Parsing
S
Shannon Serrao
Merkle Analytics
S
Soumitra Chatterjee
Merkle Analytics
D
Dorina Strori
Merkle Analytics
A
Abhishek Sharma
Merkle Analytics
Nathan Miller
Nathan Miller
Professor, Georgetown University
Industrial OrganizationAntitrust Economics