What Does an LLM-Agent Leaderboard Rank Actually Compare?

📅 2026-09-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了LLM-agent排行榜排名的实际意义,通过定义比较目标、检查共同支持并使用不确定性规则评估差异,揭示了排名相近时的不确定性及标签和规则对系统选择的影响。
📝 Abstract
An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent. Public evaluation logs may not support that conclusion when systems differ in task mixture, label source, release detail, or cost rule. We study what leaderboard scores estimate and when they justify pairwise superiority conclusions. Our estimand-aware pairwise procedure states the comparison target and measurement source, checks common support, and evaluates the supported difference using a stated uncertainty rule and practical margin. Controlled checks evaluate the decision labels under known finite-sample conditions and show why uncertainty must be included when judging sensitivity to target reweighting. Across SWE-bench, AgentRewardBench, and tau2-bench, close rank differences are often unresolved; proxy labels and utility rules can also change which system is selected. DataAgentBench and Open Agent show what remains estimable from coarser public records. A leaderboard score summarizes a released evaluation, whereas a fine-grained superiority claim additionally depends on the estimand and uncertainty rule used to interpret the difference.
Problem

Research questions and friction points this paper is trying to address.

LLM-agent leaderboard
pairwise superiority
estimand-aware
Innovation

Methods, ideas, or system contributions that make the work stand out.

leaderboard scores
pairwise comparison
estimand-aware procedure
uncertainty rule
common support
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wei-Jung Huang
Independent Researcher