Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of AI leaderboards to adaptive manipulation attacks, which compromise ranking credibility. To mitigate this, it proposes a certified corruption budget mechanism that introduces the first anytime-valid certificates. By integrating statistical inference, game-theoretic mechanism design, and robust confidence interval techniques, the approach distinguishes between fabrication and tampering scenarios, providing statistical guarantees for pairwise claims at arbitrary times against unbounded adversaries. The contributions include achieving optimal growth rates alongside privacy-preserving variants. Empirical evaluations on Chatbot Arena demonstrate that the proposed method remains effective where conventional approaches fail, confirming the system’s capacity to withstand approximately 2,000 fabricated votes.
📝 Abstract
Public leaderboards for AI models are read continuously, and attackers can see every published standing. Vote rigging, selective disclosure of private variants, and benchmark contamination can each move a ranking. Existing guarantees assume genuine records or bound the corruption per step, which an attacker who corrupts in bursts evades. We introduce the certified corruption budget, a tolerance $\widehat{B}_t$ computed after $t$ records and published with each pairwise claim. With probability at least $1-α$, simultaneously at all times, the claim is correct or more than $\widehat{B}_t$ records were corrupted. It holds against attackers who watch every certificate, with no bound on their budget. Forged records and records altered once seen require different certificates: the certificate for forgeries fails, with probability approaching one, against an attacker who flips votes it has seen, while one that charges roughly twice as much per record remains valid, with constant bets even against attackers who see the future, and no smaller charge is valid at every level. The certified budget grows nearly as fast as any valid method allows: with a win fraction $\frac{1}{2}+δ$, each new record adds close to $2δ$ to the number of forged records the claim can withstand ($δ$ flipped). Publishing the best of $V$ private variants costs only an amount growing like $\log V$. In replays on 1.8 million Chatbot Arena votes, a few hundred rigged votes make standard confidence intervals certify false orderings, while ours stays valid. On real votes, our certificate shows that clearly separated models withstand about 2,000 forged votes.
Problem

Research questions and friction points this paper is trying to address.

leaderboard integrity
vote rigging
adaptive attacks
certified corruption budget
benchmark contamination
Innovation

Methods, ideas, or system contributions that make the work stand out.

Certified Corruption Budget
Anytime-Valid Inference
Adaptive Rigging
Leaderboard Integrity
Vote Manipulation
🔎 Similar Papers
No similar papers found.