MESHA: Mechanism-Enforced Sequential Halving for Strategic Linear Bandits

📅 2026-07-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the issue of strategic manipulation in linear bandits, where arms may misreport their feature vectors to increase selection probability. To counter this, the authors propose the MESHA algorithm, which integrates uniform sampling with a round-wise “grim trigger condition” (GTC) to effectively suppress strategic behavior and eliminate severely distorted reports. The study establishes, for the first time, reliability guarantees for best-arm identification (BAI) under strategic linear settings, proving theoretically that conventional optimal-design-based sampling methods can entirely overlook the true best arm in such scenarios. Through Nash equilibrium analysis and derivation of an upper bound on the failure probability under fixed-budget constraints, the paper demonstrates that MESHA remains effective across all Nash equilibria. Empirical results confirm its significant superiority over existing baseline methods.
📝 Abstract
We design and analyze \underline{M}echanism-\underline{E}nforced \underline{S}equential \underline{HA}lving (MESHA), an algorithm for Best Arm Identification (BAI) in strategic linear bandits. In this setting, each arm may strategically misreport its feature vector to maximize the probability of being identified as the best arm, when rewards are generated from the arms' true but unobservable features. The design of MESHA applies the naïve uniform sampling rule and an epoch-wise Grim Trigger Condition (GTC): the former reduces the impact of arms' strategic behaviours and the latter eliminates arms whose reported features severely deviate from the ground truth. Considering an arbitrary Nash Equilibrium, we prove that any arm would attempt to pass the GTC check to maximize its identified probability and derive an upper bound on the failure probability of MESHA within a fixed budget $T$. We also show that state-of-the-art linear BAI algorithms with $G$-optimal design would fail in such strategic environment, as the optimal design (OD)-based sampling rule based on strategically reported features may {\it starve} the optimal arm of any sampling budget. Finally, extensive numerical experiments indicate that MESHA outperforms baselines that rely on OD-based sampling rules as well as the feature-agnostic baselines, corroborating the efficacy of MESHA.
Problem

Research questions and friction points this paper is trying to address.

Strategic Linear Bandits
Best Arm Identification
Feature Misreporting
Nash Equilibrium
Optimal Design
Innovation

Methods, ideas, or system contributions that make the work stand out.

strategic linear bandits
best arm identification
mechanism design
Grim Trigger Condition
uniform sampling
🔎 Similar Papers
2024-07-24arXiv.orgCitations: 4
💼 Related Jobs
No related jobs found.
X
Xin Li
Data Science and Analytics Thrust, Hong Kong University of Science and Technology (Guangzhou)
Z
Zixin Zhong
Data Science and Analytics Thrust, Hong Kong University of Science and Technology (Guangzhou)