Meritocratic Fairness in Budgeted Combinatorial Multi-armed Bandits via Shapley Values

📅 2026-05-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of ensuring fairness in full-bandit feedback combinatorial multi-armed bandits, where individual arm contributions are unobservable. To this end, the authors propose the K-Shapley value—a unique extension of the Shapley value under combinatorial constraints—that satisfies symmetry, linearity, null player, and efficiency axioms. Building upon this notion, they design the K-SVFair-FBF algorithm, which simultaneously learns an unknown valuation function and achieves merit-based fairness. Notably, this is the first algorithm to jointly guarantee fairness and robustness to noise under full-bandit feedback. Theoretically, it achieves a fairness regret upper bound of $O(T^{3/4})$. Empirical evaluations demonstrate that K-SVFair-FBF significantly outperforms existing baselines in federated learning and social influence maximization tasks, effectively balancing fairness and utility.
📝 Abstract
We propose a new framework for meritocratic fairness in budgeted combinatorial multi-armed bandits with full-bandit feedback (BCMAB-FBF). Unlike semi-bandit feedback, the contribution of individual arms is not received in full-bandit feedback, making the setting significantly more challenging. To compute arm contributions in BCMAB-FBF, we first extend the Shapley value, a classical solution concept from cooperative game theory, to the $K$-Shapley value, which captures the marginal contribution of an agent restricted to a set of size at most $K$. We show that $K$-Shapley value is a unique solution concept that satisfies Symmetry, Linearity, Null player, and efficiency properties. We next propose K-SVFair-FBF, a fairness-aware bandit algorithm that adaptively estimates $K$-Shapley value with unknown valuation function. Unlike standard bandit literature on full bandit feedback, K-SVFair-FBF not only learns the valuation function under full feedback setting but also mitigates the noise arising from Monte Carlo approximations. Theoretically, we prove that K-SVFair-FBF achieves $O(T^{3/4})$ regret bound on fairness regret. Through experiments on federated learning and social influence maximization datasets, we demonstrate that our approach achieves fairness and performs more effectively than existing baselines.
Problem

Research questions and friction points this paper is trying to address.

Meritocratic Fairness
Budgeted Combinatorial Multi-armed Bandits
Full-bandit Feedback
Shapley Values
Fairness in Bandits
Innovation

Methods, ideas, or system contributions that make the work stand out.

K-Shapley value
meritocratic fairness
budgeted combinatorial multi-armed bandits
full-bandit feedback
fairness-aware learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shradha Sharma
Indian Institute of Technology Ropar
Swapnil Dhamal
Swapnil Dhamal
Indian Institute of Technology Ropar
Game TheorySocial NetworksTransport PlanningBlockchainMulti-Armed Bandits
S
Shweta Jain
Indian Institute of Technology Ropar