Score
Designs and implements algorithms and systems that compute and emit per-update marginal contribution estimates for individual participants in an online or streaming setting, producing incremental reward signals without requiring global aggregation. Includes methods for conditioning those estimates on admissibility or other constraints so each emitted signal reflects an agent's marginal impact at each update.
This paper addresses slow wealth accumulation and gradient explosion—leading to overly conservative parameter updates—in nonparametric sequential hypothesis testing. We propose an online learning algorithm based on the optimistic interior-point method within a betting-based testing framework. The hypothesis test is formulated as a betting game, where evidence against the null hypothesis is quantified by online wealth accumulation; the null is rejected once wealth exceeds a predefined threshold. To our knowledge, this is the first work to integrate the optimistic interior-point method into the “testing-as-betting” paradigm, overcoming the conservatism of traditional Online Newton Step (ONS) approaches that require halving the decision space, thereby enabling safe, full-interior updates. By unifying online convex optimization with sequential probability ratio test theory, we rigorously control the Type-I error rate. Experiments demonstrate that our method significantly accelerates null rejection under the alternative hypothesis, reducing average stopping time by 30–50%.
This paper investigates competitive behavior among content creators (e.g., YouTube, TikTok influencers) under heterogeneous content quality in the attention economy. Addressing the scarcity of platform resources and user attention—both allocated proportionally to creators’ quality-weighted contributions—we propose the Proportional Payoff Allocation Game (PPA-Game), the first divisible-resource game incorporating heterogeneous quality weights. We prove the universal existence of pure Nash equilibria (PNE). Further, we integrate multi-player multi-armed bandits (MP-MAB) with online learning to design the first distributed learning algorithm achieving a logarithmic regret upper bound of $O(log^{1+eta} T)$. Theoretical analysis and stochastic simulations confirm both high PNE occurrence rates and substantial long-term payoff improvement. Our framework provides a provably optimal, dynamic decision mechanism for attention resource allocation.
This paper addresses the inefficiency of standard mechanisms (e.g., VCG) in multidimensional type environments where agents hold both private preferences and shared, uncertain state information affecting common values. To restore social efficiency, we propose a novel mechanism design framework that integrates posterior behavioral data (e.g., user feedback) into incentive-compatible allocation. Our key innovation is the first incorporation of a state estimator directly within a VCG-style mechanism, yielding a theory of implementation grounded in posterior equilibrium. The framework unifies three canonical settings: full revelation, affine utilities, and consistent estimation—achieving exact social optimality in the first two, and asymptotic optimality in the third, with estimation error decaying at an explicit rate as estimator accuracy improves. Methodologically, we bridge Bayesian mechanism design, state estimation theory, and VCG extensions. We validate the framework through formal models of digital advertising auctions and LLM-based human–AI interaction.
Existing incentive mechanisms fail to adapt to learning agents whose strategies evolve continuously over time. Method: We propose a two-timescale adaptive incentive mechanism driven by individual externality—the difference between an agent’s marginal cost and the operator’s marginal cost—and updated at a rate slower than the agents’ learning dynamics. Leveraging two-timescale stochastic approximation and differential game theory, we develop a technical framework comprising externality modeling, fixed-point analysis, and convergence proof. Contribution/Results: We establish the first general incentive framework decoupled from agents’ learning dynamics; rigorously guarantee that the Nash equilibrium coincides with the socially optimal solution; and unify treatment across atomic aggregative games and nonatomic routing games. We prove that every fixed point corresponds to an optimal incentive and derive sufficient conditions for global convergence. Numerical validation confirms these conditions hold and convergence is rapid in both canonical game settings.
This study addresses the pervasive issue of attribution bias in upper-funnel advertising, where incremental campaign effects are often misattributed to downstream channels—a phenomenon colloquially termed “assist miscredit”—leading to distorted ROAS metrics and flawed marketing mix models. To resolve this, the authors propose a novel individual-level incrementality-based measurement framework that embeds an intent-to-treat (ITT) experiment within real-world ad-exposed audiences. By extending the PIE (Probabilistic Impact Estimation) framework to the individual level and integrating audience-level natural randomization with machine learning–based response mapping, the method enables unbiased estimation of each channel’s true causal contribution to conversions. Crucially, it preserves the full conversion path while eliminating attribution bias, thereby delivering granular, actionable insights for optimal budget allocation.