🤖 AI Summary
This study addresses the neglect of content externalities in online ad auctions and the inherent difficulty of jointly achieving global welfare optimization with strict strategyproofness. We propose a learnable, globally optimized auction framework grounded in the Myersonian Ironed Revenue Transformation (MIRT) mechanism class. Methodologically, a Transformer generates candidate information flows, while hard attention combined with reinforcement learning enables bid-independent, high-welfare range selection, unifying candidate generation and bid-aware selection within a single training process. Theoretically, incentive compatibility is guaranteed via pseudo-dimension analysis. Empirically, the proposed approach significantly outperforms existing baselines in social welfare while preserving exact strategyproofness, thereby validating both its learnability and practical deployability.
📝 Abstract
Modern online platforms commonly rank ads and organic content separately before blending them into a feed displayed to the user, overlooking externalities: an item's click-through rate depends on its surrounding content, not only on its own position. Recent learning-based feed generation mechanisms model some of these cross-type interactions to globally optimize for the whole feed's welfare. However, these approaches either fix the ordering of organic content, or lack exact strategyproofness guarantees for bidders. To combat these shortfalls, we introduce the Maximal-in-Range Transformer (MIRT) mechanism class, which uses a transformer to generate a range of candidate feeds that jointly order ads and organic content, and selects the welfare-maximizing feed in the range. However, there is a tension: strategyproofness requires the generated range to be bid-independent, even though a candidate feed's welfare depends linearly on the bids. Our key technical contribution is a reinforcement learning approach that incorporates both candidate generation and bid-aware selection into training, enabling a bid-independent transformer to learn to generate high-welfare ranges by accounting for both individual feed quality and the collective quality of the range. Additionally, we bound the pseudo-dimension of the MIRT class under hard attention, showing that near-optimal expected welfare is learnable with sample complexity polynomial in the transformer size and only logarithmic in the range size. Empirically, MIRT outperforms the previous non-strategyproof state-of-the-art feed models while remaining exactly strategyproof. Our results show that transformer-based auctions can deliver externality-aware whole-feed optimization without sacrificing exact incentive compatibility, removing a major obstacle to their practical deployment.