Generative Optimization for Incentivized Advertising with Global Level Constraints

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of optimizing continuous incentive allocations in rewarded advertising under strict global return-on-investment (ROI) constraints, while accounting for non-Markovian dynamics such as high-frequency user interactions, delayed feedback, and user fatigue. The authors propose GOAL, a novel framework that formulates incentive allocation as a conditional sequence generation task. It employs a hierarchical causal state encoder to capture both local behavioral dynamics and long-range dependencies in user responses. To handle diverse ROI constraints without retraining, they introduce Safe Constrained Policy Optimization (SCPO), a single policy learning algorithm that generalizes across varying constraint thresholds. Experiments on large-scale real-world data and synthetic environments demonstrate that GOAL significantly improves long-term revenue and user retention while substantially reducing constraint violations.
📝 Abstract
Incentivized advertising allocates monetary or virtual rewards to drive user engagement, where a key challenge is optimizing continuous incentive magnitudes under strict global constraints. This problem is complicated by high-frequency interactions, delayed feedback, and non-Markovian user dynamics such as fatigue, which limit the effectiveness of existing uplift modeling and constrained reinforcement learning approaches. To address these challenges, we propose GOAL, a constraint-aware generative framework that formulates incentive allocation as a conditional sequence generation problem. GOAL directly generates incentive magnitudes conditioned on user histories and system-level global pressure, and integrates a hierarchical causal state encoder to capture both local behavioral dynamics and long-range dependencies. To enable flexible constraint control, we introduce \textbf{S}afe \textbf{C}onstrained \textbf{P}olicy \textbf{O}ptimization (SCPO), which learns a single generative policy that generalizes across a spectrum of ROI constraints without retraining. Experiments on large-scale real-world data and a synthetic fatigue-aware environment show that GOAL improves long-term revenue and user retention while substantially reducing ROI violation rates compared to strong baselines.
Problem

Research questions and friction points this paper is trying to address.

incentivized advertising
global constraints
incentive optimization
non-Markovian dynamics
delayed feedback
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Optimization
Incentivized Advertising
Global Constraints
Constrained Reinforcement Learning
Conditional Sequence Generation
🔎 Similar Papers
No similar papers found.