Reinforcement Learning for Execution under Dynamic Fees in a Closed-Loop DEX Simulator

📅 2026-07-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of empirical evidence for optimal execution in automated market makers (AMMs) under dynamic trading fees. The authors propose a closed-loop simulator grounded in equilibrium theory, incorporating constant-product liquidity pools with dynamic fees, fee-sensitive noise trading flows, and an analytical arbitrage mechanism, thereby enabling counterfactual analysis in a controlled environment. Using deep Q-networks (DQN), they compare reinforcement learning against multiple benchmark strategies and find that DQN significantly outperforms traditional scheduling and planning methods only in dynamic-fee settings. Across 1,000 held-out random seeds, DQN consistently reduces execution shortfall across all order sequencing configurations, achieving a 13.3 basis point improvement in the agent-last setting, while showing no significant advantage under constant fees.
📝 Abstract
Trader-facing dynamic fees are increasingly proposed for automated market makers (AMMs), but historical data do not identify how order flow would respond: trader-facing fees do not vary, trader types are latent, and a replayed tape is not a sequential decision environment. We therefore construct a minimal closed-loop simulator in which the missing signal exists by construction: two constant-product pools repriced by an equilibrium-inspired dynamic-fee rule, fee-sensitive noise flow, and closed-form CEX--AMM arbitrage. Equilibrium is used as a closure principle, not as an object the trader learns. Against a tuned benchmark ladder of schedule, planning, lookahead, and tabular policies, a small DQN is the only evaluated valid policy whose paired improvement over tuned one-step routing excludes zero. On a reserved final block of 1{,}000 seeds with completion forced to 1.0 for every policy, it reduces implementation shortfall under every tested intra-step ordering, by $13.3\bps$ of order notional under the pre-specified agent-last ordering, and the edge is concentrated in, and learned from, dynamic-fee environments: under constant fees the paired difference is indistinguishable from zero. The result is model-conditioned counterfactual evidence about execution control in AMMs, not evidence about historical traders, equilibrium play, or deployable profit.
Problem

Research questions and friction points this paper is trying to address.

dynamic fees
automated market makers
order flow response
reinforcement learning
execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Dynamic Fees
Automated Market Makers
Closed-Loop Simulation
Execution Optimization
🔎 Similar Papers
No similar papers found.