Inference Auctions

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of computationally constrained LLM inference and the inability of fixed pricing to accommodate users' heterogeneous latency tolerances by proposing a cost-effective inference auction mechanism. Methodologically, we design an auction model enabling users to bid for prioritized service, develop a prompt pricing algorithm that incentivizes truthful bidding, and construct an automated bidding agent that maximizes utility under budget constraints. The proposed framework is integrated into SGLang to balance low latency with high cache utilization. Experimental results demonstrate that our mechanism preserves the performance advantages of SGLang while significantly improving overall system social welfare.
📝 Abstract
When inference demand exceeds available compute capacity, model providers must decide which requests should be served first. Users have different tolerances for delay from an LLM API, but current priority pricing schemes compress these differences into coarse fixed-price service tiers. We design an inference auction that allows users to bid for faster service. Our auction allocates priority in an economically efficient way without sacrificing latency, and we develop fast algorithms for implementing prices that incentivize truthful bidding. We also design an autobidding agent for our inference auction, where users specify an inference budget and the autobidder dynamically adjusts its bids over time to maximize user utility subject to the budget constraint. Experiments validate the practicality of our auction: it increases system welfare while maintaining the cache utilization and latency advantages of SGLang, a state-of-the-art inference serving framework.
Problem

Research questions and friction points this paper is trying to address.

inference serving
priority pricing
resource allocation
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Inference Auction
Mechanism Design
Autobidding Agent
Priority Allocation
LLM Serving
🔎 Similar Papers