Pricing Time, Not Just Tokens: Latency-Aware Mechanism Design for LLM Inference

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency in existing LLM inference pricing, which treats tokens as homogeneous commodities and neglects user latency preferences. To overcome this, we construct a service market model incorporating time preferences, leveraging hardware tiering to enable user self-selection and thereby reducing three-dimensional screening to one-dimensional mechanism design. By integrating virtual value techniques with physical GPU modeling under compute- and bandwidth-bound regimes, we empirically calibrate our framework using H100/B200 cluster data. We propose a separation theorem proving that discrete hardware tiers endogenously resolve time-preference heterogeneity, providing theoretical justification for flat API pricing. Experiments demonstrate that this theorem holds across 83% of configurations, and that two-tier pricing increases profits by 26–66% over single-tier schemes while significantly optimizing resource allocation.
📝 Abstract
The economic theory of LLM pricing treats tokens as a homogeneous commodity considering aggregate token count as the main features buyers and sellers consider. We model inference as a service market where buyers have three-dimensional private information - willingness-to-pay, task volume, and time preference - and utility depends on latency slack alongside token quantities. Our main result is a separation theorem: discrete hardware tiers induce endogenous self-selection on time preferences, reducing three-dimensional screening to standard one-dimensional screening within each tier. We derive the cost structure from GPU inference physics - compute-bound prefill and bandwidth-bound decode - and characterize optimal tiered mechanisms via virtual-value techniques. Optimal per-task prices are volume-independent, providing theoretical grounding for flat per-token API pricing. We verify the mechanism empirically by calibrating to 8-GPU clusters of H100 and B200 hardware. The separation theorem holds in 83% of 105 tested configurations overall, rising to 96% at economically relevant WTP scales. A seller adopting two-tier pricing under the optimal mechanism captures 26-66% higher profit than the best single-tier alternative, with gains driven by efficient cross-tier allocation in regimes where hardware costs are a significant fraction of per-request value.
Problem

Research questions and friction points this paper is trying to address.

LLM inference pricing
latency-aware mechanism design
time preference
hardware tiering
service market
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latency-Aware Pricing
Mechanism Design
Separation Theorem
LLM Inference
Tiered Hardware
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.