Sharpening Tax in Post-Training

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the observation that reinforcement learning-based post-training, while improving single-pass inference accuracy of large language models, polarizes the agent task solution space and severely compromises solution coverage during test-time scaling. To quantify this trade-off, we introduce the "Sharpening Tax" metric and develop a Bayesian optimization-based adaptive temperature sampler (PTGS) that dynamically adjusts sampling strategies to balance single-pass accuracy with multi-round exploration diversity. Experiments across 42 cases validate the ubiquity of the Sharpening Tax and demonstrate that PTGS effectively mitigates this cost, simultaneously achieving significant improvements in both task success rates under repeated sampling and single-pass inference accuracy.
📝 Abstract
An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning Post-Training
Sharpening Tax
Solution Coverage
Test-time Scalability
Agentic Tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sharpening Tax
Posterior-Tempered Group Sampling
Test-Time Scalability
Reinforcement Learning Post-Training
Agentic Tasks
🔎 Similar Papers
No similar papers found.