Do Small Language Models Learn to Negotiate? A Controlled Scaling Study of RL-Trained Sellers

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether small language models can acquire negotiation capabilities through reinforcement learning and examines the associated scaling effects. Employing the GRPO algorithm with programmatic utility rewards, we train Gemma-based seller models of varying parameter sizes in bilateral multi-issue negotiation scenarios. Our results demonstrate that learning rate exerts a critical influence on the negotiation performance of smaller models. With optimized hyperparameters, a 4.5B-parameter model achieves performance comparable to frontier large language models while requiring only single-GPU deployment. This work challenges the prevailing assumption that small models are inherently incapable of learning to negotiate, confirming that meticulous hyperparameter tuning can substantially enhance their competitive negotiation proficiency.
📝 Abstract
LLM agents are starting to own the full customer experience. Soon, LLMs may be selling and buying on behalf of companies and customers respectively. Small models are more cost-efficient at scale, but can reinforcement learning train them into competent sellers? We train four Gemma 4 checkpoints (2.3B to 31B effective parameters) with GRPO on a programmatic utility reward for bilateral multi-issue bargaining, and evaluate every arm on the same 1,152 negotiations against two frontier buyers it never saw in training. With the same learning rate ($10^{-6}$) for every size, the gain of the RL model over its base rises from $+0.001$ at 2.3B to $+0.078$ at 31B. Each size was trained once and the two smallest checkpoints use a different architecture, so we fit no scaling law. Tripling the learning rate, with the same or fewer training steps, improves on the shared rate at every size by $+0.032$ (2.3B) to $+0.081$ (4.5B). In exploratory comparisons with two frontier models run as sellers, the 12B seller trained at the tripled rate scores above both, though its untrained base already scores as high as they do. The 4.5B seller at that rate shows no detectable difference from either and fits on one 48 GB GPU. A further 2.3B arm at ten times the shared rate raises pooled score, but its gain concentrates on the evaluation buyer that shares a model family with the training pool. These results suggest tuning the learning rate before concluding that a small model cannot learn to negotiate, and testing against buyers from more than one model family.
Problem

Research questions and friction points this paper is trying to address.

Small Language Models
Reinforcement Learning
Negotiation
Bilateral Bargaining
Scaling Study
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Small Language Models
GRPO
Automated Negotiation
Scaling Study
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.