PhantomEnvironments: Training LLM Agents in Fictional Worlds

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scaling bottlenecks in training large language model (LLM) agents, including high environmental costs, hallucination susceptibility, and benchmark contamination. To overcome these challenges, this work proposes a multi-turn reinforcement learning environment set within a fictional world generated via rule-based synthesis. The approach enables zero-marginal-cost generation of training data without LLM involvement, utilizing templated article retrieval to guide agents through multi-hop question-answering tasks. Experimental results demonstrate that this synthetic environment substantially enhances transferability to real-world scenarios. Notably, agents trained exclusively on synthetic data outperform models trained with authentic data on real-world benchmarks. Furthermore, the proposed method elicits emergent search strategies that scale linearly with task difficulty.
📝 Abstract
Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environments yield agents that transfer to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity reveals that hop count drives transfer more than constraints or comparisons: even the simplest rule-generated environments are a surprisingly effective, free resource for training generalizable LLM agents.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
reinforcement learning
synthetic environments
multi-hop search
benchmark contamination
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Synthetic Environments
Multi-hop Search
Transfer Learning
Emergent Scaling
🔎 Similar Papers
No similar papers found.