GlyphBench: A Playground for Language-Model Reinforcement Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of efficient and reproducible environments for reinforcement learning (RL) post-training of language model agents by proposing GlyphBench, a Unicode grid-game benchmark comprising over 360 tasks. This environment provides a unified evaluation interface to systematically investigate how observation modalities and RL configurations affect agent performance. Methodologically, fine-tuning upon Qwen3.5, this work introduces a characterized spatial observation representation and demonstrates its superiority over conventional textual and pixel-based inputs. Experimental results show that the proposed approach achieves 63.48% accuracy on Reasoning Gym, significantly outperforming baselines. Furthermore, it reveals that game-based reasoning exhibits stronger transferability than mathematical or coding tasks, establishing a novel paradigm for the RL post-training of intelligent agents.
📝 Abstract
We introduce GlyphBench, an environment suite for reinforcement learning (RL) post-training of language-model agents, with over 360 tasks spanning diverse games. GlyphBench renders spatial observations as two-dimensional Unicode grids and connects training, evaluation, and trajectory replay through a unified interface designed to support efficient and reproducible research. We use GlyphBench to study how observation interfaces, reasoning effort, and agent harnesses affect performance, and how RL configurations shape learning dynamics. Our results show that glyph observations outperform native text and pixels in our Craftax experiments, with further gains on several BALROG environments. RL on 100 GlyphBench tasks improves Qwen3.5-4B on held-out Reasoning Gym problems, reaching 63.48% accuracy and outperforming the base model, a math-trained baseline, and a code-trained baseline. These experiments provide empirical evidence that reasoning gains from gameplay can yield stronger transfer than math or code. Together, these results highlight GlyphBench's value as a testbed for systematic research on how language-model agents learn, interact, and generalize.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Language Model Agents
Post-training
Generalization
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Language Model Agents
Glyph Observations
Transfer Learning
Reasoning