Witness: Discovery, Deciphering, and Epiphany in Interactive Puzzle Environments

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the capability bottlenecks of large language models in autonomously discovering unknown rules within interactive environments. To this end, it introduces WITNESS, a benchmark featuring a 2D grid-based interactive environment that decouples compositional rule generalization from novel primitive discovery tasks. The work further designs an agent generation pipeline integrating reinforcement learning training with a multi-model comparative evaluation framework. Empirical findings reveal that rule acquisition constitutes a core challenge for frontier models, as the best-performing model solves only 24% of private test instances, whereas providing ground-truth rules yields substantial performance improvements. Moreover, reinforcement learning doubles the problem-solving efficiency of smaller models and successfully transfers to external benchmarks.
📝 Abstract
Automated science needs agents that can work out the rules of an unfamiliar environment by interacting with it. Interactive rule-discovery puzzles offer a controlled setting for studying this ability: an agent infers hidden rules through experimentation and uses what it has inferred to reach a stated goal. We ask what limits current language models on these puzzles and whether reinforcement learning (RL) improves performance on rules held out from training. To study both, we introduce WITNESS, a 2D grid-based puzzle environment with ground-truth ASCII observations and controlled access to rules. An agentic pipeline generates games for WitnessGym, the RL training suite, and WitnessBench, comprising public validation and private test games. The validation set separately tests new compositions of trained rule primitives and primitives absent from training. Under a shared harness, the best of 18 frontier proprietary and open-weight models solves only 24\% of private test level slots, with scores sensitive to the observation interface and agent configuration. Providing ground-truth rules raises Opus-5's validation RHAE-L5 (relative human action efficiency over the first five levels) from 59.9 to 97.8, whereas a 27B open-weight model gains only 2.1 points and remains limited even with the rules provided. RL on WitnessGym raises the 27B model's private test RHAE-L5 from 2.1 to 5.4 and yields a mean gain of 4.1 points on four external discovery benchmarks. Together, these results point to rule acquisition as a major difficulty for frontier models like Opus-5 while smaller models further struggle on rule-based execution, and indicate that RL on hidden-rule puzzles transfers to broader rules and real-world tasks beyond training. Benchmark is available at: https://witnessbench.ai
Problem

Research questions and friction points this paper is trying to address.

rule discovery
interactive puzzle environments
language model limitations
reinforcement learning generalization
automated science
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interactive Rule Discovery
Reinforcement Learning
Puzzle Environment
Agentic Pipeline
Generalization
🔎 Similar Papers
2024-03-15arXiv.orgCitations: 2