WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of large language model (LLM) agents in generating physics solvers for simulating dynamical systems by introducing the first benchmark tailored to this task. Drawing upon foundational literature in computer graphics, we construct an evaluation suite comprising 168 tasks that assess agents' comprehensive ability to synthesize executable solvers from code scaffolding, encompassing physical understanding, mathematical reasoning, and software engineering. Furthermore, we define a three-dimensional evaluation framework incorporating execution checks, visual fidelity, and physical plausibility. Experimental results demonstrate that state-of-the-art models achieve an overall score of only 48.7%, revealing that generating accurate and physically consistent solvers remains a formidable challenge.
📝 Abstract
LLM-based agents are increasingly advancing scientific and engineering problem solving, with physics simulation emerging as a challenging yet practical testbed for reproducing complex physical phenomena with application in embodied AI, games and films. As the workhorse of such simulation, a solver computes how the state of a dynamic system evolves over time. Building such solvers requires physical understanding to identify appropriate models, mathematical reasoning to formulate the underlying dynamics, and software engineering to implement them as executable code, yet this capability of LLM agents remains underexplored. To this end, we introduce WorldSolver, a benchmark of 168 simulation tasks derived from physical phenomena in 61 classic computer graphics papers, spanning 7 physical domains. Each task contains a code scaffold that provides a fixed simulation environment for the scene, with the solver implementation left for the agent to complete. Specifically, we evaluate them along three dimensions: Execution Checks for successful execution, Visual Fidelity for reproducing the intended dynamic behavior in the rendered simulation, and Physical Plausibility for physics-grounded verification of the generated dynamics. Experiments on frontier agents reveal that producing executable solvers is difficult itself, and satisfying visual and physical correctness is even harder. GPT-5.6-Sol and Claude-Opus-5 perform comparatively better than the other evaluated agents, yet achieve overall scores of only 48.7% and 46.7%, respectively. WorldSolver is an early step toward agentic solver generation, and we hope it helps drive progress toward agents that can faithfully simulate the dynamic physical world. Code is available at https://github.com/sirujiang/WorldSolver.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
physics simulation
solver generation
physical dynamics
benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Agents
Physics Simulation
Solver Generation
Benchmark
Code Scaffolding
🔎 Similar Papers
No similar papers found.