SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?

📅 2025-12-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current AI agents lack standardized evaluation for extracting biological insights from real-world spatial omics data. Method: We introduce SpatialBench—the first benchmark for spatial biology—comprising 146 verifiable questions across five experimental platforms and seven analytical task types. We propose the first systematic evaluation paradigm for spatial biology agents, emphasizing task-platform coupling and identifying Harness—a unified framework integrating tool orchestration, prompt engineering, control-flow logic, and execution environment—as the primary determinant of agent performance. Our implementation leverages multimodal large language models, custom toolchains, deterministic automated scoring, and realistic data workflow modeling. Results: State-of-the-art models achieve only 20–38% accuracy on SpatialBench; however, targeted Harness optimization yields substantial performance gains. SpatialBench establishes a reproducible, transparent, and diagnosable standard for evaluating and iteratively improving spatial biology agents.

Technology Category

Knowledge Representation and Reasoning: Geometric, Spatial, and Temporal ReasoningCognitive Modeling & Cognitive Systems: Agent ArchitecturesPlanning, Routing, and Scheduling: Optimization of Spatio-temporal Systems

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
📝 Abstract
Spatial transcriptomics assays are rapidly increasing in scale and complexity, making computational analysis a major bottleneck in biological discovery. Although frontier AI agents have improved dramatically at software engineering and general data analysis, it remains unclear whether they can extract biological insight from messy, real-world spatial datasets. We introduce SpatialBench, a benchmark of 146 verifiable problems derived from practical spatial analysis workflows spanning five spatial technologies and seven task categories. Each problem provides a snapshot of experimental data immediately prior to an analysis step and a deterministic grader that evaluates recovery of a key biological result. Benchmark data on frontier models shows that base model accuracy remains low (20-38% across model families), with strong model-task and model-platform interactions. Harness design has a large empirical effect on performance, indicating that tools, prompts, control flow, and execution environment should be evaluated and improved as first-class objects. SpatialBench serves both as a measurement tool and a diagnostic lens for developing agents that can interact with real spatial datasets faithfully, transparently, and reproducibly.
Problem

Research questions and friction points this paper is trying to address.

Evaluating AI agents' ability to analyze messy real-world spatial transcriptomics data.
Benchmarking performance across diverse spatial technologies and analysis tasks.
Identifying factors like harness design that affect agent reliability and reproducibility.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmark with 146 verifiable spatial analysis problems
Evaluates AI agents on messy real-world spatial biology data
Focuses on tools, prompts, control flow as first-class objects
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
LatchBio
K
Kenny Workman
LatchBio, San Francisco, CA
Z
Zhen Yang
LatchBio, San Francisco, CA
H
Harihara Muralidharan
LatchBio, San Francisco, CA
H
Hannah Le
LatchBio, San Francisco, CA