The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

๐Ÿ“… 2026-09-24
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the limitation of existing benchmarks in evaluating agents' assistant capabilities for multi-step retrieval, synthesis, and final artifact generation. To this end, it proposes KNOWS, a benchmark that assesses agents' reasoning decomposition and visuospatial understanding through open-ended browser tasks. Methodologically, systematic task design rules are formulated, and a hybrid evaluator combining deterministic checks with large language model judgments is introduced to balance evaluation richness and automation. Experimental results demonstrate that frontier agents achieve a complete success rate of less than 3% on complex tasks, revealing significant limitations in tool invocation, visual comprehension, and long-horizon reasoning.
๐Ÿ“ Abstract
Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to produce a coherent final product. Such workflows demand reasoning and synthesis, decomposition of complex tasks, as well as visual and spatial understanding. To study agents on workflows like these, we introduce KNOWS, a benchmark of open-ended, complex, browser-based tasks that jointly evaluate these capabilities, with each task culminating in a produced artifact. To write tasks, we develop a task design rubric and a protocol for ensuring that tasks meet the requirements. Each task is paired with an evaluator, a program that combines deterministic checks with LLM judgments to balance the richness, reliability, and automation tradeoff inherent to agent evaluation. We evaluate and analyze frontier computer-use agents and browser-based harnesses. They achieve moderate scores on partial-success metrics, but the best performer fully succeeds in fewer than 3% of our complex, long-horizon tasks. Failures on visual steps render the resulting artifacts unusable, even when agents complete more than 50% of other evaluation steps. Our results expose limitations of current agents acting as end-to-end assistants, and call for progress on tool use, visual understanding, and long-horizon reasoning.
Problem

Research questions and friction points this paper is trying to address.

Web Agents
Benchmarking
Knowledge Synthesis
Long-horizon Reasoning
Visual Understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Web Agents Benchmark
Knowledge Synthesis
Hybrid Evaluation
Long-horizon Reasoning
Visual Understanding