π€ AI Summary
This study addresses the limitations of existing benchmarks in evaluating agentsβ persistent retrieval capabilities across multilingual, multimodal, and complex reasoning chain settings. We construct a stress-test benchmark comprising 423 human-verified challenging questions that require agents to integrate heterogeneous evidence, such as videos and documents, across 13 languages to locate answers. To ensure difficulty stems from open-web evidence mining rather than parametric knowledge, we filter out questions solvable by model parameters alone. Evaluation is conducted through both native search and shared external retrieval frameworks under a unified agent protocol. This work pioneers a multilingual, multimodal evaluation paradigm, effectively quantifying the performance bottlenecks and exploratory effort of current models in complex web-browsing tasks.
π Abstract
We introduce HyperBrowseComp, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers. Questions are designed to be extremely challenging. Each question targets a concise, publicly verifiable answer whose discovery requires locating obscure evidence, following multi-step clue chains, or inspecting heterogeneous sources such as videos, scanned documents, images, or maps. Easier questions are filtered out by evaluating them with models without internet access to reduce the likelihood that they can be answered with parametric knowledge alone. We evaluate several models using provider-native search and a shared external retrieval harness under a common agent protocol. To contextualize model performance and effort, we also conduct a human evaluation on a sample of the questions. HyperBrowseComp provides a challenging testbed for persistent information seeking across languages and evidence modalities, with difficulty arising from discovering and connecting evidence on the open web.