HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of existing benchmarks in evaluating agents’ persistent retrieval capabilities across multilingual, multimodal, and complex reasoning chain settings. We construct a stress-test benchmark comprising 423 human-verified challenging questions that require agents to integrate heterogeneous evidence, such as videos and documents, across 13 languages to locate answers. To ensure difficulty stems from open-web evidence mining rather than parametric knowledge, we filter out questions solvable by model parameters alone. Evaluation is conducted through both native search and shared external retrieval frameworks under a unified agent protocol. This work pioneers a multilingual, multimodal evaluation paradigm, effectively quantifying the performance bottlenecks and exploratory effort of current models in complex web-browsing tasks.
πŸ“ Abstract
We introduce HyperBrowseComp, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers. Questions are designed to be extremely challenging. Each question targets a concise, publicly verifiable answer whose discovery requires locating obscure evidence, following multi-step clue chains, or inspecting heterogeneous sources such as videos, scanned documents, images, or maps. Easier questions are filtered out by evaluating them with models without internet access to reduce the likelihood that they can be answered with parametric knowledge alone. We evaluate several models using provider-native search and a shared external retrieval harness under a common agent protocol. To contextualize model performance and effort, we also conduct a human evaluation on a sample of the questions. HyperBrowseComp provides a challenging testbed for persistent information seeking across languages and evidence modalities, with difficulty arising from discovering and connecting evidence on the open web.
Problem

Research questions and friction points this paper is trying to address.

Web-Browsing Agents
Multilingual Benchmark
Multimodal Reasoning
Information Seeking
Stress Test
Innovation

Methods, ideas, or system contributions that make the work stand out.

Web-Browsing Agents
Multilingual Benchmark
Multimodal Reasoning
Information Seeking
Stress Test
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
Alham Fikri Aji
Alham Fikri Aji
MBZUAI, Monash Indonesia
MultilingualityLow-resource NLPLanguage ModelingMachine Translation
F
Faiz Rizki Ramadhan
Mohamed bin Zayed University of Artificial Intelligence
Z
Zayd M. K. Zuhri
Mila – Quebec Artificial Intelligence Institute
S
Seung Hun Eddie Han
Inception AI
Ryandito Diandaru
Ryandito Diandaru
Master's Student, MBZUAI
NLP
Q
Qinrong Cui
Mohamed bin Zayed University of Artificial Intelligence
Jan Christian Blaise Cruz
Jan Christian Blaise Cruz
MBZUAI, McGill University, Mila - Quebec AI Institute
Natural Language ProcessingTranslationMultilingualityLow-resource LanguagesCode Switching
Badrinath Chandana
Badrinath Chandana
Mohamed bin Zayed University of Artificial Intelligence
P
Peerawat Chomphooyod
Mohamed bin Zayed University of Artificial Intelligence
A
Ahmed Attia
Mohamed bin Zayed University of Artificial Intelligence
Jonibek Mansurov
Jonibek Mansurov
PhD student in NLP, MBZUAI
NLP
E
Emilio Villa-Cueva
Mohamed bin Zayed University of Artificial Intelligence
C
Canh Duong Nguyen
Mohamed bin Zayed University of Artificial Intelligence
I
Imran Turganov
Mohamed bin Zayed University of Artificial Intelligence
M
Minghao Wu
Alibaba Group
Peerat Limkonchotiwat
Peerat Limkonchotiwat
Research Fellow, AI Singapore, National University of Singapore
Evaluation and BenchmarkRepresentation LearningLarge Language ModelMultilingual Learning
Irina Nikishina
Irina Nikishina
Postdoc @ University of Hamburg
Natural Language ProcessingRAGTaxonomiesQuestion Answering