CruxBench: A Benchmark of Information Discovery

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of evaluating "information discovery" in large language models (LLMs), specifically their capacity to identify pivotal sub-questions for decomposing complex tasks. To this end, it proposes an evaluation benchmark grounded in the Value of Information (VOI), which quantifies the extent to which posing questions updates predictive beliefs. This framework introduces a novel contamination-resistant, open-ended design that leverages future events to generate ground truth and compute VOI, thereby enabling scalable model assessment. The findings demonstrate that VOI correlates strongly with overall model capability. Furthermore, the analysis reveals that frontier LLMs perform only marginally better than random baselines on information discovery, underscoring the substantial difficulty of this task.
📝 Abstract
Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels. But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions -- which we call cruxes -- whose answers provide key steps on the path toward solving the target problem. To evaluate this capability of information discovery, we introduce CruxBench, a benchmark that grades LLM-generated questions by their Value of Information (VOI): how much a model-proposed crux updates beliefs about a target forecasting question. CruxBench enjoys a rare combination of three key properties: it is (1) contamination-resistant by construction, since ground truth is generated by future world events; (2) open-ended, admitting unbounded and complex text-based submissions rather than one correct numeric answer; and (3) grounded, with informativeness measured against quantified changes in real-world beliefs. We evaluate a diverse set of eight models on 293 target forecasting questions and find that VOI correlates highly with independent measures of model capability (r=0.90) and captures cruxes' usefulness for answering target questions. However, information discovery remains challenging even for frontier LLMs, which only narrowly outperform a random-timing baseline.
Problem

Research questions and friction points this paper is trying to address.

Information Discovery
Large Language Models
Benchmark Evaluation
Value of Information
Problem Decomposition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Value of Information
Information Discovery
Benchmark
Contamination-resistant
Open-ended Evaluation
🔎 Similar Papers
No similar papers found.