VeRA: Verified Reasoning Data Augmentation at Scale

📅 2026-01-23
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究通过创建可执行的任务族来更新推理基准,解决其新鲜度和难度问题,提高任务质量与挑战性。
📝 Abstract
The main issue with most evaluation schemes today is their"static"nature: the same problems are reused repeatedly, allowing for memorization, format exploitation, and eventual saturation. To measure genuine AI progress, we need evaluation that is robust by construction, not by post-hoc detection. In response, we propose VeRA (Verified Reasoning Data Augmentation), a framework that converts benchmark problems into executable specifications, comprising (i) a natural language template with placeholder slots, (ii) a coherent generator that samples valid configurations, and (iii) a deterministic verifier that validates parameters and calculates the corresponding correct answers for each configuration. From a single seed problem, VeRA automatically creates unlimited verified variants with reliable labels at near-zero marginal cost without human involvement. VeRA operates in two complementary modes. VeRA-E (equivalent) rewrites problems while keeping the underlying logic intact, useful for detecting memorization versus genuine reasoning. VeRA-H (hardened) systematically increases complexity while remaining verifiable, enabling reliable creation and labelling of fresh difficult tasks at the boundary of intelligence. Evaluating 16 frontier models with VeRA, we find: (i) VeRA-E improves evaluation quality and reveals contamination patterns. (ii) VeRA-H enables human-free generation of hard tasks with reliable labels. (iii) VeRA establishes verified benchmarks as a general paradigm. VeRA reconceptualizes benchmarks from static objects used until exhausted, to executable specifications generating fresh, verified instances on demand, enhancing robustness and cost-effectiveness for evaluation. With VeRA, we envision that evaluation in any verifiable domain can scale indefinitely without sacrificing label integrity. To stimulate future research, we have open-sourced all code and datasets.
Problem

Research questions and friction points this paper is trying to address.

reasoning benchmarks
freshness
headroom
Innovation

Methods, ideas, or system contributions that make the work stand out.

executable specifications
benchmark renewal
reasoning benchmarks
task families
human auditing
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Z
Zerui Cheng
ByteDance Seed, Princeton University
Jiashuo Liu
Jiashuo Liu
Tsinghua University
Robust OptimizationOOD GeneralizationData-Centric AI
C
Chunjie Wu
ByteDance Seed
J
Jiayang Sun
J
Jianzhu Yao
Princeton University
Pramod Viswanath
Pramod Viswanath
Forrest G. Hamrick Professor in Engineering
BlockchainsWireless communication
G
Ge Zhang
ByteDance Seed
W
Wenhao Huang
ByteDance Seed