HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases

πŸ“… 2026-09-24
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing language model evaluation benchmarks in capturing the complexities of enterprise deployment, which leads to distorted performance assessments. To this end, this work proposes HARDEN, a method that leverages constrained evolutionary search along domain-specific complexity axes to automatically generate answer-preserving, high-difficulty evaluation variants while strictly maintaining task semantics, authenticity, and execution validity. Experimental results demonstrate that HARDEN reduces average model accuracy by 22.7%, with maximum declines reaching 49.9%. These findings indicate that the proposed approach significantly enhances evaluation rigor, providing a more realistic and effective tool for assessing the robustness of language models in practical deployment scenarios.
πŸ“ Abstract
Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments. We introduce HARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants while keeping their expected outputs fixed. HARDEN searches along generated domain-specific complexity axes while enforcing feasibility constraints such as preserving task semantics, realism, and execution validity. Across FinQA, PubMedQA, and ContractNLI and three Qwen3.5 model scales (35B-A3B, 122B-A10B, and 397B-A17B), HARDEN reduces task-model accuracy by 22.7% on average and by up to 49.9% relative to single-pass baselines using the same feasibility checks. These results show that evolutionary search can produce substantially harder valid evaluation cases.
Problem

Research questions and friction points this paper is trying to address.

language model evaluation
benchmark complexity
adversarial examples
enterprise deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Constrained Evolutionary Search
Answer-Preserving Evaluation
Complexity Axes
Feasibility Constraints
Language Model Benchmarking
A
Aditya Kumaran
Distyl AI
R
Rahul Singhal
Distyl AI
K
Karime Maamari
Distyl AI
Amine Mhedhbi
Amine Mhedhbi
Polytechnique MontrΓ©al
Data ManagementComputer SystemsAI SystemsQuery ProcessingQuery Optimization
P
Pradyumna Tambwekar
Distyl AI