ConstraintBench: Benchmarking LLM Constraint Reasoning on Direct Optimization

πŸ“… 2026-02-25
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work presents the first systematic evaluation of large language models’ ability to directly generate solutions to optimization problems that are both constraint-satisfying and near-optimal, without solver assistance. To this end, the authors introduce ConstraintBench, a benchmark spanning ten operations research domains, which maps natural language problem descriptions to structured solution outputs. Feasibility and suboptimality are rigorously assessed using Gurobi-computed ground truths and a deterministic verification mechanism. Experiments reveal that the best current model achieves a 65.0% constraint satisfaction rate, with feasible solutions reaching 89–96% of Gurobi’s optimal objective values; however, fewer than 30.5% of solutions are simultaneously feasible and near-optimal (within 0.1% optimality gap). The study also releases a complete evaluation infrastructure, addressing a critical gap in systematic assessment for this emerging capability.

Technology Category

Constraint Satisfaction and Optimization: Solvers and ToolsSearch and Optimization: Combinatorial OptimizationReasoning under Uncertainty: Stochastic Optimization

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
πŸ“ Abstract
Large language models are increasingly applied to operational decision-making where the underlying structure is constrained optimization. Existing benchmarks evaluate whether LLMs can formulate optimization problems as solver code, but leave open a complementary question. Can LLMs directly produce correct solutions to fully specified constrained optimization problems without access to a solver? We introduce ConstraintBench, a benchmark for evaluating LLMs on direct constrained optimization across 10 operations research domains, with all ground-truth solutions verified by the Gurobi solver. Each task presents a natural-language scenario with entities, constraints, and an optimization objective; the model must return a structured solution that a deterministic verifier checks against every constraint and the solver-proven optimum. We evaluate six frontier models on 200 tasks and find that feasibility, not optimality, is the primary bottleneck. The best model achieves only 65.0% constraint satisfaction, yet feasible solutions average 89 to 96% of the Gurobi-optimal objective. No model exceeds 30.5% on joint feasibility and optimality within 0.1% of the solver reference. Per-domain analysis shows large variation in difficulty, with average feasibility spanning from 83.3% in the production mix domain to 0.8% in the crew assignment domain. Further, systematic failure modes include duration constraint misunderstanding, entity hallucination, and a feasibility-optimality decoupling in facility location and vehicle routing where models achieve high feasibility but 0% optimality. ConstraintBench and all evaluation infrastructure will be publicly released.
Problem

Research questions and friction points this paper is trying to address.

constrained optimization
large language models
benchmarking
feasibility
optimality
Innovation

Methods, ideas, or system contributions that make the work stand out.

constrained optimization
large language models
benchmarking
direct reasoning
feasibility verification
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
J
Joseph Tso
Haladir Research Team
P
Preston Schmittou
Haladir Research Team
Q
Quan Huynh
Haladir Research Team
J
Jibran Hutchins
Haladir Research Team