SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Combinatorial Scheduling

📅 2026-08-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates the faithfulness and consistency of large language models (LLMs) in adhering to constraints when solving natural language scheduling problems that are semantically equivalent but differ in surface form. To this end, we introduce SCHEDBench, the first natural language benchmark for combinatorial scheduling tasks, encompassing diverse problem types such as job shop scheduling (JSP), resource-constrained project scheduling (RCPSP), workforce rostering, and timetabling. We generate varied surface formulations through lexical-syntactic rewrites and constraint reordering. Using scheduling solvers for validation and domain-specific templates for generation, our evaluation across 13 state-of-the-art LLMs reveals a critical lack of invariance: surface-level variations significantly degrade solution feasibility, with constraint reordering most frequently causing violations of hard constraints. This work is the first to expose the pronounced sensitivity of LLMs to linguistic surface form in scheduling reasoning.
📝 Abstract
This paper introduces SCHEDBench, a natural-language benchmark for evaluating combinatorial scheduling constraint faithfulness under surface-form variation. Grounded in canonical scheduling instances and solver-derived feasibility and optimality, SCHEDBench assesses whether large language models (LLMs) generate schedules with the same constraint-feasible behavior across varied natural-language (NL) surface forms. SCHEDBench spans 1,132 instances across job-shop scheduling problems (JSP), single and multi-mode resource-constrained project scheduling problems (RCPSP), nurse rostering/scheduling, and curriculum timetabling problems of varying difficulty. Instances are templated into natural language problems using domain-specific templates, themed entities, lexical-syntactic template rephrasing, and constraint-level surface-form variation, with reference solutions verified for feasibility and objective optimality. Across thirteen frontier and open-weight LLMs, we find that models are not reliably invariant to semantically equivalent renderings of the same scheduling problem. Surface-form variation reduces feasibility and induces above-noise shifts in per-instance hard-constraint violations on matched instances. Among the tested isolated axes, constraint reordering yields the clearest above-noise sensitivity.
Problem

Research questions and friction points this paper is trying to address.

constraint faithfulness
combinatorial scheduling
natural-language variation
large language models
surface-form invariance
Innovation

Methods, ideas, or system contributions that make the work stand out.

constraint faithfulness
natural-language scheduling
surface-form invariance
combinatorial optimization
LLM robustness
💼 Related Jobs
No related jobs found.
S
Shrenil Shaun Sharma
Independent Researcher, San Francisco, CA, USA
A
Avi Sharma
Department of Electrical Engineering and Computer Sciences, University of California, Berkeley