Game-Theoretic Co-Evolution for LLM-Based Heuristic Discovery

📅 2026-01-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited generalization of existing large language model (LLM)-based approaches to automatic heuristic discovery, which rely on static evaluation and are prone to overfitting under distributional shift. To overcome this, we propose the Algorithm Space Response Oracles (ASRO) framework, which uniquely integrates game theory with LLM-driven heuristic search by modeling heuristic discovery as a program-level co-evolution between a solver and an instance generator. Within a zero-sum game formulation, ASRO employs the LLM as an optimal-response oracle and dynamically constructs adversarial training curricula through mixed-strategy iterations over evolving strategy pools for both players. Evaluated across multiple combinatorial optimization tasks, ASRO substantially outperforms static-training baselines, demonstrating superior generalization and robustness both in-distribution and out-of-distribution.

Technology Category

Search and Optimization: Heuristic SearchGame Theory and Economic Paradigms: Adversarial LearningMultiagent Systems: Adversarial Agents

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationEconomics, Online Markets and Human Computation: Uses of LLMs and GenAI for marketplace design, bidding, and strategic interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Large language models (LLMs) have enabled rapid progress in automatic heuristic discovery (AHD), yet most existing methods are predominantly limited by static evaluation against fixed instance distributions, leading to potential overfitting and poor generalization under distributional shifts. We propose Algorithm Space Response Oracles (ASRO), a game-theoretic framework that reframes heuristic discovery as a program level co-evolution between solver and instance generator. ASRO models their interaction as a two-player zero-sum game, maintains growing strategy pools on both sides, and iteratively expands them via LLM-based best-response oracles against mixed opponent meta-strategies, thereby replacing static evaluation with an adaptive, self-generated curriculum. Across multiple combinatorial optimization domains, ASRO consistently outperforms static-training AHD baselines built on the same program search mechanisms, achieving substantially improved generalization and robustness on diverse and out-of-distribution instances.
Problem

Research questions and friction points this paper is trying to address.

automatic heuristic discovery
distributional shift
generalization
overfitting
combinatorial optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

game-theoretic co-evolution
automatic heuristic discovery
large language models
zero-sum game
adaptive curriculum