VLM Fine-Tuning for End-to-End Combinatorial Optimization

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of text-serialized representations in capturing spatial and relational structures within combinatorial optimization problems by proposing a general-purpose vision-language solver. The proposed method integrates textual and visual inputs, incorporating gold-standard-free visual representations, and employs a vision-language model (VLM) jointly optimized through supervised fine-tuning and verifier-guided reinforcement learning. Experimental results demonstrate that on complex tasks such as the Capacitated Vehicle Routing Problem (CVRP), visual information significantly enhances solution quality, with this advantage becoming increasingly pronounced as problem scale grows. These findings reveal the core value of visual representations in large-scale combinatorial optimization.
📝 Abstract
Large language models (LLMs) have provided a unified interface for end-to-end combinatorial optimization (CO), but textual serialization alone may obscure spatial and relational structures that are important for generating effective CO solutions. This paper presents a general-purpose vision-language solver that augments textual instance descriptions with input-derived visual representations. A single vision-language model (VLM) is applied across different CO tasks and trained using supervised fine-tuning followed by verifier-guided reinforcement learning. While the visual inputs contain no gold solutions or solution-derived information, our experiments show that the VLM generally improves solution quality over its text-only counterpart, with particularly clear gains on more complex CO problems such as CVRP and JSSP. The advantage of visual information is more pronounced at large problem scales.
Problem

Research questions and friction points this paper is trying to address.

Combinatorial Optimization
Vision-Language Model
Large Language Models
Spatial Structure
End-to-End Solving
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Model
Combinatorial Optimization
Reinforcement Learning
Supervised Fine-Tuning
Visual Representation
🔎 Similar Papers
No similar papers found.