🤖 AI Summary
This study addresses the limitations of text-serialized representations in capturing spatial and relational structures within combinatorial optimization problems by proposing a general-purpose vision-language solver. The proposed method integrates textual and visual inputs, incorporating gold-standard-free visual representations, and employs a vision-language model (VLM) jointly optimized through supervised fine-tuning and verifier-guided reinforcement learning. Experimental results demonstrate that on complex tasks such as the Capacitated Vehicle Routing Problem (CVRP), visual information significantly enhances solution quality, with this advantage becoming increasingly pronounced as problem scale grows. These findings reveal the core value of visual representations in large-scale combinatorial optimization.
📝 Abstract
Large language models (LLMs) have provided a unified interface for end-to-end combinatorial optimization (CO), but textual serialization alone may obscure spatial and relational structures that are important for generating effective CO solutions. This paper presents a general-purpose vision-language solver that augments textual instance descriptions with input-derived visual representations. A single vision-language model (VLM) is applied across different CO tasks and trained using supervised fine-tuning followed by verifier-guided reinforcement learning. While the visual inputs contain no gold solutions or solution-derived information, our experiments show that the VLM generally improves solution quality over its text-only counterpart, with particularly clear gains on more complex CO problems such as CVRP and JSSP. The advantage of visual information is more pronounced at large problem scales.