🤖 AI Summary
This work investigates the feasibility of solving square jigsaw puzzle reconstruction using language models alone—without any visual input. The puzzle reconstruction task is formulated as a sequence-to-sequence problem; a dedicated symbolic tokenizer discretizes image patches into language-like tokens, enabling “vision-free” puzzle solving. An encoder-decoder Transformer architecture is employed to model spatial relationships among patches end-to-end. To our knowledge, this is the first approach that entirely eliminates visual feature extraction and solves a canonical visual reasoning task purely through large language models, establishing a novel cross-modal reasoning paradigm. Evaluated on multiple standard jigsaw benchmarks, our method achieves state-of-the-art performance, significantly outperforming conventional vision-based approaches—including CNN- and ViT-based methods—demonstrating the strong generalization capability and untapped potential of LLMs in non-linguistic spatial reasoning tasks.
📝 Abstract
Jigsaw puzzles are primarily visual objects, whose algorithmic solutions have traditionally been framed from a visual perspective. In this work, however, we explore a fundamentally different approach: solving square jigsaw puzzles using language models, without access to raw visual input. By introducing a specialized tokenizer that converts each puzzle piece into a discrete sequence of tokens, we reframe puzzle reassembly as a sequence-to-sequence prediction task. Treated as"blind"solvers, encoder-decoder transformers accurately reconstruct the original layout by reasoning over token sequences alone. Despite being deliberately restricted from accessing visual input, our models achieve state-of-the-art results across multiple benchmarks, often outperforming vision-based methods. These findings highlight the surprising capability of language models to solve problems beyond their native domain, and suggest that unconventional approaches can inspire promising directions for puzzle-solving research.