Seq2Seq Models Reconstruct Visual Jigsaw Puzzles without Seeing Them

📅 2025-11-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the feasibility of solving square jigsaw puzzle reconstruction using language models alone—without any visual input. The puzzle reconstruction task is formulated as a sequence-to-sequence problem; a dedicated symbolic tokenizer discretizes image patches into language-like tokens, enabling “vision-free” puzzle solving. An encoder-decoder Transformer architecture is employed to model spatial relationships among patches end-to-end. To our knowledge, this is the first approach that entirely eliminates visual feature extraction and solves a canonical visual reasoning task purely through large language models, establishing a novel cross-modal reasoning paradigm. Evaluated on multiple standard jigsaw benchmarks, our method achieves state-of-the-art performance, significantly outperforming conventional vision-based approaches—including CNN- and ViT-based methods—demonstrating the strong generalization capability and untapped potential of LLMs in non-linguistic spatial reasoning tasks.

Technology Category

Computer Vision: Large Vision ModelsMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language Models

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Large language models for searchGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Jigsaw puzzles are primarily visual objects, whose algorithmic solutions have traditionally been framed from a visual perspective. In this work, however, we explore a fundamentally different approach: solving square jigsaw puzzles using language models, without access to raw visual input. By introducing a specialized tokenizer that converts each puzzle piece into a discrete sequence of tokens, we reframe puzzle reassembly as a sequence-to-sequence prediction task. Treated as"blind"solvers, encoder-decoder transformers accurately reconstruct the original layout by reasoning over token sequences alone. Despite being deliberately restricted from accessing visual input, our models achieve state-of-the-art results across multiple benchmarks, often outperforming vision-based methods. These findings highlight the surprising capability of language models to solve problems beyond their native domain, and suggest that unconventional approaches can inspire promising directions for puzzle-solving research.
Problem

Research questions and friction points this paper is trying to address.

Solving jigsaw puzzles without visual input using language models
Reframing puzzle reassembly as sequence-to-sequence prediction task
Achieving state-of-the-art results with blind transformer solvers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Using language models to solve visual puzzles
Tokenizer converts puzzle pieces into token sequences
Encoder-decoder transformers reconstruct layout without visual input
🔎 Similar Papers
2024-03-04Computer Vision and Pattern RecognitionCitations: 3