🤖 AI Summary
This work investigates interpretable visual reasoning without linguistic supervision and proposes a rectified flow–based diffusion Transformer that models image-to-image reasoning end-to-end purely in pixel space. By integrating noise-free flow matching, Euler integration, and a Transformer architecture, the method efficiently solves tasks such as maze navigation in as few as ten iterative steps. Trained on million-scale synthetic data, it achieves a 192-fold reduction in training loss and a 22.7-fold improvement in L2 error. Notably, the study reports the first observation of an “aha!”-like phase transition analogous to human insight: in 68% of cases, reasoning progresses minimally for most of the trajectory before the correct solution emerges globally and synchronously within the final 2% of steps, challenging conventional assumptions of sequential, incremental reasoning.
📝 Abstract
We present TACIT (Transformation-Aware Capturing of Implicit Thought), a diffusion-based transformer for interpretable visual reasoning. Unlike language-based reasoning systems, TACIT operates entirely in pixel space using rectified flow, enabling direct visualization of the reasoning process at each inference step. We demonstrate the approach on maze-solving, where the model learns to transform images of unsolved mazes into solutions. Key results on 1 million synthetic maze pairs include: - 192x reduction in training loss over 100 epochs - 22.7x improvement in L2 distance to ground truth - Only 10 Euler steps required (vs. 100-1000 for typical diffusion models) Quantitative analysis reveals a striking phase transition phenomenon: the solution remains invisible for 68% of the transformation (zero recall), then emerges abruptly at t=0.70 within just 2% of the process. Most remarkably, 100% of samples exhibit simultaneous emergence across all spatial regions, ruling out sequential path construction and providing evidence for holistic rather than algorithmic reasoning. This"eureka moment"pattern -- long incubation followed by sudden crystallization -- parallels insight phenomena in human cognition. The pixel-space design with noise-free flow matching provides a foundation for understanding how neural networks develop implicit reasoning strategies that operate below and before language.