🤖 AI Summary
This study investigates whether decentralized architectures can solve complex visual reasoning tasks relying exclusively on strictly local interactions. To this end, we propose a framework based on asynchronously updated Neural Cellular Automata (NCA), integrating experience replay and stochastic perturbation training to directly solve maze navigation, Sudoku, and ARC-AGI tasks through multi-step local recurrence in pixel space. Our work demonstrates that purely locally connected architectures are capable of emergent global reasoning, exhibiting superior out-of-distribution generalization and robust damage recovery. Furthermore, the proposed approach supports dynamic compute allocation and significantly enhances inference efficiency through redundant trajectory pruning.
📝 Abstract
Modern AI architectures used to solve visual reasoning tasks typically rely heavily on global connectivity and synchronization. As biological systems demonstrate, though, sophisticated computation can be performed in a more decentralized fashion. In this work, we test the reasoning capabilities of Neural Cellular Automata (NCAs), networks of recurrent cells that use strictly local connectivity and asynchronous updates. NCAs have been extensively studied in artificial life experiments, but it is unclear whether they can perform complex multi-step reasoning. We show that NCAs produce spatio-temporal dynamics capable of solving challenging visual reasoning tasks, including large mazes, Sudoku, and ARC-AGI-1. Furthermore, we provide evidence that NCAs generalize out-of-distribution when running with larger grids, longer rollouts, or parallel trials; and that the latter can be made more efficient via pruning of redundant trajectories. We find that these generalization capabilities depend on training with sample replay and stochastic perturbations, and that stochasticity remains beneficial at test time. Finally, we show that NCAs are robust reasoners capable of dynamically modulating compute to recover efficiently from damage, and that they can scale to solve reasoning in raw pixel space.