🤖 AI Summary
Classical continuous neural networks struggle to learn discrete formal rules—such as modular arithmetic, non-Abelian groups, and systematic linguistic composition—with precision, often relying on large-scale parameters and exhibiting stochastic generalization. This work proposes a native quantum architecture that leverages the intrinsic properties of multi-qubit systems as inductive bias, integrating parameterized geometric phase embedding, SU(2) wave interference, and a novel quantum attention mechanism. With only 5–6 qubits and 551–1,650 trainable parameters, the model achieves mathematically exact generalization—termed “crystallization”—on tasks including Z₁₁ modular arithmetic, the S₄ permutation group, and the SCAN benchmark, markedly surpassing conventional “aha”-moment generalization. Experimental validation on IBM quantum hardware demonstrates an accuracy of 97.5%.
📝 Abstract
Classical continuous-space neural networks fundamentally struggle to lock into exact mathematical symmetries, such as modular arithmetic and non-commutative algebra. To approximate these discrete logical rules, they often rely on massive parameter scaling, resulting in stochastic instability even after delayed generalization phenomena known as grokking. Here, we introduce the Universal Quantum Transformer (UQT), a fundamentally novel, quantum-native computing architecture that uses the physical properties of multi-qubit systems as a universal inductive bias for exact mathematical and algebraic reasoning. Rather than translating classical neural mechanisms, our framework relies entirely on parameterized geometric phase embedding and $SU(2)$ wave-interference. We demonstrate that the quantum attention circuit, operating on a highly compact 5-qubit substrate, perfectly learns two highly distinct formal classes: cyclic modular arithmetic ($\mathbb{Z}_{11}$) and non-Abelian algebra (the $S_4$ permutation group). While classical attention-based networks exhibit stochastic instability at convergence, the UQT achieves mathematically exact, deterministic generalization. We refer to this phenomenon as crystallization: a step beyond the well-known phenomenon of grokking. Crucially, this framework yields massive computational and memory advantages by theoretically bypassing the quadratic bottleneck of classical self-attention, and by logarithmically compressing the required representation dimension to eliminate the massive over-parameterization inherent to classical networks. Finally, we deploy this architecture on noisy intermediate-scale quantum (NISQ) hardware, proving its viability on current IBM Quantum computers. These results establish parameterized quantum topology as a universally superior physical substrate for exact artificial intelligence.