🤖 AI Summary
Standard Transformers lack the inductive bias necessary for iteratively traversing implicit relational structures in reasoning tasks. To address this limitation, this work proposes the Graph Machine architecture, which introduces an explicit edge mechanism as a novel inductive bias into neural networks. By integrating edge-augmented attention and an edge-centric referencing mechanism, the model enables dynamic, differentiable construction and updating of relational graphs. Evaluated on the Sudoku benchmark, the proposed method significantly outperforms standard Transformers. Ablation studies confirm that the edge mechanism is crucial for performance gains and further reveal that the model automatically learns compact geometric relational representations.
📝 Abstract
Transformers provide a powerful architecture for global content-based matching, but reasoning problems may benefit from a stronger inductive bias toward iterative traversal of latent relations. We introduce Graph Machine, an architecture with two explicit edge-based mechanisms: Edge-augmented attention, in which edges modulate attention between nodes, and edge-centric referral, in which nodes exchange addresses to update their edges. Conceptually, this enables the model to dynamically and differentiably construct and revise relational graphs across layers. We study this inductive bias using Sudoku under controlled settings and find that Graph Machine outperforms Transformer baselines, with ablation studies and mechanistic analysis attributing the gains to the edge mechanisms. Surprisingly, we found that the model discovers a compact edge-based construction for Sudoku geometry. Our results support explicit edge mechanisms as a promising architectural design, motivating broader evaluation.