🤖 AI Summary
This study addresses the challenge that mechanistic interpretability in large language models fails to scale with architectural size. To overcome this limitation, it proposes the Stream Recurrent Model (SRM), which exposes internal computational structures through multi-stream recurrent refinement and interactive latent streams, thereby achieving hierarchical reasoning that balances performance with transparency. The proposed framework enables fine-grained analysis of dynamic causal contributions and routing behaviors. Notably, SRM attains single-parameter performance comparable to GPT-2 while revealing structured division-of-labor patterns across streams. By elucidating these inter-stream dynamics, this work establishes a foundational approach for scalable interpretability in large-scale architectures.
📝 Abstract
Mechanistic interpretability seeks to make verifiable statements about the internal behavior of large language models (LLMs). Many interpretability techniques struggle to scale with the increasing size and depth of architectures. Our solution to this is to introduce smaller models with structures that lend themselves to interpretability. In this work, we introduce the Stream Recursion Model (SRM), a modification of the Hierarchical Reasoning Model (HRM) designed to expose internal computational structure while remaining scalable. SRM organizes computation into multiple interacting latent streams that are updated through recursive refinement, enabling direct analysis of stream dynamics, causal contribution, and routing behavior. SRM achieves performance comparable to GPT-2 on a per-parameter basis. Our analysis reveals consistent and distinct behavior across streams, indicating structured specialization and interaction. These results suggest that SRM provides a practical architectural foundation for scalable mechanistic interpretability and opens up promising avenues for future research in both reasoning performance and interpretability.