🤖 AI Summary
Existing automata extraction methods for recurrent neural networks (RNNs) fail to model RNN behaviors over extremely large or infinite alphabets, hindering interpretability. Method: We propose the first passive lattice automaton extraction framework for RNNs operating over **infinite alphabets**, integrating abstract interpretation (to over-approximate the state space), passive grammar inference, symbolic execution, and an equivalence query mechanism to extract semantics-preserving lattice automata from trained RNNs. Contributions/Results: (1) First automata learning approach supporting regular languages over infinite alphabets; (2) A novel infinite-alphabet extension of Tomita grammars as a principled benchmark; (3) State-of-the-art accuracy on standard Tomita tasks and successful validation on infinite-alphabet cases. This work bridges neural-symbolic learning, formal verification, and interpretable AI.
📝 Abstract
We present a passive automata learning algorithm that can extract automata from recurrent networks with very large or even infinite alphabets. Our method combines overapproximations from the field of Abstract Interpretation and passive automata learning from the field of Grammatical Inference. We evaluate our algorithm by first comparing it with the state-of-the-art automata extraction algorithm from Recurrent Neural Networks trained on Tomita grammars. Then, we extend these experiments to regular languages with infinite alphabets, which we propose as a novel benchmark.