🤖 AI Summary
This work systematically investigates the formal language recognition capabilities of Unique Hard Attention Transformers (UHAT), establishing precise boundaries for their expressive power. We develop, for the first time, a unified theoretical framework covering key variants—including masked/unmasked attention, finite/infinite image sets, and general/bilinear scoring functions—by characterizing lower bounds via first-order logic and upper bounds via circuit complexity (AC⁰). Leveraging tools from formal language theory, logical modeling, and circuit complexity analysis, we rigorously compare and order the recognition capacities of multiple UHAT classes, providing provable tight bounds. Our results deliver the first rigorous, logically grounded, and computationally precise characterization of hard-attention Transformers’ formal language recognition abilities, thereby filling a critical gap in the theoretical understanding of Transformer limits—specifically addressing the previously unanalyzed hard attention mechanism.
📝 Abstract
This note is a survey of various results on the capabilities of unique hard attention transformers encoders (UHATs) to recognize formal languages. We distinguish between masked vs. non-masked, finite vs. infinite image and general vs. bilinear attention score functions. We recall some relations between these models, as well as a lower bound in terms of first-order logic and an upper bound in terms of circuit complexity.