Comparison of different Unique hard attention transformer models by the formal languages they can recognize

📅 2025-06-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work systematically investigates the formal language recognition capabilities of Unique Hard Attention Transformers (UHAT), establishing precise boundaries for their expressive power. We develop, for the first time, a unified theoretical framework covering key variants—including masked/unmasked attention, finite/infinite image sets, and general/bilinear scoring functions—by characterizing lower bounds via first-order logic and upper bounds via circuit complexity (AC⁰). Leveraging tools from formal language theory, logical modeling, and circuit complexity analysis, we rigorously compare and order the recognition capacities of multiple UHAT classes, providing provable tight bounds. Our results deliver the first rigorous, logically grounded, and computationally precise characterization of hard-attention Transformers’ formal language recognition abilities, thereby filling a critical gap in the theoretical understanding of Transformer limits—specifically addressing the previously unanalyzed hard attention mechanism.

Technology Category

Knowledge Representation and Reasoning: Computational Complexity of ReasoningMachine Learning: Hardware-aware MLNatural Language Processing: (Large) Language Models

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Large language models for searchWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
This note is a survey of various results on the capabilities of unique hard attention transformers encoders (UHATs) to recognize formal languages. We distinguish between masked vs. non-masked, finite vs. infinite image and general vs. bilinear attention score functions. We recall some relations between these models, as well as a lower bound in terms of first-order logic and an upper bound in terms of circuit complexity.
Problem

Research questions and friction points this paper is trying to address.

Compare UHAT models' formal language recognition capabilities
Analyze masked vs. non-masked and finite vs. infinite variants
Establish logic and circuit complexity bounds for UHATs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unique hard attention transformers encoders (UHATs)
Masked vs. non-masked attention models
Bilinear vs. general attention score functions
🔎 Similar Papers