Learning Decision-Stump Thresholds in Context: Dynamics of Softmax Attention

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how gradient-based pretraining learns decision thresholds in two-parameter Softmax attention models. Methodologically, by leveraging large-resolution initialization and statistical learning theory, it reveals the divergence mechanism of synergistic parameters and proposes a novel theoretical framework that extends the effective training regime through gradient error control. In terms of contributions, this work establishes error bounds for frozen estimators, characterizes the temporal decay dynamics of population threshold errors, and delineates the inherent limitations of single-head architectures. Collectively, these findings provide rigorous theoretical support for understanding the threshold learning dynamics within attention mechanisms.
📝 Abstract
Estimating a decision threshold requires locating observations near an unknown boundary. We study how gradient-based pretraining learns this statistical rule in a two-parameter softmax-attention model with a fixed feature and inequality direction. Pretraining uses labeled contexts and their true thresholds; a fresh threshold must be inferred from context alone. Under a large-resolution initialization, constant-step gradient descent on $m$ tasks with $n$ examples each produces a frozen estimator with error $\widetilde O((m\wedge n)^{-1}+N^{-1})$ for each fixed interior threshold and every fresh-context size $N$. The two terms separate finite-pretraining accuracy from fresh-context localization. The mechanism is coordinated parameter divergence: population training calibrates the relative label and feature scores, then increases the attention scale as $t^{1/4}$, giving population threshold error $O(t^{-1/4})$. To transfer this mechanism to a fixed finite corpus, we control gradient errors relative to the shrinking directions of progress at successive parameter scales. This certifies a growing training interval without requiring long-time tracking of the population trajectory. We also identify the boundary limitation of the one-head model and explain statistically what a reflected symmetrization could achieve.
Problem

Research questions and friction points this paper is trying to address.

decision-stump thresholds
softmax attention
pretraining dynamics
context-based inference
parameter divergence
🔎 Similar Papers
No similar papers found.