Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过调整Maia-3棋类转换器的技能水平,发现提高技能会促使更深层次注意力层的参与,特别是在特定战术中。
📝 Abstract
Chess involves complex reasoning in a deterministic environment, which makes it a useful setting for studying the mechanisms of computation inside transformers. The Maia-3 chess transformer takes Elo, a measure of competitive chess skill, as an input to the pre-trained network, so we can vary the skill the network is conditioned on with no change to its weights. Here we investigate how turning this skill dial affects self-attention. Ablating every attention head at every Elo from 700 to 2500, we find 1) increasing skill pushes the causal center of mass of the computation deeper, monotonically, for every chess piece and move type we measured; 2) the depth migration is much greater for specific tactics, especially knight forks, than for other move types; 3) the migration consists of deeper heads getting recruited for more specialized computations while one shared shallow head keeps a roughly constant contribution. These results may shed light on how conditioning inputs redistribute computation in larger transformers.
Problem

Research questions and friction points this paper is trying to address.

Chess Transformer
Skill Level
Attention Layers
Elo Rating
Computation Redistribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

skill level
attention layers
chess transformer
Elo rating
computation redistribution
🔎 Similar Papers
No similar papers found.