Encoded but Not Decoded: Layer-Localized Evidence for a Three-Level Gap in LLM Syntax

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of existing behavioral evaluations to distinguish whether syntactic failures in large language models stem from unencoded internal representations or undecoded outputs. To resolve this, we construct a three-tier evaluation framework—probe recovery, LM head reading, and behavioral deployment—combining linear probing with activation patching to systematically assess seven models on multilingual control-dependency benchmarks. We demonstrate that probe recovery rates consistently match or exceed downstream performance, introducing the concept of recoverability surplus. Furthermore, our analysis reveals that instruction tuning exacerbates decoding rather than encoding degradation. This disparity concentrates in subject-control tasks, is driven by surface-level shortcut biases, and exhibits layer-localized characteristics.
📝 Abstract
A language model can fail a syntactic test in two distinct ways: by not encoding the relevant structure, or by encoding it but failing to use it at the output. Behavioral evaluation alone cannot tell these apart. We propose a three-level evaluation framework (behavioral deployment, LM-head readout, and probe recoverability) measured on the same items under the same binary decision. Using a compact trilingual (English, Chinese, German) control-dependency benchmark, we find that probe recoverability exceeds or equals LM-head readout, which in turn exceeds or equals behavioral deployment, across seven models and all three languages in the aggregate. The recoverability surplus is never negative across all 14 (model, task) conditions. The disconnect concentrates in subject-control, where a nearest-noun heuristic gives the wrong answer. The single largest gap (0.653) appears on Qwen3-0.6B Instruct in question answering. The gap persists at Qwen3-14B Instruct. Instruction tuning degrades deployment more than encoding in percentage terms. We rule out option-position bias, late-layer erasure, output-formatting artifacts, and probe-training variance. The pattern is consistent with decoding that favors surface shortcuts, and the behavior-probe gap measures the strength of that preference. Activation patching shows the gap is layer-localized. Under instruction tuning, the LM-head-decoded layer shifts approximately ten layers later than the probe-decoded layer. These findings argue that behavioral evaluation understates what models encode, while probing alone overstates what they deploy.
Problem

Research questions and friction points this paper is trying to address.

syntactic evaluation
encoding-decoding gap
large language models
behavioral assessment
probe recoverability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Three-level evaluation framework
Probe recoverability
Layer-localized gap
Activation patching
Instruction tuning
🔎 Similar Papers
No similar papers found.
Z
Zhenyan Lu
College of International Studies, National University of Defense Technology, Nanjing, China
He Wang
He Wang
Nanjing University of Science and Technology
Image Fusion and Reconstruction
X
Xiaohui Huang
College of International Studies, National University of Defense Technology, Nanjing, China