Global Ranks Survive, Selected Heads Shift: BOS-Sink Topology under 4-bit Weight-Only Quantization

๐Ÿ“… 2026-09-20
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
็ ”็ฉถๆŽข่ฎจไบ†4ไฝๆƒ้‡้‡ๅŒ–ไธ‹๏ผŒ้‡่ฆๆณจๆ„ๅŠ›ๅคด็š„่ฏ†ๅˆซไธŽ้‡็”จ้—ฎ้ข˜๏ผŒๅนถ้€š่ฟ‡Sinkๆ‹“ๆ‰‘ไธ€่‡ดๆ€งๆŒ‡ๆ ‡่ฏ„ไผฐๅ…ถๅฎ‰ๅ…จๆ€งๅ’Œๆœ‰ๆ•ˆๆ€งใ€‚
๐Ÿ“ Abstract
Sink-aware deployment may identify important first-token attention heads before a model is quantized, then reuse that map at the edge. We test when this shortcut is safe for 4-bit NF4 weight-only post-training quantization (PTQ). Our Sink Topology Consistency (STC) metrics separate global rank preservation, top-$k$ set overlap, and layerwise sink-mass shift, and distinguish per-input sensitivity from calibration-map transfer. Across Qwen2.5-0.5B, Qwen2.5-1.5B, and Llama-3.2-1B, global bf16-to-4-bit ranks remain high at 4,096 tokens ($ฯ_s \geq 0.980$), yet top-$k$ Jaccard overlap is only 0.619-0.793, corresponding to 76.5-88.5% membership retention. The global statistic also masks local failures: terminal Qwen layers shift by 6.2-7.9x their model means, whereas Llama-3.2-1B shows low, nearly uniform drift. Under a C4-to-LongBench shift, cross-domain overlap degrades more than the within-domain precision comparison for both Qwen models, but not for Llama-3.2-1B. Matched-domain 4-bit recalibration reaches 90% of a split-half stability plateau at the smallest tested $n=8$ for both Qwen models and $n=32$ for Llama-3.2-1B, though not as a sharp threshold; for the two Qwen models, updating only selected layers does not reach the full-map stability criterion. On Jetson Orin NX, the 16-sample workload takes seconds for the two models with valid on-device sink measurements. The practical message is precise: global rankings often transfer, but discrete head sets, layer-local policies, and cross-domain calibration should be revalidated after quantization.
Problem

Research questions and friction points this paper is trying to address.

Sink Topology
Quantization
Attention Heads
Global Ranks
Layerwise Shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sink Topology Consistency
4-bit weight-only quantization
Attention Heads Stability
Cross-Domain Calibration
๐Ÿ”Ž Similar Papers
No similar papers found.
K
Kuanlin Chen
Independent Researcher
C
Chen-Wei Kuo
National Tsing Hua University
C
Cheng-En Ou
Independent Researcher