Security-Enhanced Seed-Based Weight Quantization for Large Language Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the substantial storage overhead of large language models and the limitation that existing seed-based compression methods overlook weight sensitivity variations. To this end, we propose Seed-Q, a framework that generates weights via linear feedback shift registers (LFSRs) and introduces a metadata-free, sensitivity-aware non-uniform bit allocation strategy, accompanied by a dedicated ASIC accelerator. Experimental results demonstrate that Seed-Q achieves perplexity comparable to 4-bit quantization at lower bit widths, significantly reducing accuracy degradation. Furthermore, it effectively mitigates bit-flip attacks, with hardware evaluations confirming both low implementation overhead and high security.
📝 Abstract
Large language models (LLMs) incur substantial storage, memory-bandwidth and energy costs, motivating compact weight representations. Existing seed-based compression methods reconstruct weights from compact pseudo-random representations but do not explicitly account for the non-uniform sensitivity of model weights. We introduce Seed-Q, a security-enhanced sensitivity-aware seed-based weight compression framework that uses lightweight Linear Feedback Shift Register (LFSR)-based weight generation with non-uniform bit allocation. Our approach assigns larger representation budgets to sensitive weights while aggressively compressing less sensitive regions. Importantly, this non-uniform allocation requires no side-information: the decoder deterministically reconstructs the bit-allocation schedule, with no rung depending on the decoded weights, eliminating the need to store per-block metadata or use calibration data while preserving the baseline coding rate. Experiments across diverse LLMs show that Seed-Q matches 4-bit perplexity of SeedLM with fewer bits, while at the same 4 bits/weight it reduces both perplexity degradation and zero-shot accuracy loss relative to SeedLM. We also show that Seed-Q simultaneously achieves high security against bit-flip attacks on model parameters, as bit corruption affects multiple reconstructed weights, greatly amplifying its impact and making it easier to detect. We further implement Seed-Q in an ASIC-based accelerator and demonstrate modest hardware overhead compared to prior seed-based approaches.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Weight Quantization
Seed-Based Compression
Security
Bit-Flip Attacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Weight Quantization
Seed-Based Compression
Sensitivity-Aware Bit Allocation
Bit-Flip Attack Security
LFSR
🔎 Similar Papers
No similar papers found.