🤖 AI Summary
This study addresses the challenge of fixed-bitwidth quantization in Vision Transformer inference, which struggles to balance accuracy and efficiency at a fine granularity. To overcome this limitation, this work proposes treating stochastic computing (SC) as a dense adaptive quantizer, where bitstream length dynamically controls precision. Methodologically, we construct a GPU-accelerated library based on AND/XNOR logic and design a row-wise dynamic mixed-precision strategy, enabling flexible training-free inference. Experimental results demonstrate that this paradigm achieves performance comparable to INT quantization across diverse vision tasks while maintaining high accuracy under low average bit budgets. These findings validate the feasibility of SC as an efficient inference foundation for modern vision models.
📝 Abstract
Matrix multiplications dominate the inference cost of modern transformer-based vision models, yet existing efficiency techniques such as post-training quantization and mixed-precision inference are largely limited to the small set of fixed-width formats (INT4, INT8, BF16, and FP16) supported by conventional accelerators. We revisit stochastic computing (SC) as a way to lift this constraint: viewed as a dense adaptive quantizer, SC controls precision by bit-stream length L rather than a fixed datapath, while each multiplication reduces to a single AND/XNOR gate. We build a GPU library that emulates SC matrix multiplication at scale, exposes stream lengths as first-class kernel arguments, and evaluates SC end-to-end on image classification, object detection and instance segmentation, class-conditional image generation, and visual world-model planning. On top of this substrate, we develop a dynamic per-row mixed-precision policy that assigns stream length per token or group at matched average budget, requires no retraining, and uses the same SC hardware across schedules. Across tasks, SC remains competitive with fixed-format INT quantization at matched bit budgets, while per-row mixed precision helps maintain accuracy at lower average stream lengths. These results provide software-level feasibility evidence that SC can serve as a dense-precision substrate for fine-grained mixed-precision inference on modern vision transformers.