🤖 AI Summary
This study addresses the excessive vector register file (VRF) area, energy consumption, and register pressure in long-vector architectures caused by scalar-semantic values occupying full-width registers. To mitigate this, we propose SVRF, a microarchitectural optimization that introduces a novel scalarized VRF mechanism to redirect scalar-semantic vectors into a dedicated scalar register file, thereby avoiding full-width physical register allocation. This approach achieves efficient separation and reuse of storage resources while preserving ISA compatibility and vector renaming. Implemented at the RTL level using the RISC-V Vector Extension (RVV), experimental evaluations demonstrate that SVRF reduces VRF area by 15.8% and total power consumption by 12%, while accelerating certain workloads by 1.12×. These improvements significantly enhance execution efficiency for HPC and AI applications.
📝 Abstract
Vector processors exploit data-level parallelism to provide high computational throughput, but very long vectors can make the Vector Register File (VRF) a significant source of area, energy, and register pressure. This cost is exacerbated when vector instructions operate on values with scalar semantics: although such values are architecturally represented using vector registers, they do not require full-width storage. This paper proposes the Scalarized Vector Register File (SVRF), a microarchitectural enhancement that redirects scalar-semantics vector values to a dedicated scalar register file, avoiding unnecessary allocation and access of full-width physical vector registers. The SVRF integrates with vector register renaming while preserving compatibility with the supported vector ISA. We implement and evaluate SVRF in an Register Transfer Level (RTL)-based RISC-V Vector Extension (RVV)-compliant vector processing unit supporting very long vectors. Across representative HPC and AI/ML workloads, the SVRF reduces vector-register pressure and VRF activity without degrading performance. By exploiting the resulting reduction in VRF demand, the proposed design achieves up to 15.8% VRF area reduction and up to 6% total area reduction, while reducing total power by up to 12% in evaluated configurations. Some workloads also achieve up to 1.12X speedup.