🤖 AI Summary
This study investigates the performance and energy efficiency differences between ARM and x86-64 laptop processors, demonstrating that these disparities stem not only from instruction set architecture (ISA) but also significantly from system-level design choices. For the first time, the authors conduct a comprehensive evaluation on real-world laptop platforms—Apple M3 and AMD Ryzen 7 3750H—combining fine-grained power measurements with microarchitectural analysis. Using assembly-level benchmarks (recursive Fibonacci, integer matrix multiplication), cross-platform performance counters, and portable C-based probes, they systematically assess the impact of architectural and integration factors on energy efficiency. Results reveal that while the Ryzen platform excels in branch-intensive workloads, the Apple platform achieves substantially superior energy efficiency, reducing energy-per-computation by 5.82× and 6.38× respectively, thereby highlighting the critical role of non-ISA design elements.
📝 Abstract
ARM-based and x86-64 laptop processors differ not only in instruction-set design, but also in memory hierarchy, core organization, system integration, and power-management mechanisms. This study presents a combined architectural and experimental comparison of an Apple M3 system and an AMD Ryzen 7 3750H system. The architectural analysis contrasts AArch64's fixed-width load-store design with the variable-length, memory-operand-rich x86-64 instruction model, and discusses how register organization, calling conventions, heterogeneous core organization, memory behavior, and low-power mechanisms shape observed performance and energy characteristics. The experimental part uses two native assembly benchmarks: a recursive Fibonacci workload and an integer matrix-multiplication workload. The analysis combines repeated timing measurements, processor-energy measurements, and cross-platform microarchitectural counter measurements from matched portable-C profiling runs. The Ryzen platform is decisively faster on the branch-heavy Fibonacci benchmark, while matrix multiplication shows no meaningful timing advantage for either platform in the present measurements. In contrast, the Apple platform is markedly more energy-efficient, reducing energy-to-solution by approximately 5.82$\times$ on Fibonacci and 6.38$\times$ on matrix multiplication. These results are interpreted as platform-level findings rather than as pure ISA-only effects, reflecting differences in implementation, system integration, and measurement methodology in addition to instruction-set structure.