Machine Shape and Hierarchical Blocking: A Mathematics of Arrays Formalization, with an Open Problem in Hierarchical Shape Occupancy

📅 2026-08-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a method to automatically derive efficient tiling and prefetching schedules from hardware cache hierarchies, eliminating reliance on empirical tuning. Building upon the Mathematics of Arrays framework, it introduces a machine shape—characterized by cache capacities, bandwidths, and occupancy sequences—to model multilevel caches, and designs novel operators that map this machine shape to tiling strategies, reducing unknown parameters to a small set of level-specific occupancies. The study addresses the fundamental open question of which parameters can be directly inferred from hardware specifications. Experiments on three real machines successfully reproduce expert-tuned tile sizes, and demonstrate that certain throughput-related parameters can indeed be derived from datasheets; however, cross-architecture portability remains limited and requires further improvement.
📝 Abstract
A companion empirical study found that dense matrix multiplication block sizes calibrated on Apple M1 Pro correspond to two cache-t formulas that mispredict badly on a dierent chip's known cache sizes. This paper formalizes the question that nding raises. We extend the Mathematics of Arrays (MoA) framework's array-shape derivation operator to a new operator that derives a hierarchical, multi-level blocking and prefetch schedule from a machine's shape: an ordered sequence of cache-level capacities, bandwidths, and occupancy fractions. This operator recovers the calibrated values on every one of three real machines tested to date as a special case, reducing each machine's unknowns to a small number of level-specic occupancy fractions. We then state precisely, without claiming to resolve, the paper's central open problem: whether those fractions are derivable from more primitive properties co-tenancy, private-cache-level count, associativity, prefetcher behavior or are fundamentally per-architecture constants. Four falsiable hypotheses are stated and tested against real hardware, with mixed results. We further state two limits of the framework explicitly: it requires dedicated, non-virtualized hardware access to be well-dened at all, and it extends only partway to a distributed-memory network, where realizing a tile across nodes requires a separate choice of communication algorithm the framework does not itself make. A rst, honest attempt at extending the framework toward predicting throughput directly, not just block size, closes the paper: two terms prove derivable from a specication sheet, one requires a single measurement, and one tested across three machines does not yet transfer between them.
Problem

Research questions and friction points this paper is trying to address.

hierarchical blocking
cache occupancy
machine shape
Mathematics of Arrays
block size prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mathematics of Arrays
hierarchical blocking
cache-aware scheduling
machine shape
occupancy fraction
L
Lenore M Mullin
College of Nanotechnology, Science, and Engineering, University at Albany, SUNY