A Unified Metric Architecture for AI Infrastructure: A Cross-Layer Taxonomy Integrating Performance, Efficiency, and Cost

📅 2025-11-25
📈 Citations: 0
Influential: 0
📄 PDF

career value

205K/year
🤖 AI Summary
AI infrastructure confronts multidimensional physical and economic constraints—including power, thermal management, water usage, interconnect bandwidth, memory capacity, and data throughput—while existing metrics (e.g., PUE, TCO) are siloed and fail to capture the coupled trade-offs among energy efficiency, performance, and cost, hindering cross-layer co-optimization. To address this, we propose a unified measurement architecture grounded in a 6×3 cross-layer taxonomy—spanning facility, network, compute, storage, software, and application layers, each annotated with physical, computational, and economic semantics—and introduce the Measurement Propagation Graph (MPG) to enable, for the first time, system-level, three-dimensional relational modeling. Leveraging systematic literature review, meta-analysis, and graph-based modeling, our framework integrates heterogeneous, multi-source metrics. It supports benchmarking, capacity planning, and total cost of ownership analysis, substantially enhancing interpretability of AI cluster efficiency frontiers and enabling rigorous multi-objective optimization.

Technology Category

Application Category

📝 Abstract
The growth of large-scale AI systems is increasingly constrained by infrastructure limits: power availability, thermal and water constraints, interconnect scaling, memory pressure, data-pipeline throughput, and rapidly escalating lifecycle cost. Across hyperscale clusters, these constraints interact, yet the main metrics remain fragmented. Existing metrics, ranging from facility measures (PUE) and rack power density to network metrics (all-reduce latency), data-pipeline measures, and financial metrics (TCO series), each capture only their own domain and provide no integrated view of how physical, computational, and economic constraints interact. This fragmentation obscures the structural relationships among energy, computation, and cost, preventing a coherent optimization across sector and how bottlenecks emerge, propagate, and jointly determine the efficiency frontier of AI infrastructure. This paper develops an integrated framework that unifies these disparate metrics through a three-domain semantic classification and a six-layer architectural decomposition, producing a 6x3 taxonomy that maps how various sectors propagate across the AI infrastructure stack. The taxonomy is grounded in a systematic review and meta-analysis of all metrics with economic and financial relevance, identifying the most widely used measures, their research intensity, and their cross-domain interdependencies. Building on this evidence base, the Metric Propagation Graph (MPG) formalizes cross-layer dependencies, enabling systemwide interpretation, composite-metric construction, and multi-objective optimization of energy, carbon, and cost. The framework offers a coherent foundation for benchmarking, cluster design, capacity planning, and lifecycle economic analysis by linking physical operations, computational efficiency, and cost outcomes within a unified analytic structure.
Problem

Research questions and friction points this paper is trying to address.

Unifies fragmented metrics across AI infrastructure layers
Integrates physical, computational, and economic constraints into one framework
Enables multi-objective optimization of energy, carbon, and cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified cross-layer taxonomy for AI infrastructure metrics
Metric Propagation Graph formalizes cross-layer dependencies
Integrated framework links physical, computational, and economic constraints