Application-Driven Architecture Exploration for Cross-Layer Heterogeneous Systems

📅 2026-07-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of efficiently exploring the vast and physically constrained design space of cross-layer heterogeneous systems to support mixed AI and high-performance computing (HPC) workloads. To this end, the authors propose CHASE, a novel framework that decouples hardware architecture design from task mapping. CHASE leverages hierarchical type graphs for system modeling, a topology-aware mapper, and a telemetry-guided optimizer to enable application-driven architecture search under deployment constraints. Experimental results demonstrate that CHASE achieves geometric mean speedups of 6.20× and 2.12× on sparse computing and large language model workloads, respectively, while reducing mapping time by 60.5% on average and converging to near-global-optimal solutions within 64 iterations.
📝 Abstract
AI and HPC infrastructure increasingly serves workload portfolios that combine dense tensor computation, sparse kernels, large memory footprints, and communication-intensive collectives. Supporting these portfolios requires coordinated choices across accelerators, memory tiers, scale-up fabrics, and cluster networks. The resulting Cross-layer Heterogeneous System (XHS) design space is difficult to explore: hardware choices change legal task mappings, while rack power, switch radix, cabling, and cost constraints invalidate many candidates. We present CHASE, an application-driven framework that searches physically feasible XHS architectures through the workloads they must execute. CHASE represents candidates as hierarchical typed graphs and rejects designs that violate deployment constraints. It avoids intractable joint hardware-mapping search with a decoupled two-level loop: an inner mapper translates hardware-independent workload DAGs into topology-aware event traces, a calibrated event-driven simulator evaluates each mapping, and an outer telemetry-guided optimizer evolves the hardware graph. We evaluate CHASE on sparse-computing and LLM workloads. Its mapper remains within 6.06% of exhaustive optima while reducing mapping time by 60.5% on average relative to PEFT. Compute-model errors average 4.4-7.5%, and communication validation reproduces key trends across physical platforms. The outer search reaches near-global optima within 64 iterations. End-to-end case studies show that sparse workloads favor criticality-aware heterogeneous pods, whereas LLM inference favors scale-up islands; the resulting designs deliver 6.20$\times$ and 2.12$\times$ geomean speedups, respectively, while reducing cost and power relative to the baselines.
Problem

Research questions and friction points this paper is trying to address.

Cross-layer Heterogeneous System
Architecture Exploration
Workload Portfolio
Deployment Constraints
Hardware-Software Co-design
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-layer heterogeneous systems
application-driven architecture exploration
hierarchical typed graphs
decoupled two-level optimization
topology-aware mapping
🔎 Similar Papers
No similar papers found.
Yuchen Fan
Yuchen Fan
Shanghai AI Laboratory & Shanghai Jiao Tong University
NLPLarge Language ModelsEvaluation
M
Minghong Sun
Tsinghua University
J
Jikui Ma
Tsinghua University
Y
Yunpeng Xu
Tsinghua University
S
Shunyu Mao
Tsinghua University
L
Liu He
Tsinghua University
S
Shunan Dong
Tsinghua University
J
Jiahao Yang
Tsinghua University
Y
Yu Zhu
Tsinghua University
Xinhao Yang
Xinhao Yang
Tsinghua University
T
Tianyan Zhong
Tsinghua University
H
Haoran Sun
Tsinghua University
D
Daoqi Liu
Tsinghua University
Z
Zongle Huang
Tsinghua University
Xinyuan Lin
Xinyuan Lin
Tsinghua University
energy-efficient AI acceleratorscalable architecture
Huazhong Yang
Huazhong Yang
Professor of Electronics Engineering, Tsinghua University
VLSI circuits and systemsmachine intelligencewireless sensor networksbeyond-CMOS computing
Maokun Li
Maokun Li
Tsinghua University
Electromagnetic theorycomputational electromagneticsfast algorithmsinverse scattering problemscontrolled source electrom
Yongpan Liu
Yongpan Liu
Professor @ Tsinghua University
Machine LearningNonvolatile Memory and ComputingEnergy Efficient VLSIEmbedded SystemDesign Methodology
Y
Yu Wang
Tsinghua University
Z
Zhenhua Zhu
Tsinghua University
H
Hongyang Jia
Tsinghua University
S
Shuwen Deng
Tsinghua University