Score
Designs, builds, and analyzes interconnect architectures and their components — including physical links, topologies, switches, and protocol layers — for high-speed, high-bandwidth, high-performance and specialized environments (e.g., cryogenic systems and on‑chip/network‑on‑chip fabrics). Develops models, simulators and benchmarks and performs performance, power and latency estimation, tuning, optimization and scheduling to meet target throughput, latency, energy and reliability constraints.
To address communication bottlenecks—both on-chip and inter-node—in post-exascale supercomputers and AI data centers, this paper proposes a unified, scalable interconnect architecture spanning chip-level and system-level hierarchies. Methodologically, it introduces a novel low-diameter network topology, fine-grained flow control mechanisms, and heterogeneous resource co-sharing strategies, tightly integrated with high-bandwidth memory and accelerator hardware. Its key contribution lies in jointly optimizing communication latency and bandwidth, substantially alleviating resource contention and improving data locality. Experimental evaluation at scale—up to 1,000 accelerators—demonstrates a 32–47% reduction in communication overhead, a 2.1× increase in system throughput, and a 38% improvement in energy efficiency. The architecture thus delivers efficient, scalable interconnect support for generative AI workloads and large-scale scientific simulations.
This work addresses the lack of an efficient exploration framework in early-stage 2.5D packaging design that jointly considers package and interconnect selection. It presents the first automated interface IP generation methodology enabling co-optimization of packaging and chiplet architectures. The proposed approach rapidly evaluates power, performance, and area across diverse 2.5D packaging and communication configurations, and automatically produces standard design assets—including Verilog, Liberty, LEF, and datasheets—that comply with protocols such as UCIe. By bridging the gap between high-fidelity and highly flexible interconnect modeling, this method significantly enhances the efficiency of system architecture exploration and the accuracy of design decisions.
To address data transmission unreliability in Network-on-Chip (NoC) systems induced by power supply noise (PSN), this paper proposes a modular modeling and probabilistic verification methodology based on the Modest language. The method integrates modular router models with a hierarchical verification framework, enabling unified formal verification of functional correctness and PSN sensitivity—from individual routers up to 8×8 NoC topologies. Leveraging the Modest Toolset, we perform rigorous formal verification and statistical model checking to quantitatively assess communication consistency, functional reliability, and noise robustness. Compared to conventional approaches, our methodology significantly enhances verifiability, scalability, and model reusability at early design stages. It establishes a novel, formally grounded modeling paradigm for high-reliability NoC design under heterogeneous and dynamic traffic conditions.
Existing NoC routing designs for multicore systems overlook cache-coherence traffic, leading to inaccurate performance evaluation. To address this, this paper proposes the first coherence-aware co-optimization framework for routing and topology in NoCs. We introduce the Cache Coherence Traffic Analyzer (CCTA), a novel tool that accurately models coherence traffic under protocols such as MESI. Our framework integrates a traffic-learning-driven dynamic routing algorithm, protocol-aware adaptive topology selection, and a joint latency-energy optimization mechanism. Evaluated on standard benchmarks, our approach reduces packet latency by 10.52%, accelerates application execution time by 55.51%, and cuts total energy consumption by 49.02% over baseline methods. This work pioneers the deep integration of coherence communication modeling into joint NoC routing and topology design—enabling significant improvements in both system energy efficiency and real-time performance.
To address the challenges of assessing non-shortest-path diversity in large-scale interconnection networks and the poor scalability of conventional packet-level simulators, this paper proposes a lightweight simulation framework tailored for extreme-scale networks. By identifying memory and event-scheduling bottlenecks in mainstream simulators, we introduce three core techniques: compact data structures, lazily bound event queues, and lock-free memory pools—significantly reducing both memory footprint and synchronization overhead. Our framework enables fine-grained, packet-level simulation of data center and HPC networks with over one million endpoints on a single commodity laptop, achieving a throughput of 10 million packets per second—three orders of magnitude higher than state-of-the-art shared-memory simulators. The open-source framework supports rapid prototyping and validation of novel interconnect protocols, providing a reproducible, high-fidelity foundation for path diversity analysis and performance optimization in ultra-large-scale networks.
该论文通过分类和比较AI加速器架构,探讨了如何在通用性和专业化之间取得平衡以应对AI数据中心面临的挑战。
本文针对高密度和引脚数瓶颈导致的供电难题,提出设计多级分布式供电系统的方法,以支持复杂3D异构集成系统的可靠电力供应。
研究通过在低延迟、宽链路NoC路由器上实施三种粒度的三模冗余(TMR)来解决单事件效应(SEEs)导致的网络故障问题。
针对可重排非阻塞光互连的电路调度问题,提出Tetris算法,通过优先处理瓶颈端点并保持现有连接来减少通信延迟。
本文针对FPGA中NoC拥塞问题,通过集成拥塞成本、转弯模型路由算法及SAT求解方法,有效减少了95.1%的网络拥塞。