Analyzing 10 Petabit/s Network Data with Accelerated Associative (Token) Arrays

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of processing multi-source heterogeneous data and overcoming hardware-software scalability bottlenecks in large-scale network privacy analysis. It pioneers the deep integration of the D4M associative array mathematical framework with GPU acceleration. By leveraging the MATLAB D4M library and parallel computing techniques, the proposed approach expresses complex logic through concise algorithms, transcends hardware generational limitations, and achieves efficient parallelism across vertical, horizontal, and temporal dimensions. Validated against MIT/IEEE/Amazon benchmarks, the system demonstrates linear horizontal scaling across hundreds of GPU nodes, sustaining processing rates on the order of 10 Petabits per second. Ultimately, this work provides a highly scalable hardware-software co-design solution for ultra-large-scale network data analysis.
📝 Abstract
As networks expand and become an ever more critical infrastructure to modern society the need to analyze these networks with the highest regard for privacy is essential to ensure their proper function. Depending on the level of the network layer to be analyzed, sources and destinations can be any combination of physical, logical, or persona/agentic endpoints, which requires the ability to handle diverse data. Invaluable to these analyses are mathematical tools that enable sophisticated mathematical algorithms to be expressed succinctly while achieving scalable vertical (within a compute node), horizontal (across compute nodes), and temporal (over different generations of hardware) performance. Associative (token) array mathematics and corresponding libraries is one approach that can meet these requirements. Accelerating these libraries with GPUs enables the analysis of the largest networks. The MIT/IEEE/Amazon Anonymized Network Sensing Graph Challenge provides a venue for highlighting the applicability of accelerated associative arrays for these types of problems. The D4M associative library has been implemented in a number of languages. This work benchmarks a prototype Matlab D4M GPU accelerated implementation of the Anonymized Network Sensing challenge across a wide range of CPU and GPU hardware. Scalable performance is demonstrated within and across CPU cores, CPU nodes, and GPU nodes. Horizontal scaling across multiple nodes was linear. Running on hundreds of GPU nodes simultaneously achieved a sustained processing rate sufficient to potentially analyze a 10 Petabit/s network.
Problem

Research questions and friction points this paper is trying to address.

Network Data Analysis
Privacy-Preserving
Scalable Performance
Petabit-scale Networks
Associative Arrays
Innovation

Methods, ideas, or system contributions that make the work stand out.

Associative Arrays
GPU Acceleration
Scalability
Network Analysis
D4M
Jeremy Kepner
Jeremy Kepner
MIT Lincoln Laboratory Supercomputing Center
high performance computingsupercomputingsignal processingmatlabgraph algorithms
H
Hayden Jananthan
MIT
L
LaToya Anderson
MIT
W
William Arcand
MIT
D
David Bestor
MIT
W
William Bergeron
MIT
C
Chansup Byun
MIT
A
Alex Bonn
MIT
D
Daniel Burrill
MIT
Vijay Gadepally
Vijay Gadepally
MIT
M
Michael Houle
MIT
M
Matthew Hubbell
MIT
M
Michael Jones
MIT
Piotr Luszczek
Piotr Luszczek
University of Tennessee
High Performance ComputingPerformance Evaluation and BenchmarkingNumerical Linear Algebra
P
Peter Michaleas
MIT
L
Lauren Milechin
MIT
J
Julie Mullen
MIT
A
Andrew Prout
MIT
A
Albert Reuther
MIT
A
Antonio Rosa
MIT
C
Charles Yee
MIT
A
Alex Pentland
MIT