Score
Designs and implements algorithmic code and system components that preserve required time and space behavior as input size, concurrency, or data throughput grow. This work includes building optimized computation kernels, memory and I/O layouts, parallel and distributed execution paths, and resource-management strategies to achieve high throughput, low latency, and controlled resource usage.
This work investigates the evolutionary trajectory of algorithmic space complexity across 118 core problems in computer science, encompassing over 800 algorithms. Method: Leveraging a large-scale literature survey, historical complexity data analysis, and theoretical evaluation, the study quantifies trends in memory efficiency improvements relative to hardware advances. Contribution/Results: It is the first to empirically demonstrate that, in 20% of cases, algorithmic space optimization outpaces DRAM latency reduction. The paper introduces the “time–space trade-off Pareto frontier” framework to characterize optimal algorithmic trade-offs over time. Findings confirm that memory efficiency has emerged as a critical constraint in modern algorithm design. To support reproducible research and engineering practice, the authors release an open-source algorithm knowledge base (https://algorithm-wiki.csail.mit.edu), providing standardized benchmarks and decision-support tools for both theoretical analysis and system implementation.
Manual performance optimization in hyperscale data centers is costly, error-prone, and unscalable. Method: This paper introduces the first end-to-end automated code optimization framework, integrating (i) a historical-commit-driven performance anti-pattern dictionary and (ii) a domain-finetuned large language model (LLM) to generate trustworthy refactoring proposals; optimization safety is ensured via automated validation and production-grade A/B testing. Contribution/Results: Deployed in Google’s production environment across >100 million lines of code, the framework achieves >99.5% optimization success rate, with 6,400+ validated optimization commits modifying 25,000 lines of code. It delivers an average quarterly saving of over 500,000 normalized CPU cores. This work establishes the first empirically validated, high-reliability, and scalable AI-driven performance optimization paradigm for hyperscale production systems.
Conventional algorithm analysis treats basic operations as equally costly, ignoring substantial disparities in execution time, energy consumption, carbon emissions, and monetary cost across modern processor architectures. Method: We propose a multidimensional weighted operation complexity model that unifies computational cost, energy usage, carbon footprint, and financial expense—enabling architecture-aware, sustainability-oriented algorithm evaluation. Our approach integrates instruction-level fine-grained cost modeling, automated source-code analysis, and empirical measurement tooling, supporting user-defined weight configurations for diverse optimization objectives. Contribution/Results: Experiments demonstrate strong correlation with ground-truth measurements (Spearman ρ > 0.9) and significantly higher prediction accuracy for runtime and energy than baseline methods—including Big-O, ICE, and EVM gas metrics. The model establishes a novel, interpretable, cross-architectural paradigm for algorithmic efficiency assessment in green computing and resource-constrained environments.
Manual pragma configuration in high-level synthesis (HLS) suffers from low efficiency and an exponentially large search space. Method: This paper proposes the first nonlinear programming (NLP)-based automated pragma insertion framework, jointly optimizing loop-level pragmas—including pipelining, function unit replication, and data caching. It innovatively models discrete pragma configurations as continuous, differentiable variables and constructs analytical performance/resource models with theoretical lower-bound guarantees, solved globally via NLP. Integrated with pragma semantic analysis and the Merlin compiler, and augmented by design-space pruning, the framework explores billion-scale configurations within seconds to minutes. Contribution/Results: Experimental evaluation shows kernel performance approaching hand-tuned implementations, resource estimation error <8%, and latency lower-bound error ≤12%.
In high-level synthesis (HLS), jointly optimizing code transformations, pragma insertion, and cache-blocking size selection is challenging due to tight coupling, a vast decision space, and difficulty in guaranteeing semantic correctness. Method: This paper proposes the first unified modeling framework that jointly encodes all three aspects as a single, isomorphic optimization problem—supporting “zero-transformation” decisions—and leverages HLS compiler–driven constraint derivation coupled with nonlinear programming (NLP) to automatically and correctly optimize regular loop nests. Contribution/Results: It introduces the first paradigm for co-optimizing transformations, pragmas, and blocking sizes, with built-in semantic equivalence preservation. Evaluated on multiple benchmark kernels, the approach significantly improves quality-of-results (QoR), accurately identifies cases requiring or forbidding transformations, and generates high-performance, formally verifiable optimized code.
This work addresses the challenges of workflow task composition in high-throughput, petabyte-scale data processing environments, where resource heterogeneity and execution overhead significantly impact performance. The authors propose a hybrid task composition strategy that dynamically balances task independence against execution grouping, formulated within a multi-objective optimization framework to achieve Pareto-optimal trade-offs among throughput, I/O cost, and CPU efficiency. Leveraging workflow DAG modeling and high-dimensional parameter space simulation, the approach enables policy-driven automated synthesis of workflows. Experimental results demonstrate that the proposed strategy achieves up to a 3.8× improvement in throughput and reduces network overhead by as much as 14.9× compared to baseline methods, offering a scalable workflow synthesis framework for extreme-scale scientific computing.
This work systematically evaluates the potential of large language models (LLMs) for automatic code optimization in high-performance computing (HPC), where traditional approaches often struggle to balance performance and correctness. The study introduces a novel methodology that leverages multi-level abstractions and goal-oriented prompting to guide LLMs in directly generating optimized C code. Evaluated on the PolyBench benchmark suite, this approach is compared against conventional auto-tuning frameworks that rely on schedule representations. Experimental results demonstrate that LLM-generated C code achieves superior performance and effectiveness, highlighting the critical influence of compiler optimization abstractions on LLM guidance. These findings establish a promising new direction toward verifiable, LLM-driven code optimization for HPC applications.
This study addresses the challenges AI-driven scientific workflows encounter in high-performance computing (HPC) environments—namely data intensity, resource heterogeneity, frequent iteration cycles, and I/O bottlenecks—which hinder their compatibility with conventional linear pipeline architectures. The work presents the first systematic integration of AI workflow characteristics with HPC system design, proposing a transformative framework tailored for adaptive intelligent computing environments and articulating twelve practical design principles. This framework incorporates key technologies including containerization, job array scheduling, explicit feedback mechanisms, heterogeneous resource management, and small-file I/O optimization, making it particularly well-suited for high-throughput domains such as computational biology. The resulting guidelines offer researchers actionable strategies to substantially enhance the efficiency, portability, and scalability of AI-HPC workflows.
This work addresses the lack of resource-centric computational efficiency metrics—specifically in terms of node-hours—for existing supercomputers and large-scale AI training platforms operating under high failure rates. It proposes the first efficiency evaluation framework grounded in resource consumption rather than execution time, unifying failure rate, mean time between failures, and checkpoint/restart overhead into a cohesive resource-based model. The framework extends Daly’s (2006) model to accommodate heterogeneous scientific workloads. Validated on one year of production data from the Frontier supercomputer, the approach leverages runtime log analysis, joint modeling of failures and checkpointing, and optimization algorithms to accurately quantify the expected fraction of resources usable for scientific computation and to determine optimal checkpoint intervals that minimize resource loss.
This work proposes “grid programs,” a two-dimensional computational model grounded in an integer lattice, which overcomes limitations of traditional models constrained by linear instruction sequences, named variables, and explicit memory addresses. In this paradigm, computation proceeds as an instruction pointer traverses the grid in four cardinal directions, while program state is maintained through a data stack, an address stack, and a three-pointer cyclic doubly linked list. The model enforces no variable names or syntactic constraints, relying instead on purely spatial control flow. It constitutes the first Turing-complete computational framework that is entirely free of named variables and defined solely by spatial layout. Formal operational semantics demonstrate its ability to simulate any register machine, and practical implementations—including factorial computation and string reversal—highlight its expressiveness. The approach shows promising applications in visual programming, cellular-automaton-inspired hardware, and code obfuscation resistance.