Score
Designs and analyzes mathematical and software models that represent computational demand separately from available resource supply and the abstractions that map between them, enabling composition of supply, context, and demand as independent components. Builds tooling and evaluation methods to explore cross‑layer tradeoffs, generate candidate resource or hardware specifications, and predict how changes in supply or abstraction affect system performance and capacity.
This work addresses three core challenges in hybrid computing involving foundation models (FMs) and symbolic programs: semantic misalignment, insufficient reliability, and scalability bottlenecks. Methodologically, we propose a capability-complementary hybrid computation offloading paradigm, underpinned by an infrastructure framework that supports dynamic task offloading and scheduling. The framework integrates task decomposition, resource-aware allocation, and adaptive optimization, while unifying formal verification of symbolic programs with FM-based inference interfaces. To our knowledge, this is the first approach to achieve organic synergy between FM-driven semantic understanding and the deterministic execution guarantees of symbolic programs. As a result, the system simultaneously achieves high accuracy, strong generalization, enhanced stability, and improved efficiency in large-scale data processing. This work establishes a foundational architectural blueprint for building efficient, reliable, and formally verifiable hybrid intelligent software systems.
This work addresses the challenge of resource allocation in geographically distributed and heterogeneous continuum computing infrastructures, where combinatorial explosion and limited generalization hinder effective deployment. To tackle this, the study introduces, for the first time, the pricing structures commonly found in Software-as-a-Service (SaaS) ecosystems into the resource allocation problem, formulating a unified, price-based representation of the configuration space. The authors propose PRIME, a pricing-aware analysis engine that efficiently searches for cost-optimal deployment configurations satisfying both functional and non-functional constraints. Leveraging synthetic infrastructure topologies and workload generation techniques, the project constructs a comprehensive dataset comprising 9,600 diverse scenarios, demonstrating that the proposed approach achieves both scalability and computational efficiency in complex, heterogeneous environments.
This work addresses key engineering challenges in the systematic evolution of large language models (LLMs)—including cache reuse, context management, agent scheduling, and access control—stemming from the absence of a unified architectural framework. To bridge this gap, the paper proposes the six-layer Intelligent Computing Architecture Model (ICAM), which unifies LLM-native system design through a dual-plane view comprising a probabilistic execution plane and a deterministic control plane. Grounded in three design principles—semantic locality, context budgeting, and agent acceleration—ICAM establishes interface contracts and design axioms for model-native systems. By integrating analogical analysis, architectural abstraction, and system modeling, the framework synthesizes advances in LLM-as-OS, memory management, multi-agent coordination, and security governance. System-level empirical validation confirms the efficacy of the proposed laws, elucidates parallels and divergences between model-native computing and traditional computer architecture, and outlines promising directions for future research.
Approximate computing faces fundamental challenges in jointly optimizing accuracy, energy efficiency, and performance for compute-intensive applications such as AI and digital signal processing (DSP). Method: This work proposes, for the first time, a unified classification framework and quantitative evaluation methodology integrating application-specific and microarchitectural-level approximation techniques. It establishes a full-stack approximation technology taxonomy—spanning algorithms, instruction sets, ALUs, compute-in-memory units, and configurable-precision accelerators—alongside an application-mapping model. Contribution/Results: Through systematic benchmarking of over 120 approximation techniques across 15 representative workloads—including image processing, speech recognition, and neural network inference—the study rigorously characterizes their applicability boundaries and achievable gains. The findings provide both theoretical foundations and practical guidelines for principled approximation selection and hardware-software co-design, enabling informed trade-offs in real-world deployment.
Emerging edge and cloud AI applications demand high-energy-efficiency computing, yet conventional embedded and datacenter architectures struggle to simultaneously achieve high performance and energy efficiency. Method: This work systematically surveys 15 years of approximate computing research, introducing the first full-stack taxonomy—spanning programs, compilers, circuits, accelerators, and memory—along with rigorously defined core terminology and design principles; it further proposes a unified evaluation framework for quantitative, cross-layer trade-off analysis between performance and power consumption. Contribution/Results: The study delivers the first authoritative survey on approximate computing (Part I), addressing a critical gap in systematic, domain-wide reviews. By establishing foundational taxonomies and evaluation methodologies, it provides both theoretical grounding and practical guidance for algorithm–architecture co-optimization, thereby advancing energy-efficient computing for AI workloads.
Traditional AI systems rely on fixed monolithic models, which struggle to dynamically allocate resources, decompose tasks, or update knowledge in response to varying inputs, leading to degraded performance and increased costs. This work proposes the first system-level design methodology for distributed composite AI systems, formulating a design space through workflow topologies and configuration choices and identifying eight core design patterns. The framework jointly optimizes model selection and runtime parameters, enabling task decomposition, multi-model orchestration, and explicit control logic, thereby facilitating a shift from static monolithic architectures toward dynamic, composable, and adaptive ones. Evaluated across three case studies, the approach reduces latency by up to 60% and cost by up to 71%, with only a 2.5–4 percentage point drop in accuracy.
This work addresses the significant degradation in inference throughput caused by GPU memory constraints when concurrently deploying multiple large language models on shared heterogeneous hardware, where resource scheduling, model offloading, and preemption become critical bottlenecks. Through empirical methodologies—including cross-platform performance profiling, layer-wise offloading experiments, and fine-grained decomposition of preemption overhead—the study systematically uncovers, for the first time, the nonlinear relationship between offloading and throughput decline. It further identifies model state reloading as the primary source of preemption overhead. The findings reveal that smaller models are more sensitive to reduced GPU residency, and that such overhead is jointly influenced by model architecture and hardware characteristics. These insights motivate a scheduler design that integrates model-specific sensitivity with data migration costs, offering crucial guidance for building efficient multi-model serving systems.
Existing tools struggle to efficiently and accurately perform full-stack architectural analysis of machine learning infrastructure spanning from microwatt-scale devices to gigawatt-scale data centers. This work proposes a first-principles-based analytical modeling framework that decouples computational demand from hardware supply and environmental context through a “demand–supply” abstraction. The framework introduces a “walls-of-systems” taxonomy, a dimensionally rigorous Python engine, 22 classes of system constraints, 28 composable solvers, and a typed input registry with provenance tracking to ensure unit consistency and traceability. It enables sub-second design space exploration, precisely identifies system bottlenecks, and automatically generates optimal hardware configurations covering the entire machine learning lifecycle.
This work addresses the lack of resource-centric computational efficiency metrics—specifically in terms of node-hours—for existing supercomputers and large-scale AI training platforms operating under high failure rates. It proposes the first efficiency evaluation framework grounded in resource consumption rather than execution time, unifying failure rate, mean time between failures, and checkpoint/restart overhead into a cohesive resource-based model. The framework extends Daly’s (2006) model to accommodate heterogeneous scientific workloads. Validated on one year of production data from the Frontier supercomputer, the approach leverages runtime log analysis, joint modeling of failures and checkpointing, and optimization algorithms to accurately quantify the expected fraction of resources usable for scientific computation and to determine optimal checkpoint intervals that minimize resource loss.