Score
Designs and computes quantitative estimates and bounds that relate computational cost (runtime, number of samples, or operation counts) to estimator precision (MSE or other error measures), producing cost‑to‑precision curves, sample‑size requirements, and finite‑sample or asymptotic complexity estimates. Analyzes and compares runtime/sample‑size tradeoffs to derive cost bounds used to choose or optimize algorithms and samplers.
Emerging edge and cloud AI applications demand high-energy-efficiency computing, yet conventional embedded and datacenter architectures struggle to simultaneously achieve high performance and energy efficiency. Method: This work systematically surveys 15 years of approximate computing research, introducing the first full-stack taxonomy—spanning programs, compilers, circuits, accelerators, and memory—along with rigorously defined core terminology and design principles; it further proposes a unified evaluation framework for quantitative, cross-layer trade-off analysis between performance and power consumption. Contribution/Results: The study delivers the first authoritative survey on approximate computing (Part I), addressing a critical gap in systematic, domain-wide reviews. By establishing foundational taxonomies and evaluation methodologies, it provides both theoretical grounding and practical guidance for algorithm–architecture co-optimization, thereby advancing energy-efficient computing for AI workloads.
Approximate computing faces fundamental challenges in jointly optimizing accuracy, energy efficiency, and performance for compute-intensive applications such as AI and digital signal processing (DSP). Method: This work proposes, for the first time, a unified classification framework and quantitative evaluation methodology integrating application-specific and microarchitectural-level approximation techniques. It establishes a full-stack approximation technology taxonomy—spanning algorithms, instruction sets, ALUs, compute-in-memory units, and configurable-precision accelerators—alongside an application-mapping model. Contribution/Results: Through systematic benchmarking of over 120 approximation techniques across 15 representative workloads—including image processing, speech recognition, and neural network inference—the study rigorously characterizes their applicability boundaries and achievable gains. The findings provide both theoretical foundations and practical guidelines for principled approximation selection and hardware-software co-design, enabling informed trade-offs in real-world deployment.
This work addresses the problem of deriving provably tight floating-point rounding error bounds for numerical programs featuring conditional branches, no loops, and mixed-precision arithmetic. Methodologically, it unifies the modeling of conditional control flow and precision heterogeneity via two novel quantitative metrics—“instability jumps” and “window width”—and integrates interval arithmetic, abstract interpretation, and precision-aware semantic modeling, augmented with abstraction-guided global optimization. Its key contribution is the first formal framework enabling joint, compositional analysis of conditional branching and mixed precision, achieving both high bound tightness and practical analysis efficiency. Experimental evaluation on standard benchmarks demonstrates significantly tighter error bounds compared to prior approaches. Furthermore, the framework successfully guides precision configuration—e.g., step size and search direction—in the conjugate gradient method, empirically validating its utility in supporting design-time trade-offs among accuracy, error bounds, and computational efficiency.
This paper addresses the forward error analysis of sum-product algorithms under stochastic rounding (SR). We propose a probabilistic error bounding method grounded in martingale theory. Our key contributions are threefold: (1) We introduce the first automated martingale construction framework tailored to multilinear computational structures—encompassing addition, subtraction, multiplication, and intermediate result reuse; (2) We extend SR error analysis to algorithms with structural reuse, notably Karatsuba polynomial multiplication—previously unaddressed in SR literature; (3) Leveraging the Azuma–Hoeffding inequality, we derive a tight probabilistic error bound of $O(sqrt{n},u)$, markedly improving upon the classical worst-case bound $O(n,u)$. Our framework uniformly recovers known error guarantees for pairwise summation and Horner’s method, and—crucially—provides the first rigorous SR error guarantee for Karatsuba multiplication.
This work introduces, for the first time, fundamental thermodynamic limits into basic machine learning algorithms by leveraging Landauer’s principle and information thermodynamics to quantify the minimum irreversible energy dissipation incurred by floating-point implementations of simple linear regression. By constructing an entropy production model for continuous inputs, the study analyzes the thermodynamic costs associated with both exact solutions and stochastic gradient descent. Furthermore, it derives the optimal scaling law between training set size and energy consumption under a prescribed generalization error constraint. This paper establishes the first theoretical framework characterizing the trade-off between energy efficiency and generalization in linear regression, thereby providing a physical foundation for the design of energy-aware machine learning systems.
Existing runtime monitors support only Boolean specification verification, making it infeasible to progressively approximate quantitative properties—such as average response time—over infinite traces. Method: This paper establishes the first unified formal framework for quantitative approximate monitoring, introducing quantitative monitors whose estimates monotonically improve as observation prefixes grow, and rigorously modeling the trade-off between estimation accuracy and resource consumption (specifically, register count). Contribution/Results: We prove that register count strictly determines the theoretical upper bound on achievable accuracy; moreover, each additional register strictly increases the attainable precision—demonstrating an irreducible, non-compensatory relationship between resources and accuracy. Our framework conservatively extends classical Boolean monitoring theory while ensuring soundness. The proposed approach provides provably optimal, resource-bounded approximate monitoring for critical performance metrics, enabling verifiable, deployment-aware runtime assurance.
This project addresses the energy efficiency constraints and precision challenges introduced by heterogeneous accelerators in scientific computing. With "energy consumption per trusted solution" as its core objective, it establishes a mixed-precision computing framework. This work innovatively proposes a "reckless yet responsible" computing paradigm that integrates novel number formats, floating-point emulation, hardware-software co-design, and multi-level resource management to effectively balance aggressive low-precision arithmetic with system-level detection and verification. Furthermore, the project systematically reviews the technological landscape and development trajectories of this field, distills a list of open problems, and provides comprehensive design guidelines. Ultimately, it offers both a theoretical foundation and practical reference for next-generation energy-efficient scientific computing.
This study addresses the computational bottlenecks in scientific computing arising from the infeasibility of exact algorithms for large-scale problems. Through a systematic evaluation of approximation methods across 118 core algorithmic problems—integrating complexity analysis, taxonomies of approximation algorithms, and historical context—the work presents the first large-scale empirical evidence demonstrating that only approximately 20% of these problems derive substantial benefit from approximation. Notably, one-quarter of exponential-time-hard problems admit polynomial-time approximation schemes, and the adoption of approximation strategies increases the proportion of linear-time solvable problems by 23%. By quantifying the trade-offs between accuracy and efficiency, this research offers theoretical insights to guide the design of AI-driven and high-performance algorithms.
This study addresses the absence of preregisterable minimum detectable effect (MDE) standards in quantized model evaluation, which hinders disentangling quantization noise from other sources of variation. Adapting paired binomial sample size methodology to 4-bit quantization benchmarks, the work proposes a conservative upper-bound formula for MDE to enable effect-size budgeting in experimental design and validates its efficacy through pilot audits. Innovatively integrating a preregistrable MDE framework with the Miettinen test, FP16–NF4 disagreement rate modeling, cross-model and cross-benchmark experiments, and variance decomposition by prompt template, the analysis reveals that most NF4–FP16 performance differences fall below the MDE threshold. In contrast, prompt templates induce performance fluctuations of 2–10 percentage points—substantially exceeding quantization effects—thereby underscoring the critical need to control for prompt-induced variability.
This study addresses the loss of precision caused by rounding errors in finite element computations, which remains difficult to analyze a priori. We propose the first automated a posteriori rounding error estimation framework that leverages running error analysis to track numerical and error propagation in real time. Built upon the FEniCS Form Compiler, this method achieves the first automated error estimation within finite element kernels through a C++ backend, custom arithmetic types, and templated kernel generation techniques. Experimental results demonstrate that the framework successfully detects catastrophic cancellation with only a 2–4× performance overhead. By effectively supporting mixed-precision design and numerical debugging, this work establishes a new paradigm for ensuring reliability in scientific computing.
Quantum program verification on early fault-tolerant hardware faces critical challenges due to scarce measurement resources and tight measurement budgets. Method: We propose the first unified, program-level measurement budgeting framework that systematically links theoretical error bounds—based on trace distance, fidelity, error probability, and the quantum Chernoff bound—with practical testing strategies: inversion testing, swap testing, and chi-square testing. The framework supports scalable analysis, from single-gate verification to full-program validation. Contribution/Results: We quantify substantial measurement overhead differences among strategies: inversion testing is optimal; swap testing incurs roughly 2× overhead; chi-square testing is simple but costly. Noise and fine-grained circuit decomposition further escalate costs. To address this, we introduce coarse-grained partitioning and weighted budget allocation, achieving superior trade-offs between verification accuracy and hardware resource consumption. Our framework establishes a computationally tractable, deployable paradigm for quantifying and allocating measurement resources in quantum software verification.
Existing approaches to program resource analysis struggle to simultaneously achieve the completeness of static analysis and the worst-case coverage afforded by dynamic analysis. To address this limitation, this work proposes a hybrid analysis method that integrates dynamic symbolic execution with mixed-integer linear programming to systematically enumerate execution paths within a bounded input space and derive empirically sound upper bounds on maximum resource consumption. This approach represents the first deep integration of dynamic symbolic execution and linear programming for inferring tight and effective worst-case resource bounds for functional programs. The prototype tool CompAS demonstrates both practical utility and theoretical guarantees in estimating resource usage on complex programs.