gpu acceleration

Designs and implements data-parallel algorithms, compute kernels, and runtime components to execute efficiently on graphics processing units, including memory layouts, thread/block mappings, and synchronization. Builds and integrates GPU-accelerated modules into software workflows and analyzes performance bottlenecks using profiling, optimization, and tuning techniques.

gpuacceleration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the unpredictability of GPU kernel execution times caused by inter-kernel data dependencies and resource contention, which undermines real-time guarantees. Focusing on DAG-structured GPU tasks, the paper proposes a scheduling approach that does not rely on kernel priority assumptions. By explicitly modeling kernel dependencies and co-scheduling kernel-level parallelism, the method derives tight and safe worst-case execution time (WCET) bounds. Built upon standard CUDA APIs, the approach requires only DAG task modeling, parallel kernel scheduling, and response time analysis, without additional hardware or software support. Experimental results on both synthetic and real-world benchmarks demonstrate that the proposed method reduces WCET and measured execution time by up to 32.8% and 21.3%, respectively, compared to existing techniques.

DAGdata dependenciesGPU tasks

Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs

Jan 16, 2025
JE
Jonah Ekelund
🏛️ KTH Royal Institute of Technology

Frequent fine-grained kernel launches on GPUs incur substantial launch overhead, severely limiting performance in scientific computing. To address this, we propose a synergistic optimization combining iterative batching and CUDA Graph unrolling: multiple iterations are grouped into batches and statically unrolled into a single CUDA Graph, thereby eliminating redundant kernel launch overhead. We further introduce the first platform-agnostic criterion for selecting the optimal batch size and develop a generalizable analytical performance model. Evaluated on skeleton applications, our approach achieves over 1.4× speedup. It demonstrates significant and robust performance improvements across real-world iterative GPU applications—including Hotspot, Hotspot3D, and an FDTD-based Maxwell solver—without requiring application-specific tuning. This work establishes a general, analytically tractable, low-overhead execution paradigm for iterative GPU computations.

CUDAGPUperformance bottleneck

Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems

Mar 13, 2025
FK
Fabian Knorr
🏛️ University of Innsbruck

SYCL programs on multi-GPU clusters suffer from high scheduling latency and substantial critical-path overhead due to implicit memory allocation, cache-coherence operations, and dependency analysis. Method: We propose the Instruction Graph—a novel intermediate representation that fully decouples scheduling from execution. Our approach integrates speculative scheduling, adaptive virtual-buffer memory allocation, and tight integration with the Celerity runtime, enabling fully concurrent scheduling of memory management, data transfers, MPI communication, and kernel launches while moving all scheduling analysis off the critical execution path. Contribution/Results: Evaluated on a production-scale 128-GPU cluster, our method achieves excellent strong scaling, drastically reduces multi-application scheduling latency, and drives critical-path overhead nearly to zero—thereby overcoming fundamental limitations of conventional static and blocking schedulers.

Enhancing memory allocation and concurrency in SYCL programs on accelerator clusters.Optimizing scheduling for high-level parallel programs on multi-GPU systems.Reducing delays in distributed-memory applications through graph-based representations.

A Performance Model for Warp Specialization Kernels

Jun 12, 2025
ZL
Zhengyang Liu
🏛️ University of Utah | NVIDIA

This work addresses the challenge of predicting execution performance for GPU warp-specialized kernels. We propose the first end-to-end performance model based on differential equations, jointly characterizing key factors including warp size, tiling dimensions, matrix dimensions, memory bandwidth, and thread divergence. The model is rigorously validated through both architectural analysis and empirical CUDA kernel measurements, augmented by a detailed bandwidth model. Its key innovation lies in the first application of differential equations to warp-level performance modeling, enabling quantitative characterization of the mapping between warp-level parallelism structures and performance bottlenecks. Experimental evaluation demonstrates a prediction error of less than 8.2%. The model supports compiler-driven auto-tuning and adaptive parameter configuration, achieving a 17% improvement in energy efficiency for sparse computation and GEMM workloads.

Develops a performance model for warp specialization kernelsOptimizes GPU applications via compiler and kernel tuningPredicts execution time using differential equations and validation

GPU-Accelerated Algorithms for Process Mapping

Oct 14, 2025
PS
Petr Samoldekin
🏛️ Heidelberg University

This work addresses the classical task graph mapping problem onto processing units in supercomputers, aiming to balance computational load and minimize inter-task communication overhead. For the first time, GPU acceleration is introduced into this domain, yielding two parallel algorithms: (1) a hierarchical multi-partitioning framework accelerated on GPUs, and (2) a GPU-accelerated multilevel graph partitioning implementation integrating optimized coarsening and refinement strategies. Experiments demonstrate speedups of up to 598× over state-of-the-art CPU-based solvers, with a geometric mean speedup of 77.6×; Algorithm (1) incurs only ~10% increase in communication cost while maintaining competitive solution quality. The core contribution is the establishment of a novel GPU-parallel paradigm for task mapping—breaking through long-standing performance bottlenecks inherent in traditional CPU-centric approaches.

GPU-accelerated algorithms balance computational workload and minimize communication costsHierarchical multisection partitions task graphs using supercomputer hierarchyMultilevel graph partitioning pipeline accelerates coarsening and refinement phases

Latest Papers

What's happening recently
View more

This work addresses the critical gap in existing performance analysis tools, which largely lack support for virtual GPUs (vGPUs) and thus struggle to diagnose performance bottlenecks in vGPU-accelerated applications. To tackle this limitation, the authors design and implement a novel vGPU performance profiling tool tailored for Intel GVT-g. The tool introduces, for the first time, a fine-grained software tracing mechanism coupled with runtime data collection to generate multidimensional performance metrics. These metrics are then presented through synchronized multi-view visualizations that capture comprehensive vGPU behavioral characteristics. By enabling detailed insight into vGPU execution dynamics, this approach significantly enhances observability and debugging efficiency in GPU-virtualized environments, effectively filling a longstanding void in vGPU performance analysis tooling.

debuggingGPU virtualizationperformance analysis

This work addresses the throughput limitations of GPU systems caused by host-device synchronization latency and kernel scheduling overhead, which hinder efficient utilization of compute cores and copy engines. The authors propose a CUDA runtime framework tailored for task-parallel pipelining, which innovatively integrates multi-stream scheduling, event-chain triggering, work stealing, and stream-level buffer management. This design ensures memory safety across concurrent tasks while substantially reducing kernel launch intervals and synchronization overhead. Implemented using CUDA Graphs, the framework achieves 1.15–1.44× speedup over state-of-the-art baselines on real-world workloads and reduces scheduling overhead by 18%–54%.

CUDA graphGPU performancehardware underutilization

This work addresses the challenge that existing GPU performance models struggle to accurately simulate highly optimized large language model (LLM) kernels, which rely on fine-grained scheduling and compute–memory overlap, as traditional simulators are computationally expensive while analytical models are overly coarse. To bridge this gap, the authors propose a tile-centric simulation framework that, for the first time, models LLM kernels as tile-level dependency graphs. By combining an automated frontend for graph construction with a graph-driven backend simulator, the approach efficiently captures execution dependencies and overlapping behaviors. The framework supports GEMM, attention mechanisms, and end-to-end LLM inference, achieving average absolute percentage errors of 1.22%–8.71% on A100/H100 GPUs and successfully generalizing to the Blackwell architecture. This advancement significantly enhances simulation accuracy and scalability, facilitating effective hardware–software co-design for LLMs.

GPU simulationhardware-software co-designlarge language models

Hot Scholars

MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
FT

Federico Tombari

Google, TU Munich
Computer VisionMachine Learning3D Perception
PA

Pieter Abbeel

UC Berkeley | Covariant
RoboticsMachine LearningAI
ZX

Zexiang Xu

Hillbot
Computer VisionComputer GraphicsFoundation Models
ML

Mingrui Li

Dalian University of Technology
SLAM3D VisionRobotics