Score
Designs, implements, and analyzes algorithms and parallel kernels that map particle-based quantities onto grid-based fields (particle-to-grid or p2g mapping), optimizing interpolation, data layout, and reductions to maximize throughput and resource utilization while removing computational bottlenecks; ensures these p2g transfers preserve numerical accuracy across CPU, GPU, or other parallel hardware implementations.
Addressing the challenge of simultaneously achieving high compression throughput and controllable reconstruction error for irregular, large-scale particle-based simulations and point-cloud data on modern GPUs, this paper proposes a hardware-aware, error-bounded lossy compression framework. Our method introduces an innovative four-stage parallel pipelined architecture that jointly optimizes kernel scheduling, memory access patterns, and Streaming Multiprocessor (SM) occupancy. It integrates GPU-accelerated parallel entropy coding, adaptive per-block error control, and fine-grained memory layout optimization. Evaluated on six real-world scientific datasets, our framework outperforms five state-of-the-art GPU compressors: it achieves up to 8× higher end-to-end throughput while delivering superior compression ratios and reconstruction fidelity—approaching the theoretical hardware throughput limit of contemporary GPUs.
This work addresses the classical task graph mapping problem onto processing units in supercomputers, aiming to balance computational load and minimize inter-task communication overhead. For the first time, GPU acceleration is introduced into this domain, yielding two parallel algorithms: (1) a hierarchical multi-partitioning framework accelerated on GPUs, and (2) a GPU-accelerated multilevel graph partitioning implementation integrating optimized coarsening and refinement strategies. Experiments demonstrate speedups of up to 598× over state-of-the-art CPU-based solvers, with a geometric mean speedup of 77.6×; Algorithm (1) incurs only ~10% increase in communication cost while maintaining competitive solution quality. The core contribution is the establishment of a novel GPU-parallel paradigm for task mapping—breaking through long-standing performance bottlenecks inherent in traditional CPU-centric approaches.
GPU code optimization remains a critical bottleneck in high-performance computing and large model training and inference, as existing approaches struggle to consistently approach hardware performance limits. This work proposes a two-stage GPU kernel tuner: it first transforms the original kernel into a parameterized template through semantic restructuring, then optimizes the template parameters using a performance-feedback-driven constrained search strategy. By integrating an LLM-agent-guided iterative workflow with a synergistic mechanism of templated rewriting and search-based tuning, the method significantly enhances optimization stability and interpretability while reducing manual intervention. Evaluated on real-world CUDA kernels, it achieves up to 3× speedup over baseline implementations, outperforms pure LLM-based rewriting approaches, and demonstrates strong potential for extension to other backends such as OpenCL and HIP.
This work addresses the severe performance bottleneck imposed by fine-grained atomic updates in particle-in-cell simulations on conventional multicore CPUs. Targeting modern CPU architectures featuring mixed MPU-VPU SIMD units, the authors propose the first holistic co-design for current deposition, centered on matrix outer products. The approach integrates blocked matrix algorithms, a hybrid execution pipeline, an incremental sorter with O(1) amortized complexity, and a gap-compressed memory layout to optimize data access. Evaluated on laser wakefield acceleration simulations, the method achieves an 8.7× speedup in the third-order deposition kernel—reaching 83.08% of the CPU’s theoretical peak performance—and delivers a 2.63× end-to-end runtime acceleration, substantially outperforming existing GPU implementations.
This work addresses the high cost and low efficiency associated with parallelizing and modernizing legacy scientific computing codes by proposing a structured AI agent approach. By integrating large language model agents, manual prompt engineering, and continuous integration—guided by human-provided examples, guaranteed buildability, and constrained dialogue scope—the method successfully refactored the 60,000-line single-threaded Fortran MPI code NMAP-RKPM into a C++ MPI tool with OpenMP support within two phases over several months. This effort demonstrates the feasibility and substantial effectiveness of a structured AI-assisted paradigm for large-scale high-performance computing (HPC) software modernization.
Data transfer between the GPU and host memory is significantly slower than computational speed, becoming a major performance bottleneck for SPH solvers. To address this, this work proposes a host-side particle memory layout optimization tailored for GPU offloading. By analyzing GPU kernel access patterns and particle attribute types, the conventional Array-of-Structures (AoS) layout is decomposed into multiple fine-grained sub-structures (Split AoS), combined with a data compression strategy to substantially reduce the overhead of data reorganization before and after transfers. Experimental results demonstrate that the proposed approach reduces data packing time by 20%–40% and decreases overall GPU offloading latency by 12%–25%, thereby significantly enhancing heterogeneous computing efficiency.
This study investigates whether general-purpose coding agents can undertake rigorous scientific performance engineering. Leveraging Codex and Claude Code, the research introduces an “executable scientific contract” mechanism to guide and validate hypothesis-driven optimization experiments for fixed-radius nearest neighbor search algorithms, autonomously refactoring a PyTorch implementation into a dependency-free, high-performance C++/CUDA library. The results demonstrate that goal-directed agents can effectively assume the role of experimental performance engineers. The generated standalone library precisely reproduces the original results while achieving a 1.6× speedup over the baseline GPU-accelerated PyTorch implementation through its synchronous NumPy interface, with consistent performance maintained across diverse hardware architectures.
This project addresses the energy efficiency constraints and precision challenges introduced by heterogeneous accelerators in scientific computing. With "energy consumption per trusted solution" as its core objective, it establishes a mixed-precision computing framework. This work innovatively proposes a "reckless yet responsible" computing paradigm that integrates novel number formats, floating-point emulation, hardware-software co-design, and multi-level resource management to effectively balance aggressive low-precision arithmetic with system-level detection and verification. Furthermore, the project systematically reviews the technological landscape and development trajectories of this field, distills a list of open problems, and provides comprehensive design guidelines. Ultimately, it offers both a theoretical foundation and practical reference for next-generation energy-efficient scientific computing.
This study addresses the significant efficiency gap between LLM-generated GPU kernels and expert implementations, where the specific impact of design guidance on performance remains difficult to quantify. To investigate this, we construct a diagnostic benchmark and propose a hierarchical design guidance framework that establishes an evaluation mechanism mapping algorithmic insights to low-level optimizations. This approach systematically examines the capacity of LLM agents to translate multi-granularity expert knowledge into efficient kernels. Comprehensive evaluations are conducted on NVIDIA B200 hardware using paired experiments and multidimensional metrics. The results demonstrate that incorporating structured guidance elevates code correctness to 98.5% and achieves a geometric mean speedup of 2.49×, substantially narrowing the performance disparity between LLM-generated kernels and expert-crafted implementations.
本文通过将自动调优集成到Julia编写的硬件无关GPU内核中,解决了跨不同架构高效执行的难题,显著提升了性能。