particle-to-grid optimization

Designs, implements, and analyzes algorithms and parallel kernels that map particle-based quantities onto grid-based fields (particle-to-grid or p2g mapping), optimizing interpolation, data layout, and reductions to maximize throughput and resource utilization while removing computational bottlenecks; ensures these p2g transfers preserve numerical accuracy across CPU, GPU, or other parallel hardware implementations.

particle-to-gridoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

GPZ: GPU-Accelerated Lossy Compressor for Particle Data

Aug 13, 2025
RL
Ruoyu Li
🏛️ Florida State University | University of Iowa | The University of Chicago | Argonne National Laboratory | University of California, Riverside | University of Houston | Oakland University | University of Kentucky | The Ohio State University

Addressing the challenge of simultaneously achieving high compression throughput and controllable reconstruction error for irregular, large-scale particle-based simulations and point-cloud data on modern GPUs, this paper proposes a hardware-aware, error-bounded lossy compression framework. Our method introduces an innovative four-stage parallel pipelined architecture that jointly optimizes kernel scheduling, memory access patterns, and Streaming Multiprocessor (SM) occupancy. It integrates GPU-accelerated parallel entropy coding, adaptive per-block error control, and fine-grained memory layout optimization. Evaluated on six real-world scientific datasets, our framework outperforms five state-of-the-art GPU compressors: it achieves up to 8× higher end-to-end throughput while delivering superior compression ratios and reconstruction fidelity—approaching the theoretical hardware throughput limit of contemporary GPUs.

Balancing high compression ratios with real-time analyticsCompressing massive irregular particle datasets efficientlyOvercoming GPU architectural constraints in data compression

GPU-Accelerated Algorithms for Process Mapping

Oct 14, 2025
PS
Petr Samoldekin
🏛️ Heidelberg University

This work addresses the classical task graph mapping problem onto processing units in supercomputers, aiming to balance computational load and minimize inter-task communication overhead. For the first time, GPU acceleration is introduced into this domain, yielding two parallel algorithms: (1) a hierarchical multi-partitioning framework accelerated on GPUs, and (2) a GPU-accelerated multilevel graph partitioning implementation integrating optimized coarsening and refinement strategies. Experiments demonstrate speedups of up to 598× over state-of-the-art CPU-based solvers, with a geometric mean speedup of 77.6×; Algorithm (1) incurs only ~10% increase in communication cost while maintaining competitive solution quality. The core contribution is the establishment of a novel GPU-parallel paradigm for task mapping—breaking through long-standing performance bottlenecks inherent in traditional CPU-centric approaches.

GPU-accelerated algorithms balance computational workload and minimize communication costsHierarchical multisection partitions task graphs using supercomputer hierarchyMultilevel graph partitioning pipeline accelerates coarsening and refinement phases

GPU code optimization remains a critical bottleneck in high-performance computing and large model training and inference, as existing approaches struggle to consistently approach hardware performance limits. This work proposes a two-stage GPU kernel tuner: it first transforms the original kernel into a parameterized template through semantic restructuring, then optimizes the template parameters using a performance-feedback-driven constrained search strategy. By integrating an LLM-agent-guided iterative workflow with a synergistic mechanism of templated rewriting and search-based tuning, the method significantly enhances optimization stability and interpretability while reducing manual intervention. Evaluated on real-world CUDA kernels, it achieves up to 3× speedup over baseline implementations, outperforms pure LLM-based rewriting approaches, and demonstrates strong potential for extension to other backends such as OpenCL and HIP.

code refactoringGPU kernel optimizationLLM-agent-based optimization

This work addresses the severe performance bottleneck imposed by fine-grained atomic updates in particle-in-cell simulations on conventional multicore CPUs. Targeting modern CPU architectures featuring mixed MPU-VPU SIMD units, the authors propose the first holistic co-design for current deposition, centered on matrix outer products. The approach integrates blocked matrix algorithms, a hybrid execution pipeline, an incremental sorter with O(1) amortized complexity, and a gap-compressed memory layout to optimize data access. Evaluated on laser wakefield acceleration simulations, the method achieves an 8.7× speedup in the third-order deposition kernel—reaching 83.08% of the CPU’s theoretical peak performance—and delivers a 2.63× end-to-end runtime acceleration, substantially outperforming existing GPU implementations.

atomic updatescurrent depositionmany-core CPUs

This work addresses the high cost and low efficiency associated with parallelizing and modernizing legacy scientific computing codes by proposing a structured AI agent approach. By integrating large language model agents, manual prompt engineering, and continuous integration—guided by human-provided examples, guaranteed buildability, and constrained dialogue scope—the method successfully refactored the 60,000-line single-threaded Fortran MPI code NMAP-RKPM into a C++ MPI tool with OpenMP support within two phases over several months. This effort demonstrates the feasibility and substantial effectiveness of a structured AI-assisted paradigm for large-scale high-performance computing (HPC) software modernization.

code modernizationHPClegacy scientific codes

Latest Papers

What's happening recently
View more

Data transfer between the GPU and host memory is significantly slower than computational speed, becoming a major performance bottleneck for SPH solvers. To address this, this work proposes a host-side particle memory layout optimization tailored for GPU offloading. By analyzing GPU kernel access patterns and particle attribute types, the conventional Array-of-Structures (AoS) layout is decomposed into multiple fine-grained sub-structures (Split AoS), combined with a data compression strategy to substantially reduce the overhead of data reorganization before and after transfers. Experimental results demonstrate that the proposed approach reduces data packing time by 20%–40% and decreases overall GPU offloading latency by 12%–25%, thereby significantly enhancing heterogeneous computing efficiency.

data compressionGPU-data transferhost-device communication

This study investigates whether general-purpose coding agents can undertake rigorous scientific performance engineering. Leveraging Codex and Claude Code, the research introduces an “executable scientific contract” mechanism to guide and validate hypothesis-driven optimization experiments for fixed-radius nearest neighbor search algorithms, autonomously refactoring a PyTorch implementation into a dependency-free, high-performance C++/CUDA library. The results demonstrate that goal-directed agents can effectively assume the role of experimental performance engineers. The generated standalone library precisely reproduces the original results while achieving a 1.6× speedup over the baseline GPU-accelerated PyTorch implementation through its synchronous NumPy interface, with consistent performance maintained across diverse hardware architectures.

coding agentsCUDAfixed-radius nearest-neighbor

This project addresses the energy efficiency constraints and precision challenges introduced by heterogeneous accelerators in scientific computing. With "energy consumption per trusted solution" as its core objective, it establishes a mixed-precision computing framework. This work innovatively proposes a "reckless yet responsible" computing paradigm that integrates novel number formats, floating-point emulation, hardware-software co-design, and multi-level resource management to effectively balance aggressive low-precision arithmetic with system-level detection and verification. Furthermore, the project systematically reviews the technological landscape and development trajectories of this field, distills a list of open problems, and provides comprehensive design guidelines. Ultimately, it offers both a theoretical foundation and practical reference for next-generation energy-efficient scientific computing.

energy efficiencyhardware-software co-designmixed-precision computing

This study addresses the significant efficiency gap between LLM-generated GPU kernels and expert implementations, where the specific impact of design guidance on performance remains difficult to quantify. To investigate this, we construct a diagnostic benchmark and propose a hierarchical design guidance framework that establishes an evaluation mechanism mapping algorithmic insights to low-level optimizations. This approach systematically examines the capacity of LLM agents to translate multi-granularity expert knowledge into efficient kernels. Comprehensive evaluations are conducted on NVIDIA B200 hardware using paired experiments and multidimensional metrics. The results demonstrate that incorporating structured guidance elevates code correctness to 98.5% and achieves a geometric mean speedup of 2.49×, substantially narrowing the performance disparity between LLM-generated kernels and expert-crafted implementations.

BenchmarkCode GenerationExpert Design Guidance

Hot Scholars

CJ

Chenfanfu Jiang

Professor, UCLA
Computer GraphicsComputer VisionEmbodied AIRobotics
YS

Yizhou Sun

Professor, Computer Science, UCLA
Information NetworksKnowledge GraphsGraph Neural NetworksData Mining
YC

Yunuo Chen

University of California, Los Angeles
Physics-Based Simulation
YY

Yin Yang

The University of Utah
Computer GraphicsEmbodied AI3D VisionRobotics
SW

Sinan Wang

Southern University of Science and Technology
Software EngineeringSoftware TestingSoftware Analysis