Institution profile

Linnaeus University

Academic institutioneurope · se
Official website
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

T-CCL: Resource Efficient and Performant Collective Communication using Tensor Memory Accelerator

Oct 05, 2026

This study addresses the excessive SM resource consumption of multi-GPU collective communication, which constrains computational concurrency. We propose the first resource-efficient communication scheme leveraging the Tensor Memory Accelerator (TMA). By offloading data movement and reduction operations to hardware and employing asynchronous pipelining for collective communication, our approach overcomes the thread-intensive bottlenecks of traditional libraries, significantly reducing SM overhead while maintaining high bandwidth and optimizing communication-computation overlap. Experimental results demonstrate that, compared to NCCL, the proposed method achieves up to 3.42× speedup with lower SM utilization, yields a 1.25× acceleration ratio when overlapped with GEMM operations, and improves end-to-end inference throughput by 1.31× when integrated into the vLLM backend.

0 citationsRead paper

Two-Step Occupation Coding

Jul 22, 2026

This work addresses the challenging task of mapping occupational titles from free-form text to standardized classification schemes, particularly under the adverse effects of OCR-induced noise. To tackle this problem, the authors propose a two-stage decoupled architecture: first, a domain-adapted named entity recognition (NER) model precisely extracts occupational titles, and second, these extracted titles are mapped to the target taxonomy. This separation enables each stage to focus on a single, well-defined objective, substantially improving accuracy, robustness, and interpretability. The study further introduces an innovative margin-based confidence criterion—replacing conventional absolute thresholds—to refine mapping decisions. Experiments on German-language documents demonstrate that the proposed approach significantly outperforms end-to-end single-step baselines and exhibits strong potential for cross-lingual transfer. The implementation code and evaluation scripts are publicly released to facilitate reproducibility.

0 citationsRead paper

Fine-grained Computation-Communication Overlap via Tile-level Signaling and Scheduling for Mixture-of-Experts

Jul 21, 2026

This work addresses the performance bottleneck in Mixture-of-Experts (MoE) models during multi-GPU deployment, where serial execution of expert computation and all-to-all communication exposes communication latency on the critical path, limiting GPU utilization. The authors propose a producer-consumer co-design that employs block-level scheduling: persistent compute kernels prioritize processing remote critical blocks, while dedicated streaming multiprocessor (SM) partitions run persistent communication kernels that initiate fine-grained communication based on block readiness. This enables efficient overlap between computation and return-phase communication without modifying underlying operators or communication primitives. Evaluated on a 4×A100 platform, the approach achieves up to 2.74× speedup for MoE layers and 2.64× end-to-end acceleration, demonstrating consistent performance gains and correctness across diverse GEMM shapes and routing strategies.

0 citationsRead paper

Latent-Variable Learning of SPDEs via Wiener Chaos

Feb 12, 2026

This work proposes a novel method for learning the statistical structure of linear stochastic partial differential equations (SPDEs) with additive Gaussian noise directly from spatiotemporal observational data, without requiring prior knowledge of the driving noise or initial conditions. By integrating spectral Galerkin projection with truncated Wiener chaos expansion, the SPDE is reduced to a finite-dimensional parametric system of ordinary differential equations. A structured latent variable model is introduced, enabling joint estimation of latent states and stochastic forcing terms through variational inference. This approach achieves, for the first time, end-to-end learning of the stochastic structure of SPDEs and theoretically disentangles deterministic dynamics from stochastic forcing. It attains state-of-the-art performance on synthetic data across both bounded and unbounded one-dimensional spatial domains, accurately recovering the underlying stochastic dynamical structure of the SPDEs.

0 citationsRead paper
Recent publications

Latest Papers

T-CCL: Resource Efficient and Performant Collective Communication using Tensor Memory Accelerator

Oct 05, 2026

This study addresses the excessive SM resource consumption of multi-GPU collective communication, which constrains computational concurrency. We propose the first resource-efficient communication scheme leveraging the Tensor Memory Accelerator (TMA). By offloading data movement and reduction operations to hardware and employing asynchronous pipelining for collective communication, our approach overcomes the thread-intensive bottlenecks of traditional libraries, significantly reducing SM overhead while maintaining high bandwidth and optimizing communication-computation overlap. Experimental results demonstrate that, compared to NCCL, the proposed method achieves up to 3.42× speedup with lower SM utilization, yields a 1.25× acceleration ratio when overlapped with GEMM operations, and improves end-to-end inference throughput by 1.31× when integrated into the vLLM backend.

0 citationsRead paper

Two-Step Occupation Coding

Jul 22, 2026

This work addresses the challenging task of mapping occupational titles from free-form text to standardized classification schemes, particularly under the adverse effects of OCR-induced noise. To tackle this problem, the authors propose a two-stage decoupled architecture: first, a domain-adapted named entity recognition (NER) model precisely extracts occupational titles, and second, these extracted titles are mapped to the target taxonomy. This separation enables each stage to focus on a single, well-defined objective, substantially improving accuracy, robustness, and interpretability. The study further introduces an innovative margin-based confidence criterion—replacing conventional absolute thresholds—to refine mapping decisions. Experiments on German-language documents demonstrate that the proposed approach significantly outperforms end-to-end single-step baselines and exhibits strong potential for cross-lingual transfer. The implementation code and evaluation scripts are publicly released to facilitate reproducibility.

0 citationsRead paper

Fine-grained Computation-Communication Overlap via Tile-level Signaling and Scheduling for Mixture-of-Experts

Jul 21, 2026

This work addresses the performance bottleneck in Mixture-of-Experts (MoE) models during multi-GPU deployment, where serial execution of expert computation and all-to-all communication exposes communication latency on the critical path, limiting GPU utilization. The authors propose a producer-consumer co-design that employs block-level scheduling: persistent compute kernels prioritize processing remote critical blocks, while dedicated streaming multiprocessor (SM) partitions run persistent communication kernels that initiate fine-grained communication based on block readiness. This enables efficient overlap between computation and return-phase communication without modifying underlying operators or communication primitives. Evaluated on a 4×A100 platform, the approach achieves up to 2.74× speedup for MoE layers and 2.64× end-to-end acceleration, demonstrating consistent performance gains and correctness across diverse GEMM shapes and routing strategies.

0 citationsRead paper

Latent-Variable Learning of SPDEs via Wiener Chaos

Feb 12, 2026

This work proposes a novel method for learning the statistical structure of linear stochastic partial differential equations (SPDEs) with additive Gaussian noise directly from spatiotemporal observational data, without requiring prior knowledge of the driving noise or initial conditions. By integrating spectral Galerkin projection with truncated Wiener chaos expansion, the SPDE is reduced to a finite-dimensional parametric system of ordinary differential equations. A structured latent variable model is introduced, enabling joint estimation of latent states and stochastic forcing terms through variational inference. This approach achieves, for the first time, end-to-end learning of the stochastic structure of SPDEs and theoretically disentangles deterministic dynamics from stochastic forcing. It attains state-of-the-art performance on synthetic data across both bounded and unbounded one-dimensional spatial domains, accurately recovering the underlying stochastic dynamical structure of the SPDEs.

0 citationsRead paper