Accelerated Computing research group
Portable to Efficient: Auto-Tuning Hardware-Agnostic GPU Kernels in Julia

September 28th, 2026 — I travelled to Singapore together with Evelyne Ringoot to present our joint work on KernelTuner.jl at ICPP’26 in Sustain-HPC. The paper is a result from a collaboration with the group of Prof. Alan Edelman (MIT) that started a year before when I met Evelyne at ICPP’25. In this work, We address this challenge by integrating auto-tuning directly into hardware-agnostic GPU kernels written in Julia, using our new framework, KernelTuner.jl.

Abstract

Traditionally, GPU kernels have been developed and optimized within vendor-specific programming models to achieve high performance, resulting in software that is difficult to optimize and adapt across increasingly heterogeneous computing systems. Hardware-agnostic programming models offer a more sustainable approach to GPU software development by improving portability and maintainability, but achieving efficient execution across diverse architectures remains challenging. We address this challenge by integrating auto-tuning into hardware-agnostic GPU kernels written in Julia.

We rebuild the established Kernel Tuner auto-tuning framework with Julia support, enabling systematic exploration of kernel configurations for hardware-agnostic GPU kernels targeting NVIDIA, AMD, Intel, and Apple GPUs. We demonstrate this approach on hardware-agnostic singular value decomposition (SVD) as implemented in the NextLA.jl linear algebra library.

The results show that auto-tuning is essential for creating resource-efficient hardware-agnostic GPU kernels across a variety of hardware. Optimal configurations improve kernel performance by a factor of 3 × to 7 × compared to median parameter configurations, demonstrating the substantial impact of tuning on efficient hardware utilization. Moreover, auto-tuning enables hardware-agnostic kernels to surpass manually optimized implementations in efficiency while reducing developer effort. This is made widely accessible through the first auto-tuning framework for Julia, KernelTuner.jl, bringing portable yet resource-efficient GPU kernels within reach for a wide range of applications and devices.

Citation

F.J. Willemsen, E. Ringoot, Alen Edelman “Portable to Efficient: Auto-Tuning Hardware-Agnostic GPU Kernels in Julia” Workshop Proceedings of the 55th International Conference on Parallel Processing (ICPP 2026) https://doi.org/10.1145/3816891.3834891

Written by

Floris-Jan Willemsen

Floris-Jan is a Postdoc at Leiden University. Floris-Jan recently graduated from the Accelerated Computing research group after completing his PhD. He carried out his PhD research at the Netherlands eScience Center and Leiden University. His research focusses on intelligent, automated optimization of GPU software.