March 2026

Conference Paper

A Study of Performance Portability of Low-bit Fused Matrix-Vector Multiplication Kernels in SYCL

By:
Jin, Zheming
Page Number:
2129-2136
Book Title:
SC Workshops '25: Proceedings of the SC '25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis
Publication Date:
March 2026
Publisher Location:
Association for Computing Machinery, New York, New York, United States of America
Conference Name:
SC Workshops '25: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis
Conference Location:
St Louis, Missouri, United States of America
Conference Sponsor:
ACM and IEEE
View DOI Listing:
https://doi.org/10.1145/3731599.376757

Abstract

Understanding the causes of performance gaps between a portable programming model and a vendor-specific programming model is important for improving performance portability. This paper studies performance portability of low-bit fused general matrix-vector multiplication kernels in SYCL on vendors’ graphics processing units (GPUs). This work introduces the use case, explains the kernel implementations in detail, evaluates the performance of the CUDA, HIP, and SYCL kernels on datacenter, desktop, and laptop GPUs, and investigates the causes of performance gaps. The results show that loop unrolling, kernel dispatch overhead, and sum reduction contribute to the gaps.


Related Researchers