kirancodes.me
To Proof Maintenance & Beyond!

PANA: A Fine-Grained Runtime-Adaptive Load Balancing for Parallel SpMV on Multicore CPUs

Haodong Bian, Youhui Zhang, Xiang Fei, Jianqiang Huang, Xiaoying Wang

Abstract

SpMV has been widely utilized and is regarded as a significant kernel in various scientific and engineering computing applications, where its parallel performance is heavily influenced by matrix sparsity and hardware architecture. Despite extensive prior research, static partitioning strategies that narrowly target computation or memory access remain a key performance bottleneck, severely stifling performance advancement of SpMV on modern multicore CPUs.

This paper presents PANA, a fine-grained runtime-adaptive load balancing for parallel SpMV that operates at the micro-operation level. It employs fine-grained partitioning and a dynamic runtime adjustment mechanism to achieve balanced per-core computation and memory access. Experimental results demonstrate that across 2,898 SuiteSparse matrices, PANA consistently outperforms all baseline methods (CAMLB, CSR5, Merge, CVR, SpV8, MKL, and AOCL) in terms of overall performance, achieving geometric-mean speedups ranging from 1.36× to 7.72× on Intel and AMD CPUs.

Related papers