kirancodes.me
To Proof Maintenance & Beyond!

Reducing the burden of parallel loop schedulers for many-core processors

Mahwish Arif, Hans Vandierendonck

Abstract

This work proposes a low-overhead half-barrier pattern to schedule fine-grain parallel loops and considers its integration in the Intel OpenMP and Cilkplus schedulers. Experimental evaluation demonstrates that the scheduling overhead of our techniques is 43% lower than Intel OpenMP and 12.1x lower than Cilk. We observe 22% speedup on 48 threads, with a peak of 2.8x speedup.

Related papers