A compiler for throughput optimization of graph algorithms on GPUs
Abstract
Writing high-performance GPU implementations of graph algorithms can be challenging. In this paper, we argue that three optimizations called throughput optimizations are key to high-performance for this application class. These optimizations describe a large implementation space making it unrealistic for programmers to implement them by hand.
DOI 10.1145/2983990.2984015