kirancodes.me
To Proof Maintenance & Beyond!

Automating CUDA Synchronization via Program Transformation

Mingyuan Wu, Lingming Zhang, Cong Liu, Shin Hwei Tan, Yuqun Zhang

Abstract

While CUDA has been the most popular parallel computing platform and programming model for general purpose GPU computing, CUDA synchronization undergoes significant challenges for GPU programmers due to its intricate parallel computing mechanism and coding practices. In this paper, we propose AuCS, the first general framework to automate synchronization for CUDA kernel functions. AuCS transforms the original LLVM-level CUDA program control flow graph in a semantic-preserving manner for exploring the possible barrier function locations. Accordingly, AuCS develops mechanisms to correctly place barrier functions for automating synchronization in multiple erroneous (challenging-to-be-detected) synchronization scenarios, including data race, barrier divergence, and redundant barrier functions. To evaluate the effectiveness and efficiency of AuCS, we conduct an extensive set of experiments and the results demonstrate that AuCS can automate 20 out of 24 erroneous synchronization scenarios.

BibTeX
@inproceedings{Wu-al:ASE19,
  author    = {Mingyuan Wu and
               Lingming Zhang and
               Cong Liu and
               Shin Hwei Tan and
               Yuqun Zhang},
  title     = {Automating {CUDA} Synchronization via Program Transformation},
  booktitle = {ASE},
  pages     = {748--759},
  publisher = {{IEEE}},
  year      = {2019},
}

Related papers