kirancodes.me
To Proof Maintenance & Beyond!

Nested data-parallelism on the gpu

Lars Bergstrom, John H. Reppy

Abstract

Graphics processing units (GPUs) provide both memory bandwidth and arithmetic performance far greater than that available on CPUs but, because of their Single-Instruction-Multiple-Data (SIMD) architecture, they are hard to program. Most of the programs ported to GPUs thus far use traditional data-level parallelism, performing only operations that operate uniformly over vectors.

Related papers