PPoPP 2011
41 papers
- A domain-specific approach to heterogeneous parallelism
- A wait-free NCAS library for parallel applications with timing constraints
- Accelerating CUDA graph algorithms at maximum warp
- Achieving a single compute device image in OpenCL for multiple GPUs
- Active pebbles: a programming model for highly parallel fine-grained data-driven computations
- Algorithm-based recovery for HPL
- All-window profiling and composable models of cache sharing
- Auto-tuning of fast fourier transform on graphics processors
- Automatic formal verification of MPI-based parallel programs
- Automatic safety proofs for asynchronous memory operations
- COREMU: a scalable and portable parallel full-system emulator
- CSX: an extended compression format for spmv on shared memory systems
- Communicating memory transactions
- Compact data structure and scalable algorithms for the sparse grid technique
- Cooperative reasoning for preemptive execution
- Copperhead: compiling an embedded data parallel language
- Enhanced speculative parallelization via incremental recovery
- Evaluating graph coloring on GPUs
- GRace: a low-overhead mechanism for detecting data races in GPU programs
- How's the parallel computing revolution going?
- Inferring ownership transfer for efficient message passing
- Kremlin: like gprof, but for parallelization
- Lifeline-based global load balancing
- Lock-free and scalable multi-version software transactional memory
- OoOJava: software out-of-order execution
- Ordered vs. unordered: a comparison of parallelism and work-efficiency in irregular algorithms
- Programming the cloud
- Programming the memory hierarchy revisited: supporting irregular parallelism in sequoia
- QoS aware storage cache management in multi-server environments
- SCRATCH: a tool for automatic analysis of dma races
- ScalaExtrap: trace-based communication extrapolation for spmd programs
- SpiceC: scalable parallelism via implicit copying and explicit commit
- Symbolically modeling concurrent MCAPI executions
- The STAPL parallel container framework
- Thread contracts for safe parallelism
- Time skewing made simple
- Transaction communicators: enabling cooperation among concurrent transactions
- Two examples of parallel programming without concurrency constructs (PP-CC)
- ULCC: a user-level facility for optimizing shared cache performance on multicores
- Wait-free queues with multiple enqueuers and dequeuers
- Weak atomicity under the x86 memory consistency model