PPoPP 2014
46 papers
- 21st century computer architecture
- A decomposition for in-place matrix transposition
- A general technique for non-blocking trees
- A practical wait-free simulation for lock-free data structures
- A tool to analyze the performance of multithreaded programs on NUMA architectures
- Automatic semantic locking
- Beyond parallel programming with domain specific languages
- CUDA-NP: realizing nested thread-level parallelism in GPGPU applications
- Concurrency bug localization using shared memory access pairs
- Concurrency testing using schedule bounding: an empirical study
- Data structures for task-based priority scheduling
- Designing and auto-tuning parallel 3-D FFT for computation-communication overlap
- Detecting silent data corruption through data dynamic monitoring for scientific applications
- Efficient deterministic multithreading without global barriers
- Efficient search for inputs causing high floating-point errors
- Eliminating global interpreter locks in ruby through hardware transactional memory
- Extracting logical structure and identifying stragglers in parallel execution traces
- Fast concurrent lock-free binary search trees
- Fine-grain parallel megabase sequence comparison with multiple heterogeneous GPUs
- Heterogeneous computing: what does it mean for compiler research?
- In-place transposition of rectangular matrices on accelerators
- Infrastructure-free logging and replay of concurrent execution on multiple cores
- Initial study of multi-endpoint runtime for MPI+OpenMP hybrid programming model on multi-core systems
- Leveraging hardware message passing for efficient thread synchronization
- Lock contention aware thread migrations
- Optimistic transactional boosting
- PREDATOR: predictive false sharing detection
- Parallelization hints via code skeletonization
- Parallelizing dynamic programming through rank convergence
- Portable, MPI-interoperable coarray fortran
- Practical concurrent binary search trees via logical ordering
- Provably good scheduling for parallel programs that use data structures through implicit batching
- Race directed scheduling of concurrent programs
- Resilient X10: efficient failure-aware programming
- Revisiting loop fusion in the polyhedral framework
- SCCMulti: an improved parallel strongly connected components algorithm
- Singe: leveraging warp specialization for high performance on GPUs
- Task mapping stencil computations for non-contiguous allocations
- Theoretical analysis of classic algorithms on highly-threaded many-core GPUs
- Time-warp: lightweight abort minimization in transactional memory
- Towards fair and efficient SMP virtual machine scheduling
- Trace driven dynamic deadlock detection and reproduction
- Triolet: a programming system that unifies algorithmic skeleton interfaces for high-performance cluster computing
- Well-structured futures and cache locality
- X10 and APGAS at Petascale
- yaSpMV: yet another SpMV framework on GPUs