PPoPP 2015
44 papers
- A collection-oriented programming model for performance portability
- A framework for practical parallel fast matrix multiplication
- A hierarchical approach to reducing communication in parallel graph algorithms
- A library for portable and composable data locality optimizations for NUMA systems
- A parallel algorithm for global states enumeration in concurrent systems
- A programming model and runtime system for significance-aware energy-efficient computing
- An OpenACC-based unified programming model for multi-accelerator systems
- Are web applications ready for parallelism?
- Automatic scalable atomicity via semantic locking
- Barrier elision for production parallel programs
- CASTLE: fast concurrent internal binary search tree using edge-based locking
- Cache-oblivious wavefront: improving parallelism of recursive dynamic programming algorithms without losing cache-efficiency
- Combining phase identification and statistic modeling for automated parallel benchmark generation
- Decoupled load balancing
- Diagnosing the causes and severity of one-sided message contention
- Distributed memory code generation for mixed Irregular/Regular computations
- Dynamic deadlock verification for general barrier synchronisation
- Efficient and reasonable object-oriented concurrency
- Fence placement for legacy data-race-free programs via synchronization read detection
- GStream: a graph streaming processing method for large-scale graphs on GPUs
- Gunrock: a high-performance graph processing library on the GPU
- High performance locks for multi-level NUMA systems
- JAWS: a JavaScript framework for adaptive CPU-GPU work sharing
- Low-overhead software transactional memory with progress guarantees and strong semantics
- MPI+Threads: runtime contention and remedies
- More than you ever wanted to know about synchronization: synchrobench, measuring the impact of the synchronization on concurrent algorithms
- NUMA-aware graph-structured analytics
- On optimizing machine learning workloads via kernel fusion
- Optimization of asynchronous graph processing on GPU with hybrid coloring model
- PLUTO+: near-complete modeling of affine transformations for parallelism and locality
- Performance implications of dynamic memory allocators on transactional memory systems
- Predicate RCU: an RCU for scalable concurrent updates
- SYNC or ASYNC: time to fuse for distributed graph-parallel computation
- Scalable and efficient implementation of 3d unstructured meshes computation: a case study on matrix assembly
- Section based program analysis to reduce overhead of detecting unsynchronized thread communication
- SemCache++: semantics-aware caching for efficient multi-GPU offloading
- Software partitioning of hardware transactions
- Static/Dynamic validation of MPI collective communications in multi-threaded context
- The SprayList: a scalable relaxed priority queue
- The lazy happens-before relation: better partial-order reduction for systematic concurrency testing
- The lock-free k-LSM relaxed priority queue
- Tiles: a new language mechanism for heterogeneous parallelism
- Towards batched linear solvers on accelerated hardware platforms
- VirtCL: a framework for OpenCL device abstraction and management