PPoPP 2020
46 papers
- <u>G</u>PU <u>i</u>nitiated <u>O</u>penSHMEM: correct and efficient intra-kernel networking for dGPUs
- A novel data transformation and execution strategy for accelerating sparse matrix multiplication on GPUs
- A parallel sparse tensor benchmark suite on CPUs and GPUs
- A supernodal all-pairs shortest path algorithm
- A tool for top-down performance analysis of GPU-accelerated applications
- A wait-free universal construction for large objects
- ArcherGear: data race equivalencing for expeditious HPC debugging
- Breaking master-slave model between host and FPGAs
- Detecting and reproducing error-code propagation bugs in MPI implementations
- ELDA: LDA made efficient via algorithm-system codesign submission
- Fast concurrent data sketches
- Functional faults
- Identifying scalability bottlenecks for large-scale parallel programs with graph analysis
- Increasing the parallelism of graph coloring via shortcutting
- Kite: efficient and available release consistency for the datacenter
- MatRox: modular approach for improving data locality in hierarchical (Mat)rix App(Rox)imation
- Neighbor-list-free molecular dynamics on sunway TaihuLight supercomputer
- Nesting and composition in transactional data structure libraries
- No barrier in the road: a comprehensive study and optimization of ARM barriers
- Non-blocking interpolation search trees with doubly-logarithmic running time
- Nonblocking persistent software transactional memory
- Oak: a scalable off-heap allocated key-value map
- On the fly MHP analysis
- Optimizing GPU programs by partial evaluation
- Optimizing batched Winograd convolution on GPUs
- Overlapping host-to-device copy and computation using hidden unified memory
- PLUM: static parallel program locality analysis under uniform multiplexing
- Parallel and distributed bounded model checking of multi-threaded programs
- Parallel determinacy race detection for futures
- Practical parallel hypergraph algorithms
- Reflector: a fine-grained I/O tracker for HPC systems
- Restricted memory-friendly lock-free bounded queues
- Revisiting linpack algorithm on large-scale CPU-GPU heterogeneous systems
- Scalable top-k retrieval with Sparta
- Scaling concurrent queues by using HTM to profit from failed atomic operations
- Scaling out speculative execution of finite-state machines with parallel merge
- Taming unbalanced training workloads in deep learning with partial collective operations
- Testing concurrency on the JVM with lincheck
- Understand the overheads of storage data structures on persistent memory
- Understanding and optimizing persistent memory allocation
- Universal wait-free memory reclamation
- Using sample-based time series data for automated diagnosis of scalability losses in parallel programs
- XIndex: a scalable learned index for multicore data storage
- YewPar: skeletons for exact combinatorial search
- spECK: accelerating GPU sparse matrix-matrix multiplication through lightweight analysis
- waveSZ: a hardware-algorithm co-design of efficient lossy compression for scientific data