PPoPP 2022
46 papers
- A W-cycle algorithm for efficient batched SVD on GPUs
- A parallel branch-and-bound algorithm with history-based domination
- An LLVM-based open-source compiler for NVIDIA GPUs
- Asymmetry-aware scalable locking
- Automatic differentiation of parallel loops with formal methods
- Automatic synthesis of parallel unix commands and pipelines with KumQuat
- BaGuaLu: targeting brain scale pretrained models with over 37 million cores
- Bundling linked data structures for linearizable range queries
- CASE: a compiler-assisted SchEduling framework for multi-GPU systems
- Deadlock-free asynchronous message reordering in rust with multiparty session types
- Detectable recovery of lock-free data structures
- Dopia: online parallelism management for integrated CPU/GPU architectures
- Elimination (a, b)-trees with fast, durable updates
- Extending the limit of molecular dynamics with ab initio accuracy to 10 billion atoms
- FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models
- FliT: a library for simple and efficient persistent algorithms
- Hardening selective protection across multiple program inputs for HPC applications
- High performance GPU concurrent B+tree
- Interference relation-guided SMT solving for multi-threaded program verification
- Jiffy: a lock-free skip list with batch updates and snapshots
- LB-HM: load balance-aware data placement on heterogeneous memory for task-parallel HPC applications
- LOTUS: locality optimizing triangle counting
- Lock-free locks revisited
- Mashup: making serverless computing useful for HPC workflows via hybrid execution
- Multi-queues can be state-of-the-art priority schedulers
- Near-optimal sparse allreduce for distributed deep learning
- Optimizing consistency for partially replicated data stores
- Optimizing sparse computations jointly
- ParGeo: a library for parallel computational geometry
- Parallel algorithms for masked sparse matrix-matrix products
- Parallel block-delayed sequences
- PathCAS: an efficient middle ground for concurrent search data structures
- PerFlow: a domain specific framework for automatic performance analysis of parallel applications
- QGTC: accelerating quantized graph neural networks via GPU tensor core
- RTNN: accelerating neighbor search using hardware ray tracing
- Remote OpenMP offloading
- Rethinking graph data placement for graph neural network training on multiple GPUs
- Scaling graph traversal to 281 trillion edges with 40 million cores
- Stream processing with dependency-guided synchronization
- The performance power of software combining in persistence
- The problem-based benchmark suite (PBBS), V2
- TileSpGEMM: a tiled algorithm for parallel sparse general matrix-matrix multiplication on GPUs
- Towards OmpSs-2 and OpenACC interoperation
- Understanding and detecting deep memory persistency bugs in NVM programs with DeepMC
- Vapro: performance variance detection and diagnosis for production-run parallel applications
- wCQ: a fast wait-free queue with bounded memory usage