PPoPP 2023
43 papers
- 2PLSF: Two-Phase Locking with Starvation-Freedom
- A Programming Model for GPU Load Balancing
- A Scalable Hybrid Total FETI Method for Massively Parallel FEM Simulations
- AArch64 Atomics: Might They Be Harming Your Performance?
- Block-STM: Scaling Blockchain Execution by Turning Ordering Curse to a Performance Blessing
- Boosting Performance and QoS for Concurrent GPU B+trees by Combining-Based Synchronization
- CuPBoP: A Framework to Make CUDA Portable
- DSP: Efficient GNN Training with Multiple GPUs
- Dynamic N: M Fine-Grained Structured Sparse Attention Mechanism
- Efficient All-Reduce for Distributed DNN Training in Optical Interconnect Systems
- Efficient Direct Convolution Using Long SIMD Instructions
- Elastic Averaging for Efficient Pipelined DNN Training
- End-to-End LU Factorization of Large Matrices on GPUs
- Exploring the Use of WebAssembly in HPC
- Fast Parallel Exact Inference on Bayesian Networks
- Fast Symmetric Eigenvalue Decomposition via WY Representation on Tensor Core
- Fast and Scalable Channels in Kotlin Coroutines
- Generating Fast FFT Kernels on CPUs via FFT-Specific Intrinsics
- High-Performance Filters for GPUs
- High-Performance GPU-to-CPU Transpilation and Optimization via High-Level Parallel Constructs
- High-Performance and Scalable Agent-Based Simulation with BioDynaMo
- High-Throughput GPU Random Walk with Fine-Tuned Concurrent Query Processing
- Improving Energy Saving of One-Sided Matrix Decompositions on CPU-GPU Heterogeneous Systems
- Learning to Parallelize in a Shared-Memory Environment with Transformers
- Lifetime-Based Optimization for Simulating Quantum Circuits on a New Sunway Supercomputer
- Merchandiser: Data Placement on Heterogeneous Memory for Task-Parallel HPC Applications with Load-Balance Awareness
- OpenCilk: A Modular and Extensible Software Infrastructure for Fast Task-Parallel Code
- PiPAD: Pipelined and Parallel Dynamic GNN Training on GPUs
- Practically and Theoretically Efficient Garbage Collection for Multiversioning
- Provably Fast and Space-Efficient Parallel Biconnectivity
- Provably Good Randomized Strategies for Data Placement in Distributed Key-Value Stores
- Stream-K: Work-Centric Parallel Decomposition for Dense Matrix-Matrix Multiplication on the GPU
- Swift: Expedited Failure Recovery for Large-Scale DNN Training
- TDC: Towards Extremely Efficient CNNs on GPUs via Hardware-Aware Tucker Decomposition
- TGOpt: Redundancy-Aware Optimizations for Temporal Graph Attention Networks
- TL4x: Buffered Durable Transactions on Disk as Fast as in Memory
- The ERA Theorem for Safe Memory Reclamation
- The State-of-the-Art LCRQ Concurrent Queue Algorithm Does NOT Require CAS2
- Transactional Composition of Nonblocking Data Structures
- Unexpected Scaling in Path Copying Trees
- Visibility Algorithms for Dynamic Dependence Analysis and Distributed Coherence
- WISE: Predicting the Performance of Sparse Matrix Vector Multiplication with Machine Learning
- iQAN: Fast and Accurate Vector Search with Efficient Intra-Query Parallelism on Multi-Core Architectures