PPoPP 2024
45 papers
- A Holistic Approach to Automatic Mixed-Precision Code Generation and Tuning for Affine Programs
- A Row Decomposition-based Approach for Sparse Matrix Multiplication on GPUs
- AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read Mapping
- Are Your Epochs Too Epic? Batch Free Can Be Harmful
- Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication
- CPMA: An Efficient Batch-Parallel Compressed Set Without Pointers
- ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor Cores
- Exploiting Fine-Grained Redundancy in Set-Centric Graph Pattern Mining
- Extreme-scale Direct Numerical Simulation of Incompressible Turbulence on the Heterogeneous Many-core System
- Fast American Option Pricing using Nonlinear Stencils
- Fast Kronecker Matrix-Matrix Multiplication on GPUs
- FastFold: Optimizing AlphaFold Training and Inference on GPU Clusters
- Gallatin: A General-Purpose GPU Memory Manager
- GraphCube: Interconnection Hierarchy-aware Graph Processing
- INFINEL: An efficient GPU-based processing method for unpredictable large output graph queries
- Language-Agnostic Static Deadlock Detection for Futures
- Liger: Interleaving Intra- and Inter-Operator Parallelism for Distributed Large Model Inference
- Locks as a Resource: Fairly Scheduling Lock Occupation with CFL
- Memory Bounds for Concurrent Bounded Queues
- OsirisBFT: Say No to Task Replication for Scalable Byzantine Fault Tolerant Analytics
- POSTER: Accelerating High-Precision Integer Multiplication used in Cryptosystems with GPUs
- POSTER: Enabling Extreme-Scale Phase Field Simulation with In-situ Feature Extraction
- POSTER: FineCo: Fine-grained Heterogeneous Resource Management for Concurrent DNN Inferences
- POSTER: LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
- POSTER: OCToPus: Semantic-aware Concurrency Control for Blockchain Transactions
- POSTER: Optimizing Collective Communications with Error-bounded Lossy Compression for GPU Clusters
- POSTER: Optimizing Sparse Tensor Contraction with Revisiting Hash Table Design
- POSTER: ParGNN: Efficient Training for Large-Scale Graph Neural Network on GPU Clusters
- POSTER: Pattern-Aware Sparse Communication for Scalable Recommendation Model Training
- POSTER: RELAX: Durable Data Structures with Swift Recovery
- POSTER: RadiK: Scalable Radix Top-K Selection on GPUs
- POSTER: StructMG: A Fast and Scalable Structured Multigrid
- Parallel Integer Sort: Theory and Practice
- Parallel k-Core Decomposition with Batched Updates and Asynchronous Reads
- ParlayANN: Scalable and Deterministic Parallel Graph-Based Approximate Nearest Neighbor Search Algorithms
- Practical Hardware Transactional vEB Trees
- Pure: Evolving Message Passing To Better Leverage Shared Memory Within Nodes
- Recurrence Analysis for Automatic Parallelization of Subscripted Subscripts
- Scaling Up Transactions with Slower Clocks
- Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips
- Sparsity in Deep Neural Nets (Keynote)
- Tetris: Accelerating Sparse Convolution by Exploiting Memory Reuse on GPU
- Towards Scalable Unstructured Mesh Computations on Shared Memory Many-Cores
- Training one DeePMD Model in Minutes: a Step towards Online Learning
- VERLIB: Concurrent Versioned Pointers