PPoPP 2026
51 papers
- A Diagonal Block Memory-Aware Polynomial Preconditioner for Linear and Eigenvalue Solvers
- A Distributed Matrix-Block-Vector Multiplication in Presence of System Performance Variability
- APERTURE: Algorithm-System Co-optimization for Temporal Graph Network Inference
- ASM-SpMM: Unleashing the Potential of Arm SME for Sparse Matrix Multiplication Acceleration
- Accelerating Sparse Transformer Inference on GPU
- BEEMS: Boosting Machine Vision Efficiency via Computation Graph-Based Memory Smoothing
- Binary Compatible Critical Section Delegation
- CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model Training
- COCCL: A Collective Communication Library Supporting Easy Integration and Configuration of Customized Compression for Scalable LLM Training
- Cacheman: A Comprehensive Last-Level Cache Management System for Multi-tenant Clouds
- Characterizing Matrix Multiplication Units across General Parallel Patterns in Scientific Computing
- ChituDiffusion: A Data-Characteristic-Aware Serving System for Diffusion Models
- Concurrent Balanced Augmented Trees
- DTMiner: A Data-Centric System for Efficient Temporal Motif Mining
- DiggerBees: Depth First Search Leveraging Hierarchical Block-Level Stealing on GPUs
- Dynamic Detection of Inefficient Data Mapping Patterns in Heterogeneous OpenMP Applications
- ElasGNN: An Elastic Training Framework for Distributed GNN Training
- Elastor: Elastic and Efficient Model Partitioning and Checkpointing for Fault-Tolerant Distributed Training
- Exploiting Efficient Mapping and Pipelined Execution for Accelerating SpMV on Tensor Cores
- Faster and Cheaper: Pushing the Sequence Alignment Throughput with Commercial CPUs
- Fixing Non-blocking Data Structures for Better Compatibility with Memory Reclamation Schemes
- FlashAttention-T: Towards Fully Tensorized Attention by Exploiting Tensor-Vector Parallelism
- Hapax Locks: Scalable Value-Based Mutual Exclusion
- HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
- HierCut: Enabling 16-bit Format Mixed Precision for Molecular Dynamics through Hierarchical Cutoff
- High-Throughput Non-uniformly Quantized 3-bit LLM Inference
- JanusQuant: Accurate and Efficient 2-bit KV Cache Quantization for Long-Context Inference
- Laser: Unlocking Layer-Level Scheduling for Efficient Multi-SLO LLM Serving
- MetaAttention: A Unified and Performant Attention Framework across Hardware Backends
- MixFusion: A Patch-Level Parallel Serving System for Mixed-Resolution Diffusion Models
- Multiverse: Transactional Memory with Dynamic Multiversioning
- PANA: A Fine-Grained Runtime-Adaptive Load Balancing for Parallel SpMV on Multicore CPUs
- PIM-zd-tree: A Fast Space-Partitioning Index Leveraging Processing-in-Memory
- PRISM: An Efficient GPU-Based Lossy Compression Framework for Progressive Data Retrieval with Multi-Level Interpolation
- ParDiff: Efficiently Parallelizing Reverse-Mode Automatic Differentiation with Direct Indexing
- Parallel Dynamic Spatial Indexes
- Pipelonk: Accelerating End-to-End Zero-Knowledge Proof Generation on GPUs for PLONK-Based Protocols
- ROME: Maximizing GPU Efficiency for All-Pairs Shortest Path via Taming Fine-Grained Irregularities
- Rethinking Thread Scheduling under Oversubscription: A User-Space Framework for Coordinating Multi-runtime and Multi-process Workloads
- RoMeo: Mitigating Dual-dimensional Outliers with Rotated Mixed Precision Quantization
- Root-Down Exposure for Maximal Clique Enumeration on GPUs
- SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping
- Scaling GPU-to-CPU Migration for Efficient Distributed Execution on CPU Clusters
- Sharded Elimination and Combining for Highly-Efficient Concurrent Stacks
- TAC: Cache-Based System for Accelerating Billion-Scale GNN Training on Multi-GPU Platform
- Towards Singular Value Decomposition for Rank-Deficient Matrices: An Efficient and Accurate Algorithm on GPU Architectures
- Trojan Horse: Aggregate-and-Batch for Scaling Up Sparse Direct Solvers on GPU Clusters
- UFO Trees: Practical and Provably-Efficient Parallel Batch-Dynamic Trees
- VDHA: Vector-Driven Hash Aggregation for Sparse Matrix-Sparse Vector Multiplication on GPUs
- Waste-Efficient Work Stealing
- zBuffer: Zero-Copy and Metadata-Free Serialization for Fast RPC with Scatter-Gather Reflection