PPoPP 2019
60 papers
- A GPU memory efficient speed-up scheme for training ultra-deep neural networks: poster
- A coordinated tiling and batching framework for efficient GEMM on GPUs
- A distributed hypervisor for resource aggregation: poster
- A pattern based algorithmic autotuner for graph processing on GPUs
- A round-efficient distributed betweenness centrality algorithm
- A specialized B-tree for concurrent datalog evaluation
- Accelerating distributed stochastic gradient descent with adaptive periodic parameter averaging: poster
- Adaptive sparse matrix-matrix multiplication on the GPU
- Adaptive sparse tiling for sparse matrix multiplication
- Automated multi-dimensional elasticity for streaming runtimes: poster
- BASMAT: bottleneck-aware sparse matrix-vector multiplication auto-tuning on GPGPUs
- Beyond human-level accuracy: computational challenges in deep learning
- Blockchain abstract data type: poster
- Building parallel programming language constructs in the AbleC extensible C compiler framework: a PPoPP tutorial
- Checking linearizability using hitting families
- Compiler-assisted adaptive program scheduling in big.LITTLE systems: poster
- Corrected trees for reliable group communication
- Creating repeatable, reusable experimentation pipelines with popper: tutorial
- CuLDA_CGS: solving large-scale LDA problems on GPUs
- Data-flow/dependence profiling for structured transformations
- Efficient race detection with futures
- Encapsulated open nesting for STM: fine-grained higher-level conflict detection
- Engineering a high-performance GPU B-Tree
- Exploiting the input sparsity to accelerate deep neural networks: poster
- GOPipe: a granularity-oblivious programming framework for pipelined stencil executions on GPU
- GPOP: a cache and memory-efficient framework for graph processing over partitions
- GPU-based 3D cryo-EM reconstruction with key-value streams: poster
- Harmonia: a high throughput B+tree for GPUs
- High performance distributed deep learning: a beginner's guide
- High-throughput image alignment for connectomics using frugal snap judgments: poster
- Implementing parallel and concurrent tree structures
- Incremental flattening for nested data parallelism
- LOFT: lock-free transactional data structures
- Leveraging hardware TM in Haskell
- Lightweight hardware transactional memory profiling
- Lock-free channels for programming via communicating sequential processes: poster
- Making concurrent algorithms detectable: poster
- Managing application parallelism via parallel efficiency regulation: poster
- Modular transactions: bounding mixed races in space and time
- Optimizing GPU programs by register demotion: poster
- Optimizing computation-communication overlap in asynchronous task-based programs: poster
- Optimizing graph processing on GPUs using approximate computing: poster
- Performance portable C++ programming with RAJA
- Proactive work stealing for futures
- Processing transactions in a predefined order
- Profiling based out-of-core hybrid method for large neural networks: poster
- Programming quantum computers: a primer with IBM Q and D-Wave exercises
- Provably and practically efficient granularity control
- QTLS: high-performance TLS asynchronous offload framework with Intel® QuickAssist technology
- S-EnKF: co-designing for scalable ensemble Kalman filter
- SEP-graph: finding shortest execution paths for graph processing under a hybrid framework on GPU
- Scheduling HPC workloads on heterogeneous-ISA architectures: poster
- Semantics-aware scheduling policies for synchronization determinism
- Stretching the capacity of hardware transactional memory in IBM POWER architectures
- T-thinker: a task-centric distributed framework for compute-intensive divide-and-conquer algorithms
- Throughput-oriented GPU memory allocation
- Toward efficient architecture-independent algorithms for dynamic programs: poster
- Transitive joins: a sound and efficient online deadlock-avoidance policy
- VEBO: a vertex- and edge-balanced ordering heuristic to load balance parallel graph processing
- Verifying C11 programs operationally