PPoPP 2018
50 papers
- A microbenchmark to study GPU performance models
- A persistent lock-free queue for non-volatile memory
- A predictable synchronisation algorithm
- A scalable distance-1 vertex coloring algorithm for power-law graphs
- A scalable queue for work distribution on GPUs
- An effective fusion and tile size model for optimizing image processing pipelines
- Automated code acceleration targeting heterogeneous openCL devices
- Bridging the gap between deep learning and sparse matrix format selection
- Cache-tries: concurrent lock-free hash tries with constant-time operations
- Communication-avoiding parallel minimum cuts and connected components
- Designing scalable FPGA architectures using high-level synthesis
- DisCVar: discovering critical variables using algorithmic differentiation for transient faults
- Efficient parallel determinacy race detection for two-dimensional dags
- Efficient shuffle management with SCache for DAG computing frameworks
- Featherlight on-the-fly false-sharing detection
- FlashR: parallelize and scale R for machine learning using SSDs
- Graph partitioning applied to DAG scheduling to reduce NUMA effects
- Griffin: uniting CPU and GPU in information retrieval systems for intra-query parallelism
- HPVM: heterogeneous parallel virtual machine
- Harnessing epoch-based reclamation for efficient range queries
- Hierarchical memory management for mutable state
- High-performance genomic analysis framework with in-memory computing
- Interval-based memory reclamation
- Juggler: a dependence-aware task-based execution framework for GPUs
- Layrub: layer-centric GPU memory reuse and data migration in extreme-scale deep learning systems
- Lazygraph: lazy data coherency for replicas in distributed graph-parallel computation
- Making pull-based graph processing performant
- Optimizing N-dimensional, winograd-based convolution for manycore CPUs
- PAM: parallel augmented maps
- Performance challenges in modular parallel programs
- Performance modeling for GPUs using abstract kernel emulation
- Practical concurrent traversals in search trees
- Quantifying and reducing execution variance in STM via model driven commit optimization
- Reducing the burden of parallel loop schedulers for many-core processors
- Reducing transaction aborts by looking to the future
- Register optimizations for stencils on GPUs
- Register-based implementation of the sparse general matrix-matrix multiplication on GPUs
- Revealing parallel scans and reductions in sequential loops through function reconstruction
- SIMD code generation for stencils on brick decompositions
- Safe privatization in transactional memory
- SecureMR: secure mapreduce using homomorphic encryption and program partitioning
- Shared-memory parallelization of MTTKRP for dense tensors
- Stamp-it, amortized constant-time memory reclamation in comparison to five other schemes
- Strong trylocks for reader-writer locks
- Superneurons: dynamic GPU memory management for training deep neural networks
- Transparent GPU memory management for DNNs
- Two concurrent data structures for efficient datalog query processing
- VerifiedFT: a verified, high-performance precise dynamic race detector
- swSpTRSV: a fast sparse triangular solve with sparse level tile layout on sunway architectures
- vSensor: leveraging fixed-workload snippets of programs for performance variance detection