PPoPP 2016
55 papers
- A high-performance parallel algorithm for nonnegative matrix factorization
- A programming system for future proofing performance critical libraries
- A scalable lock-free hash table with open addressing
- A wait-free queue as fast as fetch-and-add
- AUTOGEN: automatic discovery of cache-oblivious parallel recursive algorithms for solving dynamic programs
- Adding approximate counters
- Affinity-aware work-stealing for integrated CPU-GPU processors
- An interval constrained memory allocator for the Givy GAS runtime
- Articulation points guided redundancy elimination for betweenness centrality
- Be my guest: MCS lock now welcomes guests
- Benchmarking weak memory models
- CUDA acceleration for Xen virtual machines in infiniband clusters with rCUDA
- Causal consistency: beyond memory
- Coarse grain parallelization of deep neural networks
- Concurrent hash tables: fast and general?(!)
- Contention-conscious, locality-preserving locks
- DSMR: a shared and distributed memory algorithm for single-source shortest path problem
- Data-centric combinatorial optimization of parallel code
- Declarative coordination of graph-based parallel programs
- Distributed Halide
- DomLock: a new multi-granularity locking technique for hierarchies
- Drinking from both glasses: combining pessimistic and optimistic tracking of cross-thread dependences
- ESTIMA: extrapolating scalability of in-memory applications
- Effect of portable fine-grained locality on energy efficiency and performance in concurrent search trees
- Efficient distributed workstealing via matchmaking
- Exploiting accelerators for efficient high dimensional similarity search
- GPU multisplit
- Generic messages: capability-based shared memory parallelism for event-loop systems
- Grain graphs: OpenMP performance analysis made easy
- Gunrock: a high-performance graph processing library on the GPU
- High performance model based image reconstruction
- Hybrid CPU-GPU scheduling and execution of tree traversals
- Improving efficacy of internal binary search trees using local recovery
- Keep calm and react with foresight: strategies for low-latency and energy-efficient elastic data stream processing
- Lease/release: architectural support for scaling contended data structures
- Merge-based sparse matrix-vector multiplication (SpMV) using the CSR storage format
- Multi-core on-the-fly SCC decomposition
- NUMA-aware scheduling and memory allocation for data-flow task-parallel applications
- OPR: deterministic group replay for one-sided communication
- On designing NUMA-aware concurrency control for scalable transactional memory
- On ordering transaction commit
- Optimistic concurrency with OPTIK
- Parallel type-checking with haskell using saturating LVars and stream generators
- Preemption-aware planning on big-data systems
- Production-guided concurrency debugging
- Refined transactional lock elision
- SPIRIT: a runtime system for distributed irregular tree applications
- Samsara parallel: a non-BSP parallel-in-time model
- Scalable adaptive NUMA-aware lock: combining local locking and remote locking for efficient concurrency
- The virtues of conflict: analysing modern concurrency
- Tidex: a mutual exclusion lock
- Unifying fixed code and fixed data mapping of load-imbalanced pipelined loops
- User-assisted storage reuse determination for dynamic task graphs
- Verification of MPI Java programs using software model checking
- Work stealing for interactive services to meet target latency