PPoPP 2012
57 papers
- A GPU implementation of inclusion-based points-to analysis
- A hybrid approach of OpenMP for clusters
- A lock-free, array-based priority queue
- A methodology for creating fast wait-free data structures
- A performance analysis framework for identifying potential benefits in GPGPU applications
- A speculation-friendly binary search tree
- A work-stealing scheduler for X10's task parallelism with suspension
- Adapting the polyhedral model as a framework for efficient speculative parallelization
- Algorithm-based fault tolerance for dense matrix factorizations
- An infrastructure for dynamic optimization of parallel programs
- An overview of CMPI: network performance aware MPI in the cloud
- An overview of Medusa: simplified graph processing on GPUs
- Automatic communication optimizations through memory reuse strategies
- Automatic datatype generation and optimization
- BDDT: : block-level dynamic dependence analysis for deterministic task-based parallelism
- CPHASH: a cache-partitioned hash table
- Collective algorithms for sub-communicators
- Communication avoiding successive band reduction
- Communication-centric optimizations by dynamically detecting collective operations
- Concurrent breakpoints
- Concurrent tries with efficient non-blocking snapshots
- DOJ: dynamically parallelizing object-oriented programs
- Deterministic parallel random-number generation for dynamic-multithreading platforms
- Efficient SIMD code generation for irregular kernels
- Efficient deadlock avoidance for streaming computation with filtering
- Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors
- Establishing a Miniapp as a programmability proxy
- Extending a C-like language for portable SIMD programming
- Faster topology-aware collective algorithms through non-minimal communication
- FlexBFS: a parallelism-aware implementation of breadth-first search on GPU
- GKLEE: concolic verification and test generation for GPUs
- GPU-based NFA implementation for memory efficient high speed regular expression matching
- Internally deterministic parallel algorithms can be fast
- LHlf: lock-free linear hashing (poster paper)
- Lock cohorting: a general technique for designing NUMA locks
- Mechanizing the expert dense linear algebra developer
- NDetermin: inferring nondeterministic sequential specifications for parallelism correctness
- OpenCL as a unified programming model for heterogeneous CPU/GPU clusters
- OpenMP-style parallelism in data-centered multicore computing with R
- Optimizing remote accesses for offloaded kernels: application to high-level synthesis for FPGA
- PARRAY: a unifying array representation for heterogeneous parallelism
- Performance analysis of parallel constraint-based local search
- Portable parallel performance from sequential, productive, embedded domain-specific languages
- Programming parallel embedded and consumer applications in OpenMP superscalar
- RACECAR: a heuristic for automatic function specialization on multi-core heterogeneous systems
- Revisiting the combining synchronization technique
- S: a scripting language for high-performance RESTful web services
- Scalable GPU graph traversal
- Scalable framework for mapping streaming applications onto multi-GPU systems
- Scalable parallel debugging with statistical assertions
- Scalable parallel minimum spanning forest computation
- Speculative parallelization on GPGPUs
- Synchronization views for event-loop actors
- The boat hull model: adapting the roofline model to enable performance prediction for parallel computing
- Using GPU's to accelerate stencil-based computation kernels for the development of large scale scientific applications on heterogeneous systems
- Verification of software barriers
- Wait-free linked-lists