PPoPP 2013
45 papers
- A peta-scalable CPU-GPU algorithm for global atmospheric simulations
- Adoption protocols for fanout-optimal fault-tolerant termination detection
- Array dataflow analysis for polyhedral X10 programs
- Automatic problem size sensitive task partitioning on heterogeneous parallel systems
- Betweenness centrality: algorithms and implementations
- Compiler aided manual speculation for high performance concurrent data structures
- Complexity analysis and algorithm design for reorganizing data to minimize non-coalesced memory accesses on GPU
- Correct and efficient work-stealing for weak memory models
- Data layout optimization for GPGPU architectures
- Data-only flattening for nested data parallelism
- Decomposition techniques for optimal design-space exploration of streaming applications
- Distributed merge trees
- Exploring different automata representations for efficient regular expression matching on GPUs
- Expressing graph algorithms using generalized active messages
- Fast concurrent queues for x86 processors
- FastLane: improving performance of software transactional memory for low thread counts
- From relational verification to SIMD loop synthesis
- Ligra: a lightweight graph processing framework for shared memory
- Morph algorithms on GPUs
- Multi-level parallel computing of reverse time migration for seismic imaging on blue Gene/Q
- NUMA-aware reader-writer locks
- Online-ABFT: an online algorithm based fault tolerance scheme for soft error detection in iterative methods
- Ownership passing: efficient distributed memory programming on multi-core systems
- Parallel programming with big operators
- Parallel schedule synthesis for attribute grammars
- Parallel suffix array and least common prefix for the GPU
- Programming with hardware lock elision
- RaceFree: an efficient multi-threading model for determinism
- Reducing contention through priority updates
- Relational algorithms for multi-bulk-synchronous processors
- Runtime elision of transactional barriers for captured memory
- Scalable data race detection for partitioned global address space programs
- Scalable deterministic replay in a parallel full-system emulator
- Scalable statistics counters
- Scheduling parallel programs by work stealing with private deques
- StreamScan: fast scan algorithms for GPUs without global barrier synchronization
- Swift/T: scalable data flow programming for many-task applications
- TeamWork: synchronizing threads globally to detect real deadlocks for multithreaded programs
- The tasks with effects model for safe concurrency
- TigerQuoll: parallel event-based JavaScript
- Towards an energy estimator for fault tolerance protocols
- Using hardware transactional memory to correct and simplify and readers-writer lock algorithm
- Work-stealing with configurable scheduling strategies
- WuKong: effective diagnosis of bugs at large system scales
- ZOOMM: a parallel web browser engine for multicore mobile devices