PPoPP 2008
44 papers
- A case study in SIMD text processing with parallel bit streams: UTF-8 to UTF-16 transcoding
- A portable runtime interface for multi-level memory hierarchies
- All-window profiling of concurrent executions
- An adaptive memory conscious approach for mining frequent trees: implications for multi-core architectures
- Assertional reasoning about data races in relaxed memory models
- Automated application-level checkpointing based on live-variable analysis in MPI programs
- Automatic data movement and computation mapping for multi-level parallel architectures with explicitly managed memories
- Cache-aware iteration space partitioning
- Compiler optimizations for parallelizing general-purpose applications under thread-level speculation
- Compiler-enhanced incremental checkpointing for OpenMP applications
- Compilers and parallel computing systems
- Concurrent GC leveraging transactional memory
- Design and implementation of a high-performance MPI for C# and the common language infrastructure
- Dynamic performance tuning of word-based software transactional memory
- Enhancing the performance of MPI-IO applications by overlapping I/O, computation and communication
- Experience on optimizing irregular computation for memory hierarchy in manycore architecture
- Experiences using adaptive concurrency in transactional memory with Lee's routing algorithm
- Extracting coarse-grain parallelism in general-purpose programs
- FastForward for efficient pipeline parallelism: a cache-optimized concurrent lock-free queue
- Formal specification of the MPI-2.0 standard in TLA+
- High performance dense linear algebra on a spatially distributed processor
- ISP: a tool for model checking MPI programs
- Massive parallel LDPC decoding on GPU
- Matrix product on heterogeneous master-worker platforms
- Modeling optimistic concurrency using quantitative dependence analysis
- Nested parallelism in transactional memory
- On the correctness of transactional memory
- Optimization principles and application performance evaluation of a multithreaded GPU using CUDA
- Performance without pain = productivity: data layout and collective communication in UPC
- Practical experiences with Java software transactional memory
- Probabilistic advanced reservations for batch-scheduled parallel machines
- Programming with tiles
- Quasi-static scheduling for safe futures
- Safer open-nested transactions through ownership
- Scalable packet classification using interpreting: a cross-platform multi-core solution
- Semantics-based distributed I/O for mpiBLAST
- Software transactional memory for large scale clusters
- Split hardware transactions: true nesting of transactions using best-effort hardware transactional memory
- SuperMatrix: a multithreaded runtime scheduling system for algorithms-by-blocks
- Toward high performance nonblocking software transactional memory
- Transactional boosting: a methodology for highly-concurrent transactional objects
- Type inference for locality analysis of distributed data structures
- Where will all the threads come from?
- ZOID: I/O-forwarding infrastructure for petascale architectures