PPoPP 2010
49 papers
- A distributed placement service for graph-structured and tree-structured data
- A practical concurrent binary search tree
- A symbolic verifier for CUDA programs
- An adaptive performance modeling tool for GPU architectures
- An optimizing compiler for GPGPU programs with input-data sharing
- Analyzing lock contention in multithreaded applications
- Application heartbeats for software performance and health
- Applying the concurrent collections programming model to asynchronous parallel dense linear algebra
- CUDAlign: using GPU to accelerate the comparison of megabase genomic sequences
- Compiler aided selective lock assignment for improving the performance of software transactional memory
- Composable thread coloring
- Continuous speculative program parallelization in software
- Data transformations enabling loop vectorization on multithreaded data parallel architectures
- Debugging programs that use atomic blocks and transactional memory
- Does cache sharing on modern CMP matter to the performance of contemporary multithreaded programs?
- Effective communication and computation overlap with hybrid MPI/SMPSs
- Exascale computing: the challenges and opportunities in the next decade
- Extreme scale computing: challenges and opportunities
- Fast tridiagonal solvers on the GPU
- Featherweight X10: a core calculus for async-finish parallelism
- GAMBIT: effective unit testing for concurrency libraries
- Helper locks for fork-join parallel programming
- Improving parallelism and locality with asynchronous algorithms
- Input-driven dynamic execution prediction of streaming applications
- Intra-application shared cache partitioning for multithreaded applications
- Is hardware innovation over?
- Is transactional programming actually easier?
- KRASH: reproducible CPU load generation on many cores machines
- Lazy binary-splitting: a run-time adaptive work-stealing scheduler
- Leveraging parallel nesting in transactional memory
- Load balancing on speed
- Model-driven autotuning of sparse matrix-vector multiply on GPUs
- Modeling advanced collective communication algorithms on cell-based systems
- Modeling transactional memory workload performance
- NOrec: streamlining STM by abolishing ownership records
- New abstractions for effective performance analysis of STM programs
- PHANTOM: predicting performance of parallel applications on large-scale parallel machines using a single node
- SLAW: a scalable locality-aware adaptive work-stealing scheduler for multi-core systems
- Scalable communication protocols for dynamic sparse data exchange
- Scaling LAPACK panel operations using parallel cache assignment
- Scheduling support for transactional memory contention management
- Structure-driven optimizations for amorphous data-parallel programs
- Supporting lock-free composition of concurrent data objects
- Symbolic prefetching in transactional distributed shared memory
- The LOFAR correlator: implementation and performance analysis
- The pilot library for novice MPI programmers
- Thread to strand binding of parallel network applications in massive multi-threaded systems
- Towards scalable and transparent parallelization of multiplayer games using transactional memory support
- Using data structure knowledge for efficient lock generation and strong atomicity