PPoPP 2006
27 papers
- "MAMA!": a memory allocator for multithreaded architectures
- A case study in top-down performance estimation for a large-scale parallel application
- Accurate and efficient runtime detection of atomicity errors in concurrent programs
- Adaptive scheduling with parallelism feedback
- Collective communication on architectures that support simultaneous communication over multiple links
- Exploiting distributed version concurrency in a transactional memory cluster
- Fast and transparent recovery for continuous availability of cluster-based servers
- Global-view abstractions for user-defined reductions and scans
- Hardware profile-guided automatic page placement for ccNUMA systems
- High-performance IPv6 forwarding algorithm for multi-core and multithreaded network processor
- Hybrid transactional memory
- McRT-STM: a high performance software transactional memory system for a multi-core runtime
- Minimizing execution time in MPI programs on an energy-constrained, power-scalable cluster
- Mobile MPI programs in computational grids
- On-line automated performance diagnosis on thousands of processes
- Optimizing irregular shared-memory applications for distributed-memory systems
- POSH: a TLS compiler that exploits program structure
- Parallel programming and code selection in fortress
- Parallel programming in modern web search engines
- Performance characterization of molecular dynamics techniques for biomolecular simulations
- Performance evaluation of adaptive MPI
- Predicting bounds on queuing delay for batch-scheduled parallel machines
- Programming for parallelism and locality with hierarchically tiled arrays
- Proving correctness of highly-concurrent linearisable objects
- RDMA read based rendezvous protocol for MPI over InfiniBand: design alternatives and benefits
- Scalable synchronous queues
- Teaching parallel computing to science faculty: best practices and common pitfalls