kirancodes.me
To Proof Maintenance & Beyond!
Venues / PPoPP /

PPoPP 2010

49 papers

  1. A distributed placement service for graph-structured and tree-structured data · Gregory Buehrer, Srinivasan Parthasarathy, Shirish Tatikonda
  2. A practical concurrent binary search tree · Nathan Grasso Bronson, Jared Casper, Hassan Chafi, Kunle Olukotun
  3. A symbolic verifier for CUDA programs · Guodong Li, Ganesh Gopalakrishnan, Robert M. Kirby, Daniel J. Quinlan
  4. An adaptive performance modeling tool for GPU architectures · Sara S. Baghsorkhi, Matthieu Delahaye, Sanjay J. Patel, William D. Gropp, Wen-mei W. Hwu
  5. An optimizing compiler for GPGPU programs with input-data sharing · Yi Yang, Ping Xiang, Jingfei Kong, Huiyang Zhou
  6. Analyzing lock contention in multithreaded applications · Nathan R. Tallent, John M. Mellor-Crummey, Allan Porterfield
  7. Application heartbeats for software performance and health · Henry Hoffmann, Jonathan Eastep, Marco D. Santambrogio, Jason E. Miller, Anant Agarwal
  8. Applying the concurrent collections programming model to asynchronous parallel dense linear algebra · Aparna Chandramowlishwaran, Kathleen Knobe, Richard W. Vuduc
  9. CUDAlign: using GPU to accelerate the comparison of megabase genomic sequences · Edans Flavius de Oliveira Sandes, Alba Cristina Magalhaes Alves de Melo
  10. Compiler aided selective lock assignment for improving the performance of software transactional memory · Sandya Mannarswamy, Dhruva R. Chakrabarti, Kaushik Rajan, Sujoy Saraswati
  11. Composable thread coloring · Dean F. Sutherland, William L. Scherlis
  12. Continuous speculative program parallelization in software · Chao Zhang, Chen Ding, Xiaoming Gu, Kirk Kelsey, Tongxin Bai, Xiaobing Feng
  13. Data transformations enabling loop vectorization on multithreaded data parallel architectures · Byunghyun Jang, Perhaad Mistry, Dana Schaa, Rodrigo Dominguez, David R. Kaeli
  14. Debugging programs that use atomic blocks and transactional memory · Ferad Zyulkyarov, Tim Harris, Osman S. Unsal, Adrián Cristal, Mateo Valero
  15. Does cache sharing on modern CMP matter to the performance of contemporary multithreaded programs? · Eddy Z. Zhang, Yunlian Jiang, Xipeng Shen
  16. Effective communication and computation overlap with hybrid MPI/SMPSs · Vladimir Marjanovic, Jesús Labarta, Eduard Ayguadé, Mateo Valero
  17. Exascale computing: the challenges and opportunities in the next decade · Tilak Agerwala
  18. Extreme scale computing: challenges and opportunities · Josep Torrellas, Bill Gropp, Jaime H. Moreno, Kunle Olukotun, Vivek Sarkar
  19. Fast tridiagonal solvers on the GPU · Yao Zhang, Jonathan Cohen, John D. Owens
  20. Featherweight X10: a core calculus for async-finish parallelism · Jonathan K. Lee, Jens Palsberg
  21. GAMBIT: effective unit testing for concurrency libraries · Katherine E. Coons, Sebastian Burckhardt, Madanlal Musuvathi
  22. Helper locks for fork-join parallel programming · Kunal Agrawal, Charles E. Leiserson, Jim Sukha
  23. Improving parallelism and locality with asynchronous algorithms · Lixia Liu, Zhiyuan Li
  24. Input-driven dynamic execution prediction of streaming applications · Farhana Aleen, Monirul Sharif, Santosh Pande
  25. Intra-application shared cache partitioning for multithreaded applications · Sai Prashanth Muralidhara, Mahmut T. Kandemir, Padma Raghavan
  26. Is hardware innovation over? · Arvind
  27. Is transactional programming actually easier? · Christopher J. Rossbach, Owen S. Hofmann, Emmett Witchel
  28. KRASH: reproducible CPU load generation on many cores machines · Swann Perarnau, Guillaume Huard
  29. Lazy binary-splitting: a run-time adaptive work-stealing scheduler · Alexandros Tzannes, George C. Caragea, Rajeev Barua, Uzi Vishkin
  30. Leveraging parallel nesting in transactional memory · João Pedro Barreto, Aleksandar Dragojevic, Paulo Ferreira, Rachid Guerraoui, Michal Kapalka
  31. Load balancing on speed · Steven Hofmeyr, Costin Iancu, Filip Blagojevic
  32. Model-driven autotuning of sparse matrix-vector multiply on GPUs · JeeWhan Choi, Amik Singh, Richard W. Vuduc
  33. Modeling advanced collective communication algorithms on cell-based systems · Qasim Ali, Samuel P. Midkiff, Vijay S. Pai
  34. Modeling transactional memory workload performance · Donald E. Porter, Emmett Witchel
  35. NOrec: streamlining STM by abolishing ownership records · Luke Dalessandro, Michael F. Spear, Michael L. Scott
  36. New abstractions for effective performance analysis of STM programs · Dhruva R. Chakrabarti
  37. PHANTOM: predicting performance of parallel applications on large-scale parallel machines using a single node · Jidong Zhai, Wenguang Chen, Weimin Zheng
  38. SLAW: a scalable locality-aware adaptive work-stealing scheduler for multi-core systems · Yi Guo, Yisheng Zhao, Vincent Cavé, Vivek Sarkar
  39. Scalable communication protocols for dynamic sparse data exchange · Torsten Hoefler, Christian Siebert, Andrew Lumsdaine
  40. Scaling LAPACK panel operations using parallel cache assignment · Anthony M. Castaldo, R. Clint Whaley
  41. Scheduling support for transactional memory contention management · Walther Maldonado, Patrick Marlier, Pascal Felber, Adi Suissa, Danny Hendler, Alexandra Fedorova + 2 more
  42. Structure-driven optimizations for amorphous data-parallel programs · Mario Méndez-Lojo, Donald Nguyen, Dimitrios Prountzos, Xin Sui, Muhammad Amber Hassaan, Milind Kulkarni + 2 more
  43. Supporting lock-free composition of concurrent data objects · Daniel Cederman, Philippas Tsigas
  44. Symbolic prefetching in transactional distributed shared memory · Alokika Dash, Brian Demsky
  45. The LOFAR correlator: implementation and performance analysis · John W. Romein, P. Chris Broekema, Jan David Mol, Rob van Nieuwpoort
  46. The pilot library for novice MPI programmers · John D. Carter, William B. Gardner, Gary Gréwal
  47. Thread to strand binding of parallel network applications in massive multi-threaded systems · Petar Radojkovic, Vladimir Cakarevic, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky + 1 more
  48. Towards scalable and transparent parallelization of multiplayer games using transactional memory support · Daniel Lupei, Bogdan Simion, Don Pinto, Matthew Misler, Mihai Burcea, William Krick + 1 more
  49. Using data structure knowledge for efficient lock generation and strong atomicity · Gautam Upadhyaya, Samuel P. Midkiff, Vijay S. Pai