kirancodes.me
To Proof Maintenance & Beyond!
Venues / PPoPP /

PPoPP 2012

57 papers

  1. A GPU implementation of inclusion-based points-to analysis · Mario Méndez-Lojo, Martin Burtscher, Keshav Pingali
  2. A hybrid approach of OpenMP for clusters · Okwan Kwon, Fahed Jubair, Rudolf Eigenmann, Samuel P. Midkiff
  3. A lock-free, array-based priority queue · Yujie Liu, Michael F. Spear
  4. A methodology for creating fast wait-free data structures · Alex Kogan, Erez Petrank
  5. A performance analysis framework for identifying potential benefits in GPGPU applications · Jaewoong Sim, Aniruddha Dasgupta, Hyesoon Kim, Richard W. Vuduc
  6. A speculation-friendly binary search tree · Tyler Crain, Vincent Gramoli, Michel Raynal
  7. A work-stealing scheduler for X10's task parallelism with suspension · Olivier Tardieu, Haichuan Wang, Haibo Lin
  8. Adapting the polyhedral model as a framework for efficient speculative parallelization · Alexandra Jimborean, Philippe Clauss, Benoît Pradelle, Luis Mastrangelo, Vincent Loechner
  9. Algorithm-based fault tolerance for dense matrix factorizations · Peng Du, Aurélien Bouteiller, George Bosilca, Thomas Hérault, Jack J. Dongarra
  10. An infrastructure for dynamic optimization of parallel programs · Albert Noll, Thomas R. Gross
  11. An overview of CMPI: network performance aware MPI in the cloud · Yifan Gong, Bingsheng He, Jianlong Zhong
  12. An overview of Medusa: simplified graph processing on GPUs · Jianlong Zhong, Bingsheng He
  13. Automatic communication optimizations through memory reuse strategies · Muthu Manikandan Baskaran, Nicolas Vasilache, Benoît Meister, Richard Lethin
  14. Automatic datatype generation and optimization · Fredrik Kjolstad, Torsten Hoefler, Marc Snir
  15. BDDT: : block-level dynamic dependence analysis for deterministic task-based parallelism · George Tzenakis, Angelos Papatriantafyllou, John Kesapides, Polyvios Pratikakis, Hans Vandierendonck, Dimitrios S. Nikolopoulos
  16. CPHASH: a cache-partitioned hash table · Zviad Metreveli, Nickolai Zeldovich, M. Frans Kaashoek
  17. Collective algorithms for sub-communicators · Anshul Mittal, Nikhil Jain, Thomas George, Yogish Sabharwal, Sameer Kumar
  18. Communication avoiding successive band reduction · Grey Ballard, James Demmel, Nicholas Knight
  19. Communication-centric optimizations by dynamically detecting collective operations · Torsten Hoefler, Timo Schneider
  20. Concurrent breakpoints · Chang-Seo Park, Koushik Sen
  21. Concurrent tries with efficient non-blocking snapshots · Aleksandar Prokopec, Nathan Grasso Bronson, Phil Bagwell, Martin Odersky
  22. DOJ: dynamically parallelizing object-oriented programs · Yong Hun Eom, Stephen Yang, James Christopher Jenista, Brian Demsky
  23. Deterministic parallel random-number generation for dynamic-multithreading platforms · Charles E. Leiserson, Tao B. Schardl, Jim Sukha
  24. Efficient SIMD code generation for irregular kernels · Seonggun Kim, Hwansoo Han
  25. Efficient deadlock avoidance for streaming computation with filtering · Jeremy D. Buhler, Kunal Agrawal, Peng Li, Roger D. Chamberlain
  26. Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors · Sara S. Baghsorkhi, Isaac Gelado, Matthieu Delahaye, Wen-mei W. Hwu
  27. Establishing a Miniapp as a programmability proxy · Andrew Stone, John M. Dennis, Michelle Strout
  28. Extending a C-like language for portable SIMD programming · Roland Leißa, Sebastian Hack, Ingo Wald
  29. Faster topology-aware collective algorithms through non-minimal communication · Paul Sack, William Gropp
  30. FlexBFS: a parallelism-aware implementation of breadth-first search on GPU · Gu Liu, Hong An, Wenting Han, Xiaoqiang Li, Tao Sun, Wei Zhou + 2 more
  31. GKLEE: concolic verification and test generation for GPUs · Guodong Li, Peng Li, Geoffrey Sawaya, Ganesh Gopalakrishnan, Indradeep Ghosh, Sreeranga P. Rajan
  32. GPU-based NFA implementation for memory efficient high speed regular expression matching · Yuan Zu, Ming Yang, Zhonghu Xu, Lin Wang, Xin Tian, Kunyang Peng + 1 more
  33. Internally deterministic parallel algorithms can be fast · Guy E. Blelloch, Jeremy T. Fineman, Phillip B. Gibbons, Julian Shun
  34. LHlf: lock-free linear hashing (poster paper) · Donghui Zhang, Per-Åke Larson
  35. Lock cohorting: a general technique for designing NUMA locks · David Dice, Virendra J. Marathe, Nir Shavit
  36. Mechanizing the expert dense linear algebra developer · Bryan Marker, Andy Terrel, Jack Poulson, Don S. Batory, Robert A. van de Geijn
  37. NDetermin: inferring nondeterministic sequential specifications for parallelism correctness · Jacob Burnim, Tayfun Elmas, George C. Necula, Koushik Sen
  38. OpenCL as a unified programming model for heterogeneous CPU/GPU clusters · Jungwon Kim, Sangmin Seo, Jun Lee, Jeongho Nah, Gangwon Jo, Jaejin Lee
  39. OpenMP-style parallelism in data-centered multicore computing with R · Lei Jiang, Pragneshkumar B. Patel, George Ostrouchov, Ferdinand Jamitzky
  40. Optimizing remote accesses for offloaded kernels: application to high-level synthesis for FPGA · Christophe Alias, Alain Darte, Alexandru Plesco
  41. PARRAY: a unifying array representation for heterogeneous parallelism · Yifeng Chen, Xiang Cui, Hong Mei
  42. Performance analysis of parallel constraint-based local search · Yves Caniou, Daniel Diaz, Florian Richoux, Philippe Codognet, Salvador Abreu
  43. Portable parallel performance from sequential, productive, embedded domain-specific languages · Shoaib Kamil, Derrick Coetzee, Scott Beamer, Henry Cook, Ekaterina Gonina, Jonathan Harper + 2 more
  44. Programming parallel embedded and consumer applications in OpenMP superscalar · Michael Andersch, Chi Ching Chi, Ben H. H. Juurlink
  45. RACECAR: a heuristic for automatic function specialization on multi-core heterogeneous systems · John Robert Wernsing, Greg Stitt
  46. Revisiting the combining synchronization technique · Panagiota Fatourou, Nikolaos D. Kallimanis
  47. S: a scripting language for high-performance RESTful web services · Daniele Bonetta, Achille Peternier, Cesare Pautasso, Walter Binder
  48. Scalable GPU graph traversal · Duane Merrill, Michael Garland, Andrew S. Grimshaw
  49. Scalable framework for mapping streaming applications onto multi-GPU systems · Huynh Phung Huynh, Andrei Hagiescu, Weng-Fai Wong, Rick Siow Mong Goh
  50. Scalable parallel debugging with statistical assertions · Minh Ngoc Dinh, David Abramson, Chao Jin, Andrew Gontarek, Bob Moench, Luiz De Rose
  51. Scalable parallel minimum spanning forest computation · Sadegh Nobari, Thanh-Tung Cao, Panagiotis Karras, Stéphane Bressan
  52. Speculative parallelization on GPGPUs · Min Feng, Rajiv Gupta, Laxmi N. Bhuyan
  53. Synchronization views for event-loop actors · Joeri De Koster, Stefan Marr, Theo D'Hondt
  54. The boat hull model: adapting the roofline model to enable performance prediction for parallel computing · Cedric Nugteren, Henk Corporaal
  55. Using GPU's to accelerate stencil-based computation kernels for the development of large scale scientific applications on heterogeneous systems · Jian Tao, Marek Blazewicz, Steven R. Brandt
  56. Verification of software barriers · Alexander Malkis, Anindya Banerjee
  57. Wait-free linked-lists · Shahar Timnat, Anastasia Braginsky, Alex Kogan, Erez Petrank