kirancodes.me
To Proof Maintenance & Beyond!
Venues / PPoPP /

PPoPP 2022

46 papers

  1. A W-cycle algorithm for efficient batched SVD on GPUs · Junmin Xiao, Qing Xue, Hui Ma, Xiaoyang Zhang, Guangming Tan
  2. A parallel branch-and-bound algorithm with history-based domination · Taspon Gonggiatgul, Ghassan Shobaki, Pinar Muyan-Özçelik
  3. An LLVM-based open-source compiler for NVIDIA GPUs · Da Yan, Wei Wang, Xiaowen Chu
  4. Asymmetry-aware scalable locking · Nian Liu, Jinyu Gu, Dahai Tang, Kenli Li, Binyu Zang, Haibo Chen
  5. Automatic differentiation of parallel loops with formal methods · Jan Hückelheim, Laurent Hascoët
  6. Automatic synthesis of parallel unix commands and pipelines with KumQuat · Jiasi Shen, Martin C. Rinard, Nikos Vasilakis
  7. BaGuaLu: targeting brain scale pretrained models with over 37 million cores · Zixuan Ma, Jiaao He, Jiezhong Qiu, Huanqi Cao, Yuanwei Wang, Zhenbo Sun + 19 more
  8. Bundling linked data structures for linearizable range queries · Jacob Nelson-Slivon, Ahmed Hassan, Roberto Palmieri
  9. CASE: a compiler-assisted SchEduling framework for multi-GPU systems · Chao Chen, Chris Porter, Santosh Pande
  10. Deadlock-free asynchronous message reordering in rust with multiparty session types · Zak Cutner, Nobuko Yoshida, Martin Vassor
  11. Detectable recovery of lock-free data structures · Hagit Attiya, Ohad Ben-Baruch, Panagiota Fatourou, Danny Hendler, Eleftherios Kosmas
  12. Dopia: online parallelism management for integrated CPU/GPU architectures · Younghyun Cho, Jiyeon Park, Florian Negele, Changyeon Jo, Thomas R. Gross, Bernhard Egger
  13. Elimination (a, b)-trees with fast, durable updates · Anubhav Srivastava, Trevor Brown
  14. Extending the limit of molecular dynamics with ab initio accuracy to 10 billion atoms · Zhuoqiang Guo, Denghui Lu, Yujin Yan, Siyu Hu, Rongrong Liu, Guangming Tan + 8 more
  15. FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models · Jiaao He, Jidong Zhai, Tiago Antunes, Haojie Wang, Fuwen Luo, Shangfeng Shi + 1 more
  16. FliT: a library for simple and efficient persistent algorithms · Yuanhao Wei, Naama Ben-David, Michal Friedman, Guy E. Blelloch, Erez Petrank
  17. Hardening selective protection across multiple program inputs for HPC applications · Yafan Huang, Shengjian Guo, Sheng Di, Guanpeng Li, Franck Cappello
  18. High performance GPU concurrent B+tree · Weihua Zhang, Chuanlei Zhao, Lu Peng, Yuzhe Lin, Fengzhe Zhang, Jinhu Jiang
  19. Interference relation-guided SMT solving for multi-threaded program verification · Hongyu Fan, Weiting Liu, Fei He
  20. Jiffy: a lock-free skip list with batch updates and snapshots · Tadeusz Kobus, Maciej Kokocinski, Pawel T. Wojciechowski
  21. LB-HM: load balance-aware data placement on heterogeneous memory for task-parallel HPC applications · Zhen Xie, Jie Liu, Sam Ma, Jiajia Li, Dong Li
  22. LOTUS: locality optimizing triangle counting · Mohsen Koohi Esfahani, Peter Kilpatrick, Hans Vandierendonck
  23. Lock-free locks revisited · Naama Ben-David, Guy E. Blelloch, Yuanhao Wei
  24. Mashup: making serverless computing useful for HPC workflows via hybrid execution · Rohan Basu Roy, Tirthak Patel, Vijay Gadepally, Devesh Tiwari
  25. Multi-queues can be state-of-the-art priority schedulers · Anastasiia Postnikova, Nikita Koval, Giorgi Nadiradze, Dan Alistarh
  26. Near-optimal sparse allreduce for distributed deep learning · Shigang Li, Torsten Hoefler
  27. Optimizing consistency for partially replicated data stores · Ivan Kuraj, Armando Solar-Lezama, Nadia Polikarpova
  28. Optimizing sparse computations jointly · Kazem Cheshmi, Michelle Mills Strout, Maryam Mehri Dehnavi
  29. ParGeo: a library for parallel computational geometry · Yiqiu Wang, Shangdi Yu, Laxman Dhulipala, Yan Gu, Julian Shun
  30. Parallel algorithms for masked sparse matrix-matrix products · Srdan Milakovic, Oguz Selvitopi, Israt Nisa, Zoran Budimlic, Aydin Buluç
  31. Parallel block-delayed sequences · Sam Westrick, Mike Rainey, Daniel Anderson, Guy E. Blelloch
  32. PathCAS: an efficient middle ground for concurrent search data structures · Trevor Brown, William Sigouin, Dan Alistarh
  33. PerFlow: a domain specific framework for automatic performance analysis of parallel applications · Yuyang Jin, Haojie Wang, Runxin Zhong, Chen Zhang, Jidong Zhai
  34. QGTC: accelerating quantized graph neural networks via GPU tensor core · Yuke Wang, Boyuan Feng, Yufei Ding
  35. RTNN: accelerating neighbor search using hardware ray tracing · Yuhao Zhu
  36. Remote OpenMP offloading · Atmn Patel, Johannes Doerfert
  37. Rethinking graph data placement for graph neural network training on multiple GPUs · Shihui Song, Peng Jiang
  38. Scaling graph traversal to 281 trillion edges with 40 million cores · Huanqi Cao, Yuanwei Wang, Haojie Wang, Heng Lin, Zixuan Ma, Wanwang Yin + 1 more
  39. Stream processing with dependency-guided synchronization · Konstantinos Kallas, Filip Niksic, Caleb Stanford, Rajeev Alur
  40. The performance power of software combining in persistence · Panagiota Fatourou, Nikolaos D. Kallimanis, Eleftherios Kosmas
  41. The problem-based benchmark suite (PBBS), V2 · Daniel Anderson, Guy E. Blelloch, Laxman Dhulipala, Magdalen Dobson, Yihan Sun
  42. TileSpGEMM: a tiled algorithm for parallel sparse general matrix-matrix multiplication on GPUs · Yuyao Niu, Zhengyang Lu, Haonan Ji, Shuhui Song, Zhou Jin, Weifeng Liu
  43. Towards OmpSs-2 and OpenACC interoperation · Orestis Korakitis, Simon Garcia De Gonzalo, Nicolas L. Guidotti, João Pedro Barreto, José C. Monteiro, Antonio J. Peña
  44. Understanding and detecting deep memory persistency bugs in NVM programs with DeepMC · Benjamin Reidys, Jian Huang
  45. Vapro: performance variance detection and diagnosis for production-run parallel applications · Liyan Zheng, Jidong Zhai, Xiongchao Tang, Haojie Wang, Teng Yu, Yuyang Jin + 2 more
  46. wCQ: a fast wait-free queue with bounded memory usage · Ruslan Nikolaev, Binoy Ravindran