kirancodes.me
To Proof Maintenance & Beyond!
Venues / PPoPP /

PPoPP 2019

60 papers

  1. A GPU memory efficient speed-up scheme for training ultra-deep neural networks: poster · Jinrong Guo, Wantao Liu, Wang Wang, Qu Lu, Songlin Hu, Jizhong Han + 1 more
  2. A coordinated tiling and batching framework for efficient GEMM on GPUs · Xiuhong Li, Yun Liang, Shengen Yan, Liancheng Jia, Yinghan Li
  3. A distributed hypervisor for resource aggregation: poster · Yubin Chen, Zhuocheng Ding, Jin Zhang, Yun Wang, Zhengwei Qi, Haibing Guan
  4. A pattern based algorithmic autotuner for graph processing on GPUs · Ke Meng, Jiajia Li, Guangming Tan, Ninghui Sun
  5. A round-efficient distributed betweenness centrality algorithm · Loc Hoang, Matteo Pontecorvi, Roshan Dathathri, Gurbinder Gill, Bozhi You, Keshav Pingali + 1 more
  6. A specialized B-tree for concurrent datalog evaluation · Herbert Jordan, Pavle Subotic, David Zhao, Bernhard Scholz
  7. Accelerating distributed stochastic gradient descent with adaptive periodic parameter averaging: poster · Peng Jiang, Gagan Agrawal
  8. Adaptive sparse matrix-matrix multiplication on the GPU · Martin Winter, Daniel Mlakar, Rhaleb Zayer, Hans-Peter Seidel, Markus Steinberger
  9. Adaptive sparse tiling for sparse matrix multiplication · Changwan Hong, Aravind Sukumaran-Rajam, Israt Nisa, Kunal Singh, P. Sadayappan
  10. Automated multi-dimensional elasticity for streaming runtimes: poster · Xiang Ni, Scott Schneider, Raju Pavuluri, Jonathan Kaus, Kun-Lung Wu
  11. BASMAT: bottleneck-aware sparse matrix-vector multiplication auto-tuning on GPGPUs · Athena Elafrou, Georgios I. Goumas, Nectarios Koziris
  12. Beyond human-level accuracy: computational challenges in deep learning · Joel Hestness, Newsha Ardalani, Gregory F. Diamos
  13. Blockchain abstract data type: poster · Emmanuelle Anceaume, Antonella Del Pozzo, Romaric Ludinard, Maria Potop-Butucaru, Sara Tucci Piergiovanni
  14. Building parallel programming language constructs in the AbleC extensible C compiler framework: a PPoPP tutorial · Travis Carlson, Eric Van Wyk
  15. Checking linearizability using hitting families · Burcu Kulahcioglu Ozkan, Rupak Majumdar, Filip Niksic
  16. Compiler-assisted adaptive program scheduling in big.LITTLE systems: poster · Marcelo Novaes, Vinicius Petrucci, Abdoulaye Gamatié, Fernando Magno Quintão Pereira
  17. Corrected trees for reliable group communication · Martin Küttler, Maksym Planeta, Jan Bierbaum, Carsten Weinhold, Hermann Härtig, Amnon Barak + 1 more
  18. Creating repeatable, reusable experimentation pipelines with popper: tutorial · Ivo Jimenez, Jay F. Lofstead, Carlos Maltzahn
  19. CuLDA_CGS: solving large-scale LDA problems on GPUs · Xiaolong Xie, Yun Liang, Xiuhong Li, Wei Tan
  20. Data-flow/dependence profiling for structured transformations · Fabian Gruber, Manuel Selva, Diogo Sampaio, Christophe Guillon, Antoine Moynault, Louis-Noël Pouchet + 1 more
  21. Efficient race detection with futures · Robert Utterback, Kunal Agrawal, Jeremy T. Fineman, I-Ting Angelina Lee
  22. Encapsulated open nesting for STM: fine-grained higher-level conflict detection · Martin Bättig, Thomas R. Gross
  23. Engineering a high-performance GPU B-Tree · Muhammad A. Awad, Saman Ashkiani, Rob Johnson, Martin Farach-Colton, John D. Owens
  24. Exploiting the input sparsity to accelerate deep neural networks: poster · Xiao Dong, Lei Liu, Guangli Li, Jiansong Li, Peng Zhao, Xueying Wang + 1 more
  25. GOPipe: a granularity-oblivious programming framework for pipelined stencil executions on GPU · Chanyoung Oh, Zhen Zheng, Xipeng Shen, Jidong Zhai, Youngmin Yi
  26. GPOP: a cache and memory-efficient framework for graph processing over partitions · Kartik Lakhotia, Rajgopal Kannan, Sourav Pati, Viktor K. Prasanna
  27. GPU-based 3D cryo-EM reconstruction with key-value streams: poster · Kunpeng Wang, Shizhen Xu, Hongkun Yu, Haohuan Fu, Guangwen Yang
  28. Harmonia: a high throughput B+tree for GPUs · Zhaofeng Yan, Yuzhe Lin, Lu Peng, Weihua Zhang
  29. High performance distributed deep learning: a beginner's guide · Dhabaleswar K. Panda, Ammar Ahmad Awan, Hari Subramoni
  30. High-throughput image alignment for connectomics using frugal snap judgments: poster · Tim Kaler, Brian Wheatman, Sarah Wooders
  31. Implementing parallel and concurrent tree structures · Yihan Sun, Guy E. Blelloch
  32. Incremental flattening for nested data parallelism · Troels Henriksen, Frederik Thorøe, Martin Elsman, Cosmin E. Oancea
  33. LOFT: lock-free transactional data structures · Avner Elizarov, Guy Golan-Gueta, Erez Petrank
  34. Leveraging hardware TM in Haskell · Ryan Yates, Michael L. Scott
  35. Lightweight hardware transactional memory profiling · Qingsen Wang, Pengfei Su, Milind Chabbi, Xu Liu
  36. Lock-free channels for programming via communicating sequential processes: poster · Nikita Koval, Dan Alistarh, Roman Elizarov
  37. Making concurrent algorithms detectable: poster · Naama Ben-David, Guy E. Blelloch, Michal Friedman, Yuanhao Wei
  38. Managing application parallelism via parallel efficiency regulation: poster · Sharanyan Srikanthan, Princeton Ferro, Sayak Chakraborti, Sandhya Dwarkadas
  39. Modular transactions: bounding mixed races in space and time · Brijesh Dongol, Radha Jagadeesan, James Riely
  40. Optimizing GPU programs by register demotion: poster · Putt Sakdhnagool, Amit Sabne, Rudolf Eigenmann
  41. Optimizing computation-communication overlap in asynchronous task-based programs: poster · Emilio Castillo, Nikhil Jain, Marc Casas, Miquel Moretó, Martin Schulz, Ramón Beivide + 2 more
  42. Optimizing graph processing on GPUs using approximate computing: poster · Somesh Singh, Rupesh Nasre
  43. Performance portable C++ programming with RAJA · David Beckingsale, Richard D. Hornung, Tom Scogland, Arturo Vargas
  44. Proactive work stealing for futures · Kyle Singer, Yifan Xu, I-Ting Angelina Lee
  45. Processing transactions in a predefined order · Mohamed M. Saad, Masoomeh Javidi Kishi, Shihao Jing, Sandeep Hans, Roberto Palmieri
  46. Profiling based out-of-core hybrid method for large neural networks: poster · Yuki Ito, Haruki Imai, Tung D. Le, Yasushi Negishi, Kiyokuni Kawachiya, Ryo Matsumiya + 1 more
  47. Programming quantum computers: a primer with IBM Q and D-Wave exercises · Frank Mueller, Greg Byrd, Patrick Dreher
  48. Provably and practically efficient granularity control · Umut A. Acar, Vitaly Aksenov, Arthur Charguéraud, Mike Rainey
  49. QTLS: high-performance TLS asynchronous offload framework with Intel® QuickAssist technology · Xiaokang Hu, Changzheng Wei, Jian Li, Brian Will, Ping Yu, Lu Gong + 1 more
  50. S-EnKF: co-designing for scalable ensemble Kalman filter · Junmin Xiao, Shijie Wang, Weiqiang Wan, Xuehai Hong, Guangming Tan
  51. SEP-graph: finding shortest execution paths for graph processing under a hybrid framework on GPU · Hao Wang, Liang Geng, Rubao Lee, Kaixi Hou, Yanfeng Zhang, Xiaodong Zhang
  52. Scheduling HPC workloads on heterogeneous-ISA architectures: poster · Mohamed Lamine Karaoui, Anthony Carno, Robert Lyerly, Sang-Hoon Kim, Pierre Olivier, Changwoo Min + 1 more
  53. Semantics-aware scheduling policies for synchronization determinism · Qi Zhao, Zhengyi Qiu, Guoliang Jin
  54. Stretching the capacity of hardware transactional memory in IBM POWER architectures · Ricardo Filipe, Shady Issa, Paolo Romano, João Barreto
  55. T-thinker: a task-centric distributed framework for compute-intensive divide-and-conquer algorithms · Da Yan, Guimu Guo, Md Mashiur Rahman Chowdhury, M. Tamer Özsu, John C. S. Lui, Weida Tan
  56. Throughput-oriented GPU memory allocation · Isaac Gelado, Michael Garland
  57. Toward efficient architecture-independent algorithms for dynamic programs: poster · Mohammad Mahdi Javanmard, Pramod Ganapathi, Rathish Das, Zafar Ahmad, Stephen L. Tschudi, Rezaul Chowdhury
  58. Transitive joins: a sound and efficient online deadlock-avoidance policy · Caleb Voss, Tiago Cogumbreiro, Vivek Sarkar
  59. VEBO: a vertex- and edge-balanced ordering heuristic to load balance parallel graph processing · Jiawen Sun, Hans Vandierendonck, Dimitrios S. Nikolopoulos
  60. Verifying C11 programs operationally · Simon Doherty, Brijesh Dongol, Heike Wehrheim, John Derrick