kirancodes.me
To Proof Maintenance & Beyond!
Venues / CGO /

CGO 2019

33 papers

  1. A Code Generator for High-Performance Tensor Contractions on GPUs · Jinsung Kim, Aravind Sukumaran-Rajam, Vineeth Thumma, Sriram Krishnamoorthy, Ajay Panyala, Louis-Noël Pouchet + 2 more
  2. A Shared BTB Design for Multicore Systems · Moumita Das, Ansuman Banerjee, Bhaskar Sardar
  3. A Tool for Performance Analysis of GPU-Accelerated Applications · Keren Zhou, John M. Mellor-Crummey
  4. Accelerating GPU Computing at Runtime with Binary Optimization · Guangli Li, Lei Liu, Xiaobing Feng
  5. An Optimization-Driven Incremental Inline Substitution Algorithm for Just-in-Time Compilers · Aleksandar Prokopec, Gilles Duboscq, David Leopoldseder, Thomas Würthinger
  6. Automatic Equivalence Checking for Assembly Implementations of Cryptography Libraries · Jay P. Lim, Santosh Nagarakatte
  7. Automatic Generation of Warp-Level Primitives and Atomic Instructions for Fast and Portable Parallel Reduction on GPUs · Simon Garcia De Gonzalo, Sitao Huang, Juan Gómez-Luna, Simon D. Hammond, Onur Mutlu, Wen-Mei Hwu
  8. Automatic Parallelization of Irregular x86-64 Loops · Brandon Neth, Michelle Mills Strout
  9. BOLT: A Practical Binary Optimizer for Data Centers and Beyond · Maksim Panchenko, Rafael Auler, Bill Nell, Guilherme Ottoni
  10. CSOD: Context-Sensitive Overflow Detection · Hongyu Liu, Sam Silvestro, Xiaoyin Wang, Lide Duan, Tongping Liu
  11. Code Generation from Formal Models for Automatic RTOS Portability · Renata Martins Gomes, Marcel Baunach
  12. Decoding CUDA Binary · Ari B. Hayes, Fei Hua, Jin Huang, Yan-Hao Chen, Eddy Z. Zhang
  13. Extending LLVM for Lightweight SPMD Vectorization: Using SIMD and Vector Instructions Easily from Any Language · Robin Kruppe, Julian Oppermann, Lukas Sommer, Andreas Koch
  14. From Loop Fusion to Kernel Fusion: A Domain-Specific Approach to Locality Optimization · Bo Qiao, Oliver Reiche, Frank Hannig, Jürgen Teich
  15. Function Merging by Sequence Alignment · Rodrigo C. O. Rocha, Pavlos Petoumenos, Zheng Wang, Murray Cole, Hugh Leather
  16. Generation of In-Bounds Inputs for Arrays in Memory-Unsafe Languages · Marcus Rodrigues, Breno Guimarães, Fernando Magno Quintão Pereira
  17. IGC: The Open Source Intel Graphics Compiler · Anupama Chandrasekhar, Gang Chen, Po-Yu Chen, Wei-Yu Chen, Junjie Gu, Peng Guo + 6 more
  18. Janus: Statically-Driven and Profile-Guided Automatic Dynamic Binary Parallelisation · Ruoyu Zhou, Timothy M. Jones
  19. Kernel Fusion/Decomposition for Automatic GPU-Offloading · Alok Mishra, Martin Kong, Barbara M. Chapman
  20. Locus: A System and a Language for Program Optimization · Thiago S. F. X. Teixeira, Corinne Ancourt, David A. Padua, William Gropp
  21. Multi-target Compiler for the Deployment of Machine Learning Models · Oscar Castro-López, Inés Fernando Vega López
  22. Optimizing RNA-RNA Interaction Computations · Swetha Varadarajan
  23. Quantifying and Reducing Execution Variance in STM via Model Driven Commit Optimization · Girish Mururu, Ada Gavrilovska, Santosh Pande
  24. Reasoning about the Node.js Event Loop using Async Graphs · Haiyang Sun, Daniele Bonetta, Filippo Schiavio, Walter Binder
  25. Smokestack: Thwarting DOP Attacks with Runtime Stack Layout Randomization · Misiker Tadesse Aga, Todd M. Austin
  26. Super-Node SLP: Optimized Vectorization for Code Sequences Containing Operators and Their Inverse Elements · Vasileios Porpodas, Rodrigo C. O. Rocha, Evgueni Brevnov, Luís F. W. Góes, Timothy G. Mattson
  27. Tensor Algebra Compilation with Workspaces · Fredrik Kjolstad, Willow Ahrens, Shoaib Kamil, Saman P. Amarasinghe
  28. Tiramisu: A Polyhedral Compiler for Expressing Fast and Portable Code · Riyadh Baghdadi, Jessica Ray, Malek Ben Romdhane, Emanuele Del Sozzo, Abdurrahman Akkas, Yunming Zhang + 3 more
  29. Transforming Query Sequences for High-Throughput B+ Tree Processing on Many-Core Processors · Ruiqin Tian, Junqiao Qiu, Zhijia Zhao, Xu Liu, Bin Ren
  30. Translating CUDA to OpenCL for Hardware Generation using Neural Machine Translation · Yonghae Kim, Hyesoon Kim
  31. Translating Traditional SIMD Instructions to Vector Length Agnostic Architectures · Sheng-Yu Fu, Wei-Chung Hsu
  32. Understanding RDMA Behavior in NUMA Systems · Jacob Nelson, Roberto Palmieri
  33. White-Box Program Tuning · Wen-Chuan Lee, Yingqi Liu, Peng Liu, Shiqing Ma, Hongjun Choi, Xiangyu Zhang + 1 more