CGO 2019
33 papers
- A Code Generator for High-Performance Tensor Contractions on GPUs
- A Shared BTB Design for Multicore Systems
- A Tool for Performance Analysis of GPU-Accelerated Applications
- Accelerating GPU Computing at Runtime with Binary Optimization
- An Optimization-Driven Incremental Inline Substitution Algorithm for Just-in-Time Compilers
- Automatic Equivalence Checking for Assembly Implementations of Cryptography Libraries
- Automatic Generation of Warp-Level Primitives and Atomic Instructions for Fast and Portable Parallel Reduction on GPUs
- Automatic Parallelization of Irregular x86-64 Loops
- BOLT: A Practical Binary Optimizer for Data Centers and Beyond
- CSOD: Context-Sensitive Overflow Detection
- Code Generation from Formal Models for Automatic RTOS Portability
- Decoding CUDA Binary
- Extending LLVM for Lightweight SPMD Vectorization: Using SIMD and Vector Instructions Easily from Any Language
- From Loop Fusion to Kernel Fusion: A Domain-Specific Approach to Locality Optimization
- Function Merging by Sequence Alignment
- Generation of In-Bounds Inputs for Arrays in Memory-Unsafe Languages
- IGC: The Open Source Intel Graphics Compiler
- Janus: Statically-Driven and Profile-Guided Automatic Dynamic Binary Parallelisation
- Kernel Fusion/Decomposition for Automatic GPU-Offloading
- Locus: A System and a Language for Program Optimization
- Multi-target Compiler for the Deployment of Machine Learning Models
- Optimizing RNA-RNA Interaction Computations
- Quantifying and Reducing Execution Variance in STM via Model Driven Commit Optimization
- Reasoning about the Node.js Event Loop using Async Graphs
- Smokestack: Thwarting DOP Attacks with Runtime Stack Layout Randomization
- Super-Node SLP: Optimized Vectorization for Code Sequences Containing Operators and Their Inverse Elements
- Tensor Algebra Compilation with Workspaces
- Tiramisu: A Polyhedral Compiler for Expressing Fast and Portable Code
- Transforming Query Sequences for High-Throughput B+ Tree Processing on Many-Core Processors
- Translating CUDA to OpenCL for Hardware Generation using Neural Machine Translation
- Translating Traditional SIMD Instructions to Vector Length Agnostic Architectures
- Understanding RDMA Behavior in NUMA Systems
- White-Box Program Tuning