CGO 2024
38 papers
- A Framework for Fine-Grained Synchronization of Dependent GPU Kernels
- A System-Level Dynamic Binary Translator Using Automatically-Learned Translation Rules
- A Tensor Algebra Compiler for Sparse Differentiation
- AXI4MLIR: User-Driven Automatic Host Code Generation for Custom AXI-Based Accelerators
- AskIt: Unified Programming Interface for Programming with Large Language Models
- BEC: Bit-Level Static Analysis for Reliability against Soft Errors
- Boosting the Performance of Multi-Solver IFDS Algorithms with Flow-Sensitivity Optimizations
- Compile-Time Analysis of Compiler Frameworks for Query Compilation
- Compiler Testing with Relaxed Memory Models
- DrPy: Pinpointing Inefficient Memory Usage in Multi-Layer Python Applications
- EasyTracker: A Python Library for Controlling and Inspecting Program Execution
- EasyView: Bringing Performance Profiles into Integrated Development Environments
- Ecmas: Efficient Circuit Mapping and Scheduling for Surface Code
- Enabling Fine-Grained Incremental Builds by Making Compiler Stateful
- Energy-Aware Tile Size Selection for Affine Programs on GPUs
- Enhancing Performance Through Control-Flow Unmerging and Loop Unrolling on GPUs
- Experiences Building an MLIR-Based SYCL Compiler
- High-Throughput, Formal-Methods-Assisted Fuzzing for LLVM
- Instruction Scheduling for the GPU on the GPU
- JITSPMM: Just-in-Time Instruction Generation for Accelerated Sparse Matrix-Matrix Multiplication
- Latent Idiom Recognition for a Minimalist Functional Array Language Using Equality Saturation
- One Automaton to Rule Them All: Beyond Multiple Regular Expressions Execution
- OptiWISE: Combining Sampling and Instrumentation for Granular CPI Analysis
- PolyTOPS: Reconfigurable and Flexible Polyhedral Scheduler
- PresCount: Effective Register Allocation for Bank Conflict Reduction
- Representing Data Collections in an SSA Form
- Retargeting and Respecializing GPU Workloads for Performance Portability
- Revamping Sampling-Based PGO with Context-Sensitivity and Pseudo-instrumentation
- Revealing Compiler Heuristics Through Automated Discovery and Optimization
- SCHEMATIC: Compile-Time Checkpoint Placement and Memory Allocation for Intermittent Systems
- SLaDe: A Portable Small Language Model Decompiler for Optimized Assembly
- Seer: Predictive Runtime Kernel Selection for Irregular Problems
- Tackling the Matrix Multiplication Micro-Kernel Generation with Exo
- TapeFlow: Streaming Gradient Tapes in Automatic Differentiation
- Unveiling and Vanquishing Goroutine Leaks in Enterprise Microservices: A Dynamic Analysis Approach
- Welcome from the Program Chairs
- Whose Baseline Compiler is it Anyway?
- oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation