CGO 2025
48 papers
- A Multi-level Compiler Backend for Accelerated Micro-kernels Targeting RISC-V ISA Extensions
- A Priori Loop Nest Normalization: Automatic Loop Scheduling in Complex Applications
- ANT-ACE: An FHE Compiler Framework for Automating Neural Network Inference
- ASDF: A Compiler for Qwerty, a Basis-Oriented Quantum Programming Language
- Accelerating LLMs using an Efficient GEMM Library and Target-Aware Optimizations on Real-World PIM Devices
- An Efficient Polynomial Multiplication Derived Implementation of Convolution in Neural Networks
- Automatic Synthesis of Specialized Hash Functions
- CUrator: An Efficient LLM Execution Engine with Optimized Integration of CUDA Libraries
- Cage: Hardware-Accelerated Safe WebAssembly
- Calibro: Compilation-Assisted Linking-Time Binary Code Outlining for Code Size Reduction in Android Applications
- Code Generation for Cryptographic Kernels using Multi-word Modular Arithmetic on GPU
- Combining MLIR Dialects with Domain-Specific Architecture for Efficient Regular Expression Matching
- CuAsmRL: Optimizing GPU SASS Schedules via Deep Reinforcement Learning
- DialEgg: Dialect-Agnostic MLIR Optimizer using Equality Saturation with Egglog
- Enhancing Deployment-Time Predictive Model Robustness for Code Analysis and Optimization
- FastFlip: Compositional SDC Resiliency Analysis
- GoFree: Reducing Garbage Collection via Compiler-Inserted Freeing
- GraalNN: Context-Sensitive Static Profiling with Graph Neural Networks
- Honey Potion: An eBPF Backend for Elixir
- Improving Native-Image Startup Performance
- IntelliGen: Instruction-Level Auto-tuning for Tensor Program with Monotonic Memory Optimization
- Janitizer: Rethinking Binary Tools for Practical and Comprehensive Security
- LLM-Vectorizer: LLM-Based Verified Loop Vectorizer
- MTE4JNI: A Memory Tagging Method to Protect Java Heap Memory from Illicit Native Code Access
- Memory Safety Instrumentations in Practice: Usability, Performance, and Security Guarantees
- Parallaft: Runtime-Based CPU Fault Tolerance via Heterogeneous Parallelism
- Pattern Matching in AI Compilers and Its Formalization
- Postiz: Extending Post-increment Addressing for Loop Optimization and Code Size Reduction
- PreFix: Optimizing the Performance of Heap-Intensive Applications
- Proteus: Portable Runtime Optimization of GPU Kernel Execution with Just-in-Time Compilation
- Qiwu: Exploiting Ciphertext-Level SIMD Parallelism in Homomorphic Encryption Programs
- Qubit Movement-Optimized Program Generation on Zoned Neutral Atom Processors
- Scalar Interpolation: A Better Balance between Vector and Scalar Execution for SuperScalar Architectures
- SkipFlow: Improving the Precision of Points-to Analysis using Primitive Values and Predicate Edges
- Speeding up the Local C++ Development Cycle with Header Substitution
- Stack Filtering: Elevating Precision and Efficiency in Rust Pointer Analysis
- Stardust: Compiling Sparse Tensor Algebra to a Reconfigurable Dataflow Architecture
- SySTeC: A Symmetric Sparse Tensor Compiler
- Synthesis of Quantum Simulators by Compilation
- Synthesis of Sorting Kernels
- Teapot: Efficiently Uncovering Spectre Gadgets in COTS Binaries
- Tensorize: Fast Synthesis of Tensor Programs from Legacy Code using Symbolic Tracing, Sketching and Solving
- The MLIR Transform Dialect: Your Compiler Is More Powerful Than You Think
- Towards Efficient Compiler Auto-tuning: Leveraging Synergistic Search Spaces
- VEGA: Automatically Generating Compiler Backends using a Pre-trained Transformer Model
- Vectron: A Dynamic Programming Auto-vectorization Framework
- Weaver: A Retargetable Compiler Framework for FPQA Quantum Architectures
- xDSL: Sidekick Compilation for SSA-Based Compilers