673 papers · page 1 of 34
Yashwanth Boda, Abhijit Chunduri, Ruchi Kumari, Awanish Pandey
Software testing is a critical stage in the software development lifecycle, ensuring that programs behave correctly and reliably across diverse environments. Unlike general software testing, compiler testing requires well-structured input programs that systematically exercise the…
Khushboo Chitre, Piyus Kedia, Rahul Purandare
Alias analysis is a technique to identify whether a memory location can be accessed in more than one way. An ideal alias analysis implementation should be both precise and scalable. However, in practice, implementations of alias analysis have to make a trade-off between precision…
Jonathan Van der Cruysse, Abd-El-Aziz Zayed, Mai Jacob Peng, Christophe Dubach
Equality saturation enables compilers to explore many semantically equivalent program variants, deferring optimization decisions to a final extraction phase. However, existing frameworks exhibit sequential execution and hard-coded saturation loops. This limits scalability and req…
Milan Cugurovic, Aleksandar Prokopec, Boris Spasojevic, Vojin Jovanovic, Milena Vujosevic-Janicic
Optimizing compilers often sacrifice binary size in pursuit of higher run-time performance. In the absence of method execution profiles, they uniformly apply performance-oriented optimizations, typically various forms of code duplication. Duplications in methods that are rarely o…
Bastian Hagedorn, Alexander Collins, Tony Mongkolsmai, Vinod Grover
The proliferation of Python DSLs for developing kernels has democratized GPU programming. While kernel development is now Python-native, performance analysis and optimization still rely on external tools and fragmented workflows.
We introduce Nsight Python, a Python profiling to…
Bastian Hagedorn, Vinod Grover
The rise of asynchronous execution and specialized concurrent thread groups has reshaped GPU programming, but it also introduces complex timing and coordination challenges. Developers must carefully manage data readiness, concurrency, and hardware-specific tensor units that curre…
Shideh Hashemian, Michael F. P. O'Boyle, Amir Shaikhha
Sparse tensor algebra plays an important role in many scientific and engineering applications, yet existing sparse libraries and compilers face challenges when the output tensor is sparse. Array-based storage formats, such as CSR, require costly memory reallocations and rely on i…
Ange-Thierry Ishimwe, Sam McDiarmid-Sterling, Zack McKevitt, Tamara Silbergleit Lehman
Transient execution attacks exploit speculative execution to leak confidential data through unauthorized transient memory accesses. We make the observation that transient attacks can be identified by one unusual memory access, the transient sensitive data access. To protect syste…
Saba Jamilan, Snehasish Kumar, Heiner Litz
Compilers apply optimizations such as function specialization and constant propagation to eliminate redundant work at compile time. However, because compilers must prove that values are constant, many profitable optimization opportunities remain unrealized. In this paper, we prop…
Shinnung Jeong, Chihyo Ahn, Huanzhi Pu, Jisheng Zhao, Hyesoon Kim, Blaise Tine
Recent efforts in open-source GPU research are opening new avenues in a domain that has long been tightly coupled with a few commercial vendors. Emerging open GPU architectures define SIMT functionality through their own ISAs, but executing existing GPU programs and optimizing pe…
Ramya Kasaraneni, V. Krishna Nandivada
Constraint-based points-to analysis using Andersen-style inclusion constraints is widely used for its convenience, generality, and precision in modeling complex program behaviors. Typically, such analyses generate constraints and resolve them by computing the transitive closure o…
Gaeun Ko, Seonyeong Heo
Tiny machine learning (TinyML) enables low-power microcontrollers to leverage the power of artificial intelligence without relying on remote computing resources. Typically, developing a TinyML application relies primarily on existing TinyML frameworks, which provide runtime APIs …
Johannes Lenfers, Sven Spehr, Justus Dieckmann, Johannes Jansen, Martin Paul Lücke, Sergei Gorlatch
This paper introduces Schedgehammer, a general-purpose auto-scheduling framework that optimizes program execution across diverse compiler infrastructures. Unlike existing auto-schedulers that are tightly coupled to specific intermediate representations or rely on template-based s…
Nikolaos Louloudakis, Ajitha Rajan
The ONNX Optimizer, part of the official ONNX repository and widely adopted for graph-level model optimizations, is used by default to optimize ONNX models. Despite its popularity, its ability to preserve model correctness has not been systematically evaluated. We present DiTOX, …
José Wesley de Souza Magalhães, Shideh Hashemian, Alexander Brauckmann, Jackson Woodruff, Elizabeth Polgreen, Michael F. P. O'Boyle
Linear algebra libraries and tensor domain-specific languages are able to deliver high performance for modern scientific and machine learning workloads. While there has been recent work in automatically translating legacy software to use these libraries/DSLs using pattern matchin…
A. Samuel Moses, V. Krishna Nandivada
May happen in Parallel (MHP) analysis is one of the most foundational analysis in the context of programs written in parallel languages like Java. The currently known techniques for doing MHP analysis of Java applications suffer from two main challenges: (i) scalability to real-w…
Niccolò Nicolosi, Gabriele Magnani, Emilio Corigliano, Davide Baroffio, Federico Reghenzani, Giovanni Agosta
With version 17, LLVM finalized the transition to opaque pointer types, eliminating explicit pointee‑type information from the Intermediate Representation (IR). Thus, starting from LLVM 17, each pointer type is represented in IR by the unique type ptr. Despite eliminating redunda…
Abdessamed Seddiki, Arab Mohammed, Zakaria Hebbal, Aimad Chabounia, Eduardo Chielle, Karima Benatchba, Challal Yacine, Djamel Eddine Menacer + 2 more
Fully Homomorphic Encryption (FHE) enables computations to be performed directly on encrypted data without requiring decryption, providing strong privacy guarantees. However, FHE remains computationally expensive, and writing efficient FHE programs is a complex, error-prone, and …
Alexander Brauckmann, Anderson Faustino da Silva, Gabriel Synnaeve, Michael F. P. O'Boyle, Jerónimo Castrillón, Hugh Leather
Data flow analysis is fundamental to modern program optimization and verification, serving as a critical foundation for compiler transformations. As machine learning increasingly drives compiler tasks, the need for models that can implicitly understand and correctly reason about …
Michael Canesche, Vanderson Martins do Rosario, Edson Borin, Fernando Magno Quintão Pereira
Tensor compilers like XLA, TVM, and TensorRT operate on computational graphs, where vertices represent operations and edges represent data flow between these operations. Operator fusion is an optimization that merges operators to improve their efficiency. This paper presents the …