736 papers · page 5 of 37
Wei Li, Dongjie He, Wenguang Chen, Jingling Xue
Context-sensitive pointer analysis tends to generate excessive spurious points-to relations, causing inefficiency and imprecision. We introduce stack filtering, a novel approach using Rust's stack object lifetime information to address this issue. It identifies and eliminates con…
Long Li, Jianxin Lai, Peng Yuan, Tianxiang Sui, Yan Liu, Qing Zhu, Xiaojing Zhang, Linjie Xiao + 2 more
Fully Homomorphic Encryption (FHE) facilitates computations on encrypted data without requiring access to the decryption key, offering substantial privacy benefits for deploying neural network applications in sensitive sectors such as healthcare and finance. Nonetheless, programm…
Zhanhao Liang, Hanming Sun, Wenhan Shang, Mengting Yuan, Jingqin Fu, Jiang Ma, Chun Jason Xue, Qingan Li
Recent Android systems have employed pre-compilation technology to boost app launch speed and runtime performance. However, this generates large OAT files that over-consume scarce memory and storage resources in mobile devices. This paper conducts an evaluation of code redundancy…
Fangzheng Lin, Zhongfa Wang, Hiroshi Sasaki
Speculative execution is crucial in enhancing modern processor performance but can introduce Spectre-type vulnerabilities that may leak sensitive information. Detecting Spectre gadgets from programs has been a research focus to enhance the analysis and understanding of Spectre at…
Alexandre Lopoukhine, Federico Ficarelli, Christos Vasiladiotis, Anton Lydike, Josse Van Delm, Alban Dutilleul, Luca Benini, Marian Verhelst + 1 more
High-performance micro-kernels must fully exploit today’s diverse and specialized hardware to deliver peak performance to deep neural networks (DNNs). While higher-level optimizations for DNNs are offered by numerous compilers (e.g., MLIR, TVM, OpenXLA), performance-critical micr…
Martin Paul Lücke, Oleksandr Zinenko, William S. Moses, Michel Steuwer, Albert Cohen
To take full advantage of a specific hardware target, performance engineers need to gain control on compilers in order to leverage their domain knowledge about the program and hardware. Yet, modern compilers are poorly controlled, usually by configuring a sequence of coarse-grain…
Zixuan Ma, Haojie Wang, Jingze Xing, Shuhong Huang, Liyan Zheng, Chen Zhang, Huanqi Cao, Kezhao Huang + 4 more
Tensor compilers play a critical role in optimizing deep neural networks (DNNs), with memory performance emerging as a key bottleneck in code generation for DNN models. Existing tensor compilers are constrained by inefficient auto-tuning algorithms. They either must deploy coarse…
Lazar Milikic, Milan Cugurovic, Vojin Jovanovic
Accurate static profile prediction is crucial for achieving optimal program performance in the absence of dynamic profiles. However, existing static profiling methods struggle to fully exploit the complex structure of the compiler’s intermediate representation and fail to effecti…
Sourena Naser Moghaddasi, Haris Smajlovic, Ariya Shajii, Ibrahim Numanagic
Dynamic programming (DP) is a fundamental algorithmic strategy that decomposes large problems into manageable subproblems. It is a cornerstone of many important computational methods in diverse fields, especially in the field of computational genomics, where it is used for sequen…
Haolin Pan, Yuanyu Wei, Mingjie Xing, Yanjun Wu, Chen Zhao
Determining the optimal sequence of compiler optimization passes is challenging due to the extensive and intricate search space. Traditional auto-tuning techniques, such as iterative compilation and machine learning methods, are often limited by high computational costs and diffi…
Radha Patel, Willow Ahrens, Saman P. Amarasinghe
Symmetric and sparse tensors arise naturally in many domains including linear algebra, statistics, physics, chemistry, and graph theory. Symmetric tensors are equal to their transposes, so in the n-dimensional case we can save up to a factor of n! by avoiding redundant operations…
Haoran Peng, Yu Zhang, Michael D. Ernst, Jinbao Chen, Boyao Ding
In a memory-managed programming language, programmers allocate memory by creating new objects, but programmers never free memory. A garbage collector (GC) periodically reclaims memory used by unreachable objects. As an optimization based on escape analysis, some memory can be fre…
Andrea Somaini, Filippo Carloni, Giovanni Agosta, Marco D. Santambrogio, Davide Conficconi
Pattern matching based on Regular Expressions (REs) is a pervasive and challenging computational kernel used in several applications to identify critical information in a data stream. Due to the sequential data dependency of REs and the increasing data volume growth, hardware acc…
Jubi Taneja, Avery Laird, Cong Yan, Madan Musuvathi, Shuvendu K. Lahiri
Vectorization is a powerful optimization technique that significantly boosts the performance of high performance computing applications operating on large data arrays. Despite decades of research on auto-vectorization, compilers frequently miss opportunities to vectorize code. On…
Meisam Tarabkhah, Mahshid Delavar, Mina Doosti, Amir Shaikhha
Quantum simulation plays a critical role in advancing our understanding and development of quantum algorithms, quantum computing hardware, and quantum information science. Despite the availability of various quantum circuit simulators, they often face challenges in terms of maint…
Lukas Trümper, Philipp Schaad, Berke Ates, Alexandru Calotoiu, Marcin Copik, Torsten Hoefler
The same computations are often expressed differently across software projects and programming languages. In particular, how computations involving loops are expressed varies due to the many possibilities to permute and compose loops. Since each variant may have unique performanc…
Marcel Ullrich, Sebastian Hack
Recently, AlphaDev has shown significant advances in the synthesis of branchless sorting kernels for arrays of lengths 3 to 5. In this paper, we propose an enumerative search technique based on A* search and present novel optimality-preserving heuristics and non-optimality-prese…
Huanting Wang, Patrick Lenihan, Zheng Wang
Supervised machine learning techniques have shown promising results in code analysis and optimization problems. However, a learning-based solution can be brittle because minor changes in hardware or application workloads – such as facing a new CPU architecture or code pattern – m…
Haoke Xu, Yulin Zhang, Zitong Cheng, Xiaoming Li
Convolution is the most time consuming computation kernel in Convolutional Neural Network (CNN) applications and the majority of Graph Neural Network (GNN) applications. To achieve good convolution performance, current NN libraries such as cuDNN usually transform the naive convol…
Abd-El-Aziz Zayed, Christophe Dubach
MLIR’s ability to optimize programs at multiple levels of abstraction is key to enabling domain-specific optimizing compilers. However, expressing optimizations remains tedious. Optimizations can interact in unexpected ways, making it hard to unleash full performance. Equality sa…