736 papers · page 16 of 37
David Leopoldseder, Lukas Stadler, Thomas Würthinger, Josef Eisl, Doug Simon, Hanspeter Mössenböck
Compilers perform a variety of advanced optimizations to improve the quality of the generated machine code. However, optimizations that depend on the data flow of a program are often limited by control-flow merges. Code duplication can solve this problem by hoisting, i.e. duplica…
Daniel Maier, Biagio Cosenza, Ben H. H. Juurlink
Many applications provide inherent resilience to some amount of error and can potentially trade accuracy for performance by using approximate computing. Applications running on GPUs often use local memory to minimize the number of global memory accesses and to speed up execution.…
Vasileios Porpodas, Rodrigo C. O. Rocha, Luís F. W. Góes
Auto-vectorizing compilers automatically generate vector (SIMD) instructions out of scalar code. The state-of-the-art algorithm for straight-line code vectorization is Superword-Level Parallelism (SLP). In this work we identify a major limitation at the core of the SLP algorithm,…
Andrea Rosà, Eduardo Rosales, Walter Binder
Task granularity, i.e., the amount of work performed by parallel tasks, is a key performance attribute of parallel applications. On the one hand, fine-grained tasks (i.e., small tasks carrying out few computations) may introduce considerable parallelization overheads. On the othe…
Probir Roy, Shuaiwen Leon Song, Sriram Krishnamoorthy, Xu Liu
In memory hierarchies, caches perform an important role in reducing average memory access latency. Minimizing cache misses can yield significant performance gains. As set-associative caches are widely used in modern architectures, capacity and conflict cache misses co-exist. Thes…
Du Shen, Shuaiwen Leon Song, Ang Li, Xu Liu
General-purpose GPUs have been widely utilized to accelerate parallel applications. Given a relatively complex programming model and fast architecture evolution, producing efficient GPU code is nontrivial. A variety of simulation and profiling tools have been developed to aid GPU…
Savvas Sioutas, Sander Stuijk, Henk Corporaal, Twan Basten, Lou J. Somers
Memory-bound applications heavily depend on the bandwidth of the system in order to achieve high performance. Improving temporal and/or spatial locality through loop transformations is a common way of mitigating this dependency. However, choosing the right combination of optimiza…
Marcos Yukio Siraichi, Vinícius Fernandes dos Santos, Caroline Collange, Fernando Magno Quintão Pereira
In May of 2016, IBM Research has made a quantum processor available in the cloud to the general public. The possibility of programming an actual quantum device has elicited much enthusiasm. Yet, quantum programming still lacks the compiler support that modern programming language…
Daniele G. Spampinato, Diego Fabregat-Traver, Paolo Bientinesi, Markus Püschel
We present SLinGen, a program generation system for linear algebra. The input to SLinGen is an application expressed mathematically in a linear-algebra-inspired language (LA) that we define. LA provides basic scalar/vector/matrix additions/multiplications and higher level operati…
Alen Stojanov, Ivaylo Toskov, Tiark Rompf, Markus Püschel
Managed language runtimes such as the Java Virtual Machine (JVM) provide adequate performance for a wide range of applications, but at the same time, they lack much of the low-level control that performance-minded programmers appreciate in languages like C/C++ . One important exa…
Luca Della Toffola, Michael Pradel, Thomas R. Gross
Software often suffers from performance bottlenecks, e.g., because some code has a higher computational complexity than expected or because a code change introduces a performance regression. Finding such bottlenecks is challenging for developers and for profiling techniques becau…
Biwei Xie, Jianfeng Zhan, Xu Liu, Wanling Gao, Zhen Jia, Xiwen He, Lixin Zhang
Sparse Matrix-vector Multiplication (SpMV) is an important computation kernel widely used in HPC and data centers. The irregularity of SpMV is a well-known challenge that limits SpMV’s parallelism with vectorization operations. Existing work achieves limited locality and vectoriz…
Feng Zhang, Jingling Xue
We propose POKER, a permutation-based vectorization approach for vectorizing multiple queries over B+-trees. Our key insight is to combine vector loads and path-encoding-based permutations to alleviate memory latency while keeping the number of key comparisons needed for a query …
Qing Zhou, Lian Li, Lei Wang, Jingling Xue, Xiaobing Feng
May-Happen-in-Parallel (MHP) analysis computes whether two statements in a multi-threaded program may execute concurrently or not. It works as a basis for many analyses and optimization techniques of concurrent programs. This paper proposes a novel approach for MHP analysis, by s…
Sam Ainsworth, Timothy M. Jones
Jayvant Anantpur, R. Govindarajan
Soham Chakraborty, Viktor Vafeiadis
Emilio G. Cota, Paolo Bonzini, Alex Bennée, Luca P. Carloni
Chris Cummins, Pavlos Petoumenos, Zheng Wang, Hugh Leather
Johannes Doerfert, Tobias Grosser, Sebastian Hack