736 papers · page 20 of 37
James Pallister, Kerstin Eder, Simon J. Hollis
Deeply embedded systems often have the tightest constraints on energy consumption, requiring that they consume tiny amounts of current and run on batteries for years. However, they typically execute code directly from flash, instead of the more energy efficient RAM. We implement …
Vasileios Porpodas, Alberto Magni, Timothy M. Jones
The need to increase performance and power efficiency in modern processors has led to a wide adoption of SIMD vector units. All major vendors support vector instructions and the trend is pushing them to become wider and more powerful. However, writing code that makes efficient us…
Erven Rohou, Bharath Narasimha Swamy, André Seznec
Interpreters have been used in many contexts. They provide portability and ease of development at the expense of performance. The literature of the past decade covers analysis of why interpreters are slow, and many software techniques to improve them. A large proportion of these …
Sunil Shrestha, Guang R. Gao, Joseph B. Manzano, Andrés Márquez, John Feo
Stencil computations are at the heart of many physical simulations used in scientific codes. Thus, there exists a plethora of optimization efforts for this family of computations. Among these techniques, tiling techniques that allow concurrent start have proven to be very efficie…
Jithendra Srinivas, Wei Ding, Mahmut T. Kandemir
To fully exploit the power of emerging multicore architectures, managing shared resources (i.e., caches) across applications and over time is critical. However, to our knowledge, most prior efforts view this problem from the OS/hardware side, and do not consider whether applicati…
Evgeniy Stepanov, Konstantin Serebryany
This paper presents MemorySanitizer, a dynamic tool that detects uses of uninitialized memory in C and C++. The tool is based on compile time instrumentation and relies on bit-precise shadow memory at run-time. Shadow propagation technique is used to avoid false positive reports …
Wai Teng Tang, Ruizhe Zhao, Mian Lu, Yun Liang, Huynh Phung Huyng, Xibai Li, Rick Siow Mong Goh
Recently, the Intel Xeon Phi coprocessor has received increasing attention in high performance computing due to its simple programming model and highly parallel architecture. In this paper, we implement sparse matrix vector multiplication (SpMV) for scale-free matrices on the Xeo…
Xiaochun Zhang, Qi Guo, Yunji Chen, Tianshi Chen, Weiwu Hu
In the era of mobile and cloud computing, cross-ISA (Instruction Set Architecture) binary translation attracts increasing attentions due to the ISA diversity of computing platforms. To easily adapt to vast guest- and host-ISAs with minimal porting efforts, existing cross-ISA bina…
Long Zheng, Xiaofei Liao, Bingsheng He, Song Wu, Hai Jin
Locks have been widely used as an effective synchronization mechanism among processes and threads. However, we observe that a large number of false inter-thread dependencies (i.e., unnecessary lock contentions) exist during the program execution on multicore processors, thereby i…
Rajkishore Barik, Rashid Kaleem, Deepak Majeti, Brian T. Lewis, Tatiana Shpeisman, Chunling Hu, Yang Ni, Ali-Reza Adl-Tabatabai
Aleksandar Brankovic, Kyriakos Stavrou, Enric Gibert, Antonio González
Pablo de Oliveira Castro, Yuriy Kashnikov, Chadi Akel, Mihail Popov, William Jalby
Milind Chabbi, Xu Liu, John M. Mellor-Crummey
Emilio Coppa, Camil Demetrescu, Irene Finocchi, Romolo Marotta
Shuhan Ding, John Earnest, Soner Önder
Tobias Grosser, Albert Cohen, Justin Holewinski, P. Sadayappan, Sven Verdoolaege
Zoltán Herczeg
Sungpack Hong, Semih Salihoglu, Jennifer Widom, Kunle Olukotun
Alexandra Jimborean, Konstantinos Koukos, Vasileios Spiliopoulos, David Black-Schaffer, Stefanos Kaxiras
Juan Carlos Juega, José Ignacio Gómez, Christian Tenllado, Francky Catthoor