1,184 papers · page 52 of 60
Amit Karwande, Xin Yuan, David K. Lowenthal
Compiled communication has recently been proposed to improve communication performance for clusters of workstations. The idea of compiled communication is to apply more aggressive optimizations to communications whose information is known at compile time. Existing MPI libraries d…
Baris M. Kazar
No abstract available.
Hyong-youb Kim, Vijay S. Pai, Scott Rixner
Programmable network interfaces provide the potential to extend the functionality of network services but lead to instruction processing overheads when compared to application-specific network interfaces. This paper aims to offset those performance disadvantages by exploiting tas…
Ting Liu, Margaret Martonosi
Sensor networks are long-running computer systems with many sensing/compute nodes working to gather information about their environment, process and fuse that information, and in some cases, actuate control mechanisms in response. Like traditional parallel systems, communication …
Collin McCurdy, Charles N. Fischer
In programming high performance applications, shared address-space platforms are preferable for fine-grained computation, while distributed address-space platforms are more suitable for coarse-grained computation. However, currently only distributed address-space systems scale be…
Luke K. McDowell, Susan J. Eggers, Steven D. Gribble
Simultaneous multithreading (SMT) represents a fundamental shift in processor capability. SMT's ability to execute multiple threads simultaneously within a single CPU offers tremendous potential performance benefits. However, the structure and behavior of software affects the ext…
Piotr Nienaltowski
No abstract available.
Robert O'Callahan, Jong-Deok Choi
We present a new method for dynamically detecting potential data races in multithreaded programs. Our method improves on the state of the art in accuracy, in usability, and in overhead. We improve accuracy by combining previously known race detection techniques -- lockset-based d…
Eli Pozniansky, Assaf Schuster
Data race detection is highly essential for debugging multithreaded programs and assuring their correctness. Nevertheless, there is no single universal technique capable of handling the task efficiently, since the data race detection problem is computationally hard in the general…
Manohar K. Prabhu, Kunle Olukotun
In this paper, we provide examples of how thread-level speculation (TLS) simplifies manual parallelization and enhances its performance. A number of techniques for manual parallelization using TLS are presented and results are provided that indicate the performance contribution o…
Diego Puppin
No abstract available.
Steven Saunders, Lawrence Rauchwerger
ARMI is a communication library that provides a framework for expressing fine-grain parallelism and mapping it to a particular machine using shared-memory and message passing library calls. The library is an advanced implementation of the RMI protocol and handles low-level detail…
Jeffrey M. Squyres
No abstract available.
Kai Tan, Duane Szafron, Jonathan Schaeffer, John Anvik, Steve MacDonald
A design pattern is a mechanism for encapsulating the knowledge of experienced designers into a re-usable artifact. Parallel design patterns reflect commonly occurring parallel communication and synchronization structures. Our tools, CO2P3S (Correct Object-Oriented Pattern-based …
Kenjiro Taura, Kenji Kaneda, Toshio Endo, Akinori Yonezawa
This paper proposes Phoenix, a programming model for writing parallel and distributed applications that accommodate dynamically joining/leaving compute resources. In the proposed model, nodes involved in an application see a large and fixed virtual node name space. They communica…
Enrique V. Carrera, Ricardo Bianchini
Efficiency and portability are conflicting objectives for cluster-based network servers that distribute the clients' requests across the cluster based on the actual content requested. Our work is based on the observation that this efficiency vs. portability tradeoff has not been …
Andrew A. Chien
No abstract available.
Ian T. Foster
No abstract available.
Fumihiko Ino, Noriyuki Fujimoto, Kenichi Hagihara
We present a new parallel computational model, named LogGPS, which captures synchronization.
Seon Wook Kim, Chong-liang Ooi, Rudolf Eigenmann, Babak Falsafi, T. N. Vijaykumar
Recent proposals for multithreaded architectures allow threads with unknown dependences to execute speculatively in parallel. These architectures use hardware speculative storage to buffer uncertain data, track data dependences and roll back incorrect executions. Because all memo…