1,184 papers · page 58 of 60
David B. Wagner, Brad Calder
A future is a language construct that allows programmers to expose parallelism in applicative languages such as MultiLisp [5] with minimal effort. In this paper we describe a technique for implementing futures, which we call leapfrogging, that reduces blocking due to load imbalan…
Donald Yeung, Anant Agarwal
This paper discusses our experience with fine-grain synchronization for a variant of the preconditioned conjugate gradient method. This algorithm represents a large class of algorithms that have been widely used but traditionally difficult to implement efficiently on vector and p…
Corinne Ancourt, François Irigoin
Supercompilers perform complex program transformations which often result in new loop bounds. This paper shows that, under the usual assumptions in automatic parallelization, most transformations on loop nests can be expressed as affine transformations on integer sets de ned by p…
David F. Bacon, Robert E. Strom
We present a transparent program transformation which converts a sequential execution of S1 ; S2 by a process in a multiprocess environment into an optimistic parallel execution of S1 and S2 . Such a transformation is valuable in the case where S1 and S2 cannot be parallelized by…
Hester Bakewell, Donna J. Quammen, Pearl Y. Wang
Vasanth Balasundaram, Geoffrey C. Fox, Ken Kennedy, Ulrich Kremer
The choice of the data domain partitioning scheme is an important factor in determining the available parallelism and hence the performance of an application on a distributed memory multiprocessor.In this paper, we present a performance estimator for statically evaluating the rel…
Jong-Deok Choi, Sang Lyul Min
Races in describes a mechanism to debug unintended data races
Vítor Santos Costa, David H. D. Warren, Rong Yang
article Andorra I: a parallel Prolog system that transparently exploits both And-and or-parallelism Share on Authors: Vítor Santos Costa View Profile , David H. D. Warren View Profile , Rong Yang View Profile Authors Info & Claims ACM SIGPLAN NoticesVolume 26Issue 7July 1991 pp 8…
Michael J. Feeley, Brian N. Bershad, Jeffrey S. Chase, Henry M. Levy
Idle workstationsin a network represent a significant computing potential.In particular, their processing power can be used by parallel-distributed programs that treat the network as a loosely-coupled multiprocessor.Our experiments with Amber show that node reconfiguration can be…
Philip J. Hatcher, Anthony J. Lapadula, Robert R. Jones, Michael J. Quinn, Ray J. Anderson
article Free Access Share on A production-quality C* compiler for Hypercube multicomputers Authors: Philip J. Hatcher View Profile , Anthony J. Lapadula View Profile , Robert R. Jones View Profile , Michael J. Quinn View Profile , Ray J. Anderson View Profile Authors Info & Claim…
Jeffrey K. Hollingsworth, R. Bruce Irvin, Barton P. Miller
The IPS-2 parallel program measurement tools pro-vide performance data from application programs, the operating system, hardware, network, and other sources. Previous versions of IPS-2 allowed programmers to collect information about an application based only on what could be col…
Dz-Ching Ju, Wai-Mee Ching
Programs written m APL implicitly contain data parallelism because the high level APL primitives denoting array o erations may be executed in parallel.Our experiment $ APL/C compiler translates ordinary APL programs into the C language with additional parallel constmcts for synch…
V. Prasad Krothapalli, P. Sadayappan
article Free Access Share on Removal of redundant dependences in DOACROSS loops with constant dependences Authors: V. P. Krothapalli View Profile , P. Sadayappan View Profile Authors Info & Claims ACM SIGPLAN NoticesVolume 26Issue 7July 1991 pp 51–60https://doi.org/10.1145/109626…
H. T. Kung, Peter Steenkiste, Marco Dimas Gubitoso, Manpreet Khaira
Several large applicationshave been paralleli,zed on Nectar, a network-based multicomputer recently developed by Carnegie Mellon.These applications were previously either too large or too complex to be easily implemented on distributed memory parallel systems.Parallelizing these …
Richard P. LaRowe Jr., James T. Wilkes, Carla Schlatter Ellis
Shared memory multiprocessors are attractive because they are programmed in a manner similar to uniprocessors. The UMA class of shared memory multiprocessors is the most attractive, from the programmer's point of view, since the programmer need not be concerned with the placement…
Monica S. Lam, Martin C. Rinard
This paper presents Jade, a language which allows a programmer to easily express dynamic coarse-grain parallelism. Starting with a sequential program, a programmer augments those sections of code to be parallelized with abstract data usage information. The compiler and run-time s…
Lee-Chung Lu
This paper presents a formal mathematical framework which unifies the existing loop transformations.This framework also includes more general classes of loop transformations, which can extract more parallelism from a class of programs than the existing techniques.We classify sche…
Allen D. Malony
Determiningthe performance behavior of parallel computations requires some form of intrusive tracing measurement.The greater the need for detailed performance data, the more intrusion the measurement will cause.Recovering actual execution performance jfrom perturbed performance m…
Ulrike Meier, Rudolf Eigenmann
We analyze the computational structure of the Conjugate Gradient algorithm and we describe its parallel implement ation on the Cedar hierarchicalmemory multiprocessor.The analysis will cover both explicit manual parallelization and automatic compilation.We report performance meas…
John M. Mellor-Crummey, Michael L. Scott
Reader-writer synchronization relaxes the constraints of mutual exclusion to permit more than one process to inspect a shared object concurrently, as long as none of them changes its value.On uniprocessors, mutual exclusion and readercessor demonstrate that our algorithms provide…