kirancodes.me
To Proof Maintenance & Beyond!

Parallelism by design: data analysis with sawzall

Robert Griesemer

Abstract

Very large data sets - telephone call records, network logs, high-resolution satellite images, or web document repositories - are not easily analyzed using traditional database techniques. They may be simply too large, grow too fast, or may not fit well in a database schema. They tend to span multiple disks and machines. On the other hand, these large data sets often have a flat and regular structure that permits distributed filtering and aggregation.

Related papers