kirancodes.me
To Proof Maintenance & Beyond!

Co-dependence Aware Fuzzing for Dataflow-Based Big Data Analytics

Ahmad Humayun, Miryung Kim, Muhammad Ali Gulzar

Abstract

Data-intensive scalable computing has become popular due to the increasing demands of analyzing big data. For example, Apache Spark and Hadoop allow developers to write dataflow-based applications with user-defined functions to process data with custom logic. Testing such applications is difficult. (1) These applications often take multiple datasets as input. (2) Unlike in SQL, there is no explicit schema for these datasets and each unstructured (or semi-structured) dataset is segmented and parsed at runtime. (3) Dataflow operators (e.g., join) create implicit co-dependence constraints between the fields of multiple datasets. An efficient and effective testing technique must analyze co-dependence among different regions of multiple datasets at the level of rows and columns and orchestrate input mutations jointly on co-dependent regions.

BibTeX
@inproceedings{Humayun-al:FSE23,
  author    = {Ahmad Humayun and
               Miryung Kim and
               Muhammad Ali Gulzar},
  title     = {Co-dependence Aware Fuzzing for {Dataflow-Based} Big Data Analytics},
  booktitle = {{ESEC/SIGSOFT} {FSE}},
  pages     = {1050--1061},
  publisher = {{ACM}},
  year      = {2023},
}

Related papers