kirancodes.me
To Proof Maintenance & Beyond!

An empirical study on crash recovery bugs in large-scale distributed systems

Yu Gao, Wensheng Dou, Feng Qin, Chushu Gao, Dong Wang, Jun Wei, Ruirui Huang, Li Zhou, Yongming Wu

Abstract

In large-scale distributed systems, node crashes are inevitable, and can happen at any time. As such, distributed systems are usually designed to be resilient to these node crashes via various crash recovery mechanisms, such as write-ahead logging in HBase and hinted handoffs in Cassandra. However, faults in crash recovery mechanisms and their implementations can introduce intricate crash recovery bugs, and lead to severe consequences.

BibTeX
@inproceedings{Gao-al:FSE18,
  author    = {Yu Gao and
               Wensheng Dou and
               Feng Qin and
               Chushu Gao and
               Dong Wang and
               Jun Wei and
               Ruirui Huang and
               Li Zhou and
               Yongming Wu},
  title     = {An empirical study on crash recovery bugs in large-scale distributed systems},
  booktitle = {{ESEC/SIGSOFT} {FSE}},
  pages     = {539--550},
  publisher = {{ACM}},
  year      = {2018},
}

Related papers