An empirical study on crash recovery bugs in large-scale distributed systems
Abstract
In large-scale distributed systems, node crashes are inevitable, and can happen at any time. As such, distributed systems are usually designed to be resilient to these node crashes via various crash recovery mechanisms, such as write-ahead logging in HBase and hinted handoffs in Cassandra. However, faults in crash recovery mechanisms and their implementations can introduce intricate crash recovery bugs, and lead to severe consequences.
BibTeX
@inproceedings{Gao-al:FSE18,
author = {Yu Gao and
Wensheng Dou and
Feng Qin and
Chushu Gao and
Dong Wang and
Jun Wei and
Ruirui Huang and
Li Zhou and
Yongming Wu},
title = {An empirical study on crash recovery bugs in large-scale distributed systems},
booktitle = {{ESEC/SIGSOFT} {FSE}},
pages = {539--550},
publisher = {{ACM}},
year = {2018},
}