Detection Is Better Than Cure: A Cloud Incidents Perspective
Abstract
Cloud providers use automated watchdogs or monitors to continuously observe service availability and to proactively report incidents when system performance degrades. Improper monitoring can lead to delays in the detection and mitigation of production incidents, which can be extremely expensive in terms of customer impacts and manual toil from engineering resources. Therefore, a systematic understanding of the pitfalls in current monitoring practices and how they can lead to production incidents is crucial for ensuring continuous reliability of cloud services.
BibTeX
@inproceedings{Ganatra-al:FSE23,
author = {Vaibhav Ganatra and
Anjaly Parayil and
Supriyo Ghosh and
Yu Kang and
Minghua Ma and
Chetan Bansal and
Suman Nath and
Jonathan Mace},
title = {Detection Is Better Than Cure: A Cloud Incidents Perspective},
booktitle = {{ESEC/SIGSOFT} {FSE}},
pages = {1891--1902},
publisher = {{ACM}},
year = {2023},
}