Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of Variance
Abstract
Deep learning (DL) training algorithms utilize nondeterminism to improve models' accuracy and training efficiency. Hence, multiple identical training runs (e.g., identical training data, algorithm, and network) produce different models with different accuracies and training times. In addition to these algorithmic factors, DL libraries (e.g., TensorFlow and cuDNN) introduce additional variance (referred to as implementation-level variance) due to parallelism, optimization, and floating-point computation.
BibTeX
@inproceedings{Pham-al:ASE20,
author = {Hung Viet Pham and
Shangshu Qian and
Jiannan Wang and
Thibaud Lutellier and
Jonathan Rosenthal and
Lin Tan and
Yaoliang Yu and
Nachiappan Nagappan},
title = {Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of Variance},
booktitle = {ASE},
pages = {771--783},
publisher = {{IEEE}},
year = {2020},
}