Can Machine Learning Pipelines Be Better Configured?
Abstract
A Machine Learning (ML) pipeline configures the workflow of a learning task using the APIs provided by ML libraries. However, a pipeline’s performance can vary significantly across different configurations of ML library versions. Misconfigured pipelines can result in inferior performance, such as poor execution time and memory usage, numeric errors and even crashes. A pipeline is subject to misconfiguration if it exhibits significantly inconsistent performance upon changes in the versions of its configured libraries or the combination of these libraries. We refer to such performance inconsistency as a pipeline configuration (PLC) issue.
BibTeX
@inproceedings{Wang-al:FSE23,
author = {Yibo Wang and
Ying Wang and
Tingwei Zhang and
Yue Yu and
Shing{-}Chi Cheung and
Hai Yu and
Zhiliang Zhu},
title = {Can Machine Learning Pipelines Be Better Configured?},
booktitle = {{ESEC/SIGSOFT} {FSE}},
pages = {463--475},
publisher = {{ACM}},
year = {2023},
}