kirancodes.me
To Proof Maintenance & Beyond!

RegMiner: mining replicable regression dataset from code repositories

Xuezhi Song, Yun Lin, Yijian Wu, Yifan Zhang, Siang Hwee Ng, Xin Peng, Jin Song Dong, Hong Mei

Abstract

In this work, we introduce a tool, RegMiner, to automate the process of collecting replicable regression bugs from a set of Git repositories. In the code commit history, RegMiner searches for regressions where a test can pass a regression-fixing commit, fail a regressioninducing commit, and pass a previous working commit again. Technically, RegMiner (1) identifies potential regression-fixing commits from the code evolution history, (2) migrates the test and its code dependencies in the commit over the history, and (3) minimizes the compilation overhead during the regression search. Our experients show that RegMiner can successfully collect 1035 regressions over 147 projects in 8 weeks, creating the largest replicable regression dataset within the shortest period, to the best of our knowledge. In addition, our experiments further show that (1) RegMiner can construct the regression dataset with very high precision and acceptable recall, and (2) the constructed regression dataset is of high authenticity and diversity. The source code of RegMiner is available at https://github.com/SongXueZhi/RegMiner, the mined regression dataset is available at https://regminer.github.io/, and the demonstration video is available at https://youtu.be/yzcM9Y4unok.

BibTeX
@inproceedings{Song-al:FSE22,
  author    = {Xuezhi Song and
               Yun Lin and
               Yijian Wu and
               Yifan Zhang and
               Siang Hwee Ng and
               Xin Peng and
               Jin Song Dong and
               Hong Mei},
  title     = {{RegMiner:} mining replicable regression dataset from code repositories},
  booktitle = {{ESEC/SIGSOFT} {FSE}},
  pages     = {1711--1715},
  publisher = {{ACM}},
  year      = {2022},
}

Related papers