kirancodes.me
To Proof Maintenance & Beyond!

Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code Models

Shuzheng Gao, Wenxin Mao, Cuiyun Gao, Li Li, Xing Hu, Xin Xia, Michael R. Lyu

Abstract

Pre-trained code models have recently achieved substantial improvements in many code intelligence tasks. These models are first pre-trained on large-scale unlabeled datasets in a task-agnostic manner using self-supervised learning, and then fine-tuned on labeled datasets in downstream tasks. However, the labeled datasets are usually limited in size (i.e., human intensive efforts), which may hinder the performance of pre-trained code models in specific tasks. To mitigate this, one possible solution is to leverage the large-scale unlabeled data in the tuning stage by pseudo-labeling, i.e., generating pseudo labels for unlabeled data and further training the pre-trained code models with the pseudo-labeled data. However, directly employing the pseudo-labeled data can bring a large amount of noise, i.e., incorrect labels, leading to suboptimal performance. How to effectively leverage the noisy pseudo-labeled data is a challenging yet under-explored problem.

BibTeX
@inproceedings{Gao-al:ICSE24,
  author    = {Shuzheng Gao and
               Wenxin Mao and
               Cuiyun Gao and
               Li Li and
               Xing Hu and
               Xin Xia and
               Michael R. Lyu},
  title     = {Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code Models},
  booktitle = {ICSE},
  pages     = {80:1--80:13},
  publisher = {{ACM}},
  year      = {2024},
}

Related papers