kirancodes.me
To Proof Maintenance & Beyond!

Bridging Pre-trained Models and Downstream Tasks for Source Code Understanding

Deze Wang, Zhouyang Jia, Shanshan Li, Yue Yu, Yun Xiong, Wei Dong, Xiangke Liao

Abstract

With the great success of pre-trained models, the pretrain-then-finetune paradigm has been widely adopted on downstream tasks for source code understanding. However, compared to costly training a large-scale model from scratch, how to effectively adapt pre-trained models to a new task has not been fully explored. In this paper, we propose an approach to bridge pre-trained models and code-related tasks. We exploit semantic-preserving transformation to enrich downstream data diversity, and help pre-trained models learn semantic features invariant to these semantically equivalent transformations. Further, we introduce curriculum learning to organize the transformed data in an easy-to-hard manner to fine-tune existing pre-trained models.

BibTeX
@inproceedings{Wang-al:ICSE22,
  author    = {Deze Wang and
               Zhouyang Jia and
               Shanshan Li and
               Yue Yu and
               Yun Xiong and
               Wei Dong and
               Xiangke Liao},
  title     = {Bridging Pre-trained Models and Downstream Tasks for Source Code Understanding},
  booktitle = {ICSE},
  pages     = {287--298},
  publisher = {{ACM}},
  year      = {2022},
}

Related papers