VarCLR: Variable Semantic Representation Pre-training via Contrastive Learning
Abstract
Variable names are critical for conveying intended program behavior. Machine learning-based program analysis methods use variable name representations for a wide range of tasks, such as suggesting new variable names and bug detection. Ideally, such methods could capture semantic relationships between names beyond syntactic similarity, e.g., the fact that the names average and mean are similar. Unfortunately, previous work has found that even the best of previous representation approaches primarily capture "relatedness" (whether two variables are linked at all), rather than "similarity" (whether they actually have the same meaning).
BibTeX
@inproceedings{Chen-al:ICSE22,
author = {Qibin Chen and
Jeremy Lacomis and
Edward J. Schwartz and
Graham Neubig and
Bogdan Vasilescu and
Claire Le Goues},
title = {{VarCLR:} Variable Semantic Representation Pre-training via Contrastive Learning},
booktitle = {ICSE},
pages = {2327--2339},
publisher = {{ACM}},
year = {2022},
}