Big code != big vocabulary: open-vocabulary models for source code
Abstract
Statistical language modeling techniques have successfully been applied to large source code corpora, yielding a variety of new software development tools, such as tools for code suggestion, improving readability, and API migration. A major issue with these techniques is that code introduces new vocabulary at a far higher rate than natural language, as new identifier names proliferate. Both large vocabularies and out-of-vocabulary issues severely affect Neural Language Models (NLMs) of source code, degrading their performance and rendering them unable to scale.
BibTeX
@inproceedings{Karampatsis-al:ICSE20,
author = {Rafael{-}Michael Karampatsis and
Hlib Babii and
Romain Robbes and
Charles Sutton and
Andrea Janes},
title = {Big code != big vocabulary: open-vocabulary models for source code},
booktitle = {ICSE},
pages = {1073--1085},
publisher = {{ACM}},
year = {2020},
}