kirancodes.me
To Proof Maintenance & Beyond!
ICSE 2020★ Distinguished Paper

Big code != big vocabulary: open-vocabulary models for source code

Rafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton, Andrea Janes

Abstract

Statistical language modeling techniques have successfully been applied to large source code corpora, yielding a variety of new software development tools, such as tools for code suggestion, improving readability, and API migration. A major issue with these techniques is that code introduces new vocabulary at a far higher rate than natural language, as new identifier names proliferate. Both large vocabularies and out-of-vocabulary issues severely affect Neural Language Models (NLMs) of source code, degrading their performance and rendering them unable to scale.

BibTeX
@inproceedings{Karampatsis-al:ICSE20,
  author    = {Rafael{-}Michael Karampatsis and
               Hlib Babii and
               Romain Robbes and
               Charles Sutton and
               Andrea Janes},
  title     = {Big code != big vocabulary: open-vocabulary models for source code},
  booktitle = {ICSE},
  pages     = {1073--1085},
  publisher = {{ACM}},
  year      = {2020},
}

Related papers