kirancodes.me
To Proof Maintenance & Beyond!

Unveiling Memorization in Code Models

Zhou Yang, Zhipeng Zhao, Chenyu Wang, Jieke Shi, Dongsun Kim, DongGyun Han, David Lo

Abstract

The availability of large-scale datasets, advanced architectures, and powerful computational resources have led to effective code models that automate diverse software engineering activities. The datasets usually consist of billions of lines of code from both open-source and private repositories. A code model memorizes and produces source code verbatim, which potentially contains vulnerabilities, sensitive information, or code with strict licenses, leading to potential security and privacy issues.

BibTeX
@inproceedings{Yang-al:ICSE24,
  author    = {Zhou Yang and
               Zhipeng Zhao and
               Chenyu Wang and
               Jieke Shi and
               Dongsun Kim and
               DongGyun Han and
               David Lo},
  title     = {Unveiling Memorization in Code Models},
  booktitle = {ICSE},
  pages     = {72:1--72:13},
  publisher = {{ACM}},
  year      = {2024},
}

Related papers