kirancodes.me
To Proof Maintenance & Beyond!

Promise and Peril of Collaborative Code Generation Models: Balancing Effectiveness and Memorization

Zhi Chen, Lingxiao Jiang

Abstract

In the rapidly evolving field of machine learning, training models with datasets from various locations and organizations presents significant challenges due to privacy and legal concerns. The exploration of effective collaborative training settings, which are capable of leveraging valuable knowledge from distributed and isolated datasets, is increasingly crucial. This study investigates key factors that impact the effectiveness of collaborative training methods in code next-token prediction, as well as the correctness and utility of the generated code, showing the promise of such methods. Additionally, we evaluate the memorization of different participant training data across various collaborative training settings, including centralized, federated, and incremental training, showing their potential risks in leaking data.

BibTeX
@inproceedings{Chen-Jiang:ASE24,
  author    = {Zhi Chen and
               Lingxiao Jiang},
  title     = {Promise and Peril of Collaborative Code Generation Models: Balancing Effectiveness and Memorization},
  booktitle = {ASE},
  pages     = {493--505},
  publisher = {{ACM}},
  year      = {2024},
}

Related papers