kirancodes.me
To Proof Maintenance & Beyond!

Contextualized Data-Wrangling Code Generation in Computational Notebooks

Junjie Huang, Daya Guo, Chenglong Wang, Jiazhen Gu, Shuai Lu, Jeevana Priya Inala, Cong Yan, Jianfeng Gao, Nan Duan, Michael R. Lyu

Abstract

Data wrangling, the process of preparing raw data for further analysis in computational notebooks, is a crucial yet time-consuming step in data science. Code generation has the potential to automate the data wrangling process to reduce analysts' overhead by translating user intents into executable code. Precisely generating data wrangling code necessitates a comprehensive consideration of the rich context present in notebooks, including textual context, code context and data context. However, notebooks often interleave multiple non-linear analysis tasks into linear sequence of code blocks, where the contextual dependencies are not clearly reflected. Directly training models with source code blocks fails to fully exploit the contexts for accurate wrangling code generation.

BibTeX
@inproceedings{Huang-al:ASE24,
  author    = {Junjie Huang and
               Daya Guo and
               Chenglong Wang and
               Jiazhen Gu and
               Shuai Lu and
               Jeevana Priya Inala and
               Cong Yan and
               Jianfeng Gao and
               Nan Duan and
               Michael R. Lyu},
  title     = {Contextualized {Data-Wrangling} Code Generation in Computational Notebooks},
  booktitle = {ASE},
  pages     = {1282--1294},
  publisher = {{ACM}},
  year      = {2024},
}

Related papers