Preference-Guided Refactored Tuning for Retrieval Augmented Code Generation
Abstract
Retrieval-augmented code generation utilizes Large Language Models as the generator and significantly expands their code generation capabilities by providing relevant code, documentation, and more via the retriever. The current approach suffers from two primary limitations: 1) information redundancy. The indiscriminate inclusion of redundant information can result in resource wastage and may misguide generators, affecting their effectiveness and efficiency. 2) preference gap. Due to different optimization objectives, the retriever strives to procure code with higher ground truth similarity, yet this effort does not substantially benefit the generator. The retriever and the generator may prefer different golden code, and this gap in preference results in a suboptimal design. Additionally, differences in parameterization knowledge acquired during pre-training result in varying preferences among different generators.
BibTeX
@inproceedings{Gao-al:ASE24,
author = {Xinyu Gao and
Yun Xiong and
Deze Wang and
Zhenhan Guan and
Zejian Shi and
Haofen Wang and
Shanshan Li},
title = {{Preference-Guided} Refactored Tuning for Retrieval Augmented Code Generation},
booktitle = {ASE},
pages = {65--77},
publisher = {{ACM}},
year = {2024},
}