kirancodes.me
To Proof Maintenance & Beyond!

JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models

Jialun Cao, Zhiyong Chen, Jiarong Wu, Shing-Chi Cheung, Chang Xu

Abstract

Code generation benchmarks such as HumanEval are widely adopted to evaluate LLMs' capabilities. However, after consolidating the latest 24 benchmarks, we noticed three significant imbalances. First, imbalanced programming language. 95.8% of benchmarks involve Python, while only 5 benchmarks involve Java, resulting in an insufficient understanding of LLMs' capability to generate Java code. Second, imbalanced code granularity. Function-/statement-level benchmarks account for over 83.3% of benchmarks. Only a mere handful extends to class-/project-levels, and all are limited to Python. Third, lacking advanced features. Existing benchmarks primarily assess basic coding skills (e.g., variables, operators, and control structures), while overlooking advanced Object-Oriented Programming (OOP) features (i.e., encapsulation, inheritance, and polymorphism). Considering the prevalence of these advanced features in real-world Java project development, constructing benchmarks to test LLMs on handling OOP features is necessary.

BibTeX
@inproceedings{Cao-al:ASE24,
  author    = {Jialun Cao and
               Zhiyong Chen and
               Jiarong Wu and
               Shing{-}Chi Cheung and
               Chang Xu},
  title     = {{JavaBench:} A Benchmark of {Object-Oriented} Code Generation for Evaluating Large Language Models},
  booktitle = {ASE},
  pages     = {870--882},
  publisher = {{ACM}},
  year      = {2024},
}

Related papers