JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
Abstract
Code generation benchmarks such as HumanEval are widely adopted to evaluate LLMs' capabilities. However, after consolidating the latest 24 benchmarks, we noticed three significant imbalances. First, imbalanced programming language. 95.8% of benchmarks involve Python, while only 5 benchmarks involve Java, resulting in an insufficient understanding of LLMs' capability to generate Java code. Second, imbalanced code granularity. Function-/statement-level benchmarks account for over 83.3% of benchmarks. Only a mere handful extends to class-/project-levels, and all are limited to Python. Third, lacking advanced features. Existing benchmarks primarily assess basic coding skills (e.g., variables, operators, and control structures), while overlooking advanced Object-Oriented Programming (OOP) features (i.e., encapsulation, inheritance, and polymorphism). Considering the prevalence of these advanced features in real-world Java project development, constructing benchmarks to test LLMs on handling OOP features is necessary.
BibTeX
@inproceedings{Cao-al:ASE24,
author = {Jialun Cao and
Zhiyong Chen and
Jiarong Wu and
Shing{-}Chi Cheung and
Chang Xu},
title = {{JavaBench:} A Benchmark of {Object-Oriented} Code Generation for Evaluating Large Language Models},
booktitle = {ASE},
pages = {870--882},
publisher = {{ACM}},
year = {2024},
}