SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation
Abstract
Automated Code Review (ACR) is crucial for software quality, yet existing benchmarks often fail to reflect real-world complexities, hindering the evaluation of modern Large Language Models (LLMs). Current benchmarks frequently focus on fine-grained code units, lack complete project context, and use inadequate evaluation metrics. To address these limitations, we introduce SWR-Bench, a new benchmark comprising 1000 manually verified Pull Requests (PRs) from GitHub, offering PR-centric review with full project context. SWR-Bench employs an objective LLM-based evaluation method that aligns strongly with human judgment (∼90% agreement) by verifying if issues from a structured ground truth are covered in generated reviews. Our systematic evaluation of mainstream ACR tools and LLMs on SWR-Bench reveals that current systems underperform, and ACR tools are more adept at detecting functional errors. Subsequently, we propose and validate a simple multi-review aggregation strategy that significantly boosts ACR performance, increasing F1 scores by up to 43.67%. Our contributions include the SWR-Bench benchmark, its objective evaluation method, a comprehensive study of current ACR capabilities, and an effective enhancement approach, offering valuable insights for advancing ACR research.
BibTeX
@article{Zeng-al:FSE26,
author = {Zhengran Zeng and
Ruikai Shi and
Keke Han and
Yixin Li and
Kaicheng Sun and
Yidong Wang and
Zhuohao Yu and
Rui Xie and
Wei Ye and
Shikun Zhang},
title = {{SWR-Bench:} Assessing {LLM} Performance in {Real-World} Code Review Comment Generation},
journal = {{PACMSE}},
volume = {3},
number = {{FSE}},
pages = {3093--3115},
year = {2026},
}