From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench

2026-08-27Software Engineering

Software EngineeringArtificial IntelligenceComputation and Language
AI summary

The authors created MCR-Bench, a new benchmark to better mimic real code reviews that happen in multiple back-and-forth rounds, unlike previous tests that treated code review as a one-time check. Their benchmark includes lots of real-world examples with detailed information about bugs and how they change during the review process. They tested popular large language models (LLMs) on this benchmark and found that these models struggle to track defects well, especially over multiple rounds. The models also have trouble catching more complex or subtle bugs and often make mistakes because they can’t remember or align details across review rounds properly.

Code ReviewLarge Language ModelsDefect DetectionDefect LifecycleMulti-round InteractionBenchmark DatasetError AnalysisTemporal MisalignmentLong-range MemorySoftware Quality
Authors
Dewu Zheng, Yanlin Wang, Xiwen Wang, Kefeng Duan, Hongyu Zhang, Xilin Liu, Yuchi Ma, Zibin Zheng
Abstract
In real-world software development, code review typically involves iterative interactions between developers and reviewers to improve software quality, making the process costly and time-consuming. Although recent work explores large language models (LLMs) for automated code review, most approaches oversimplify code review into a single-round, static decision task, which fails to capture the multi-round interactive nature and the complex problem-solving processes inherent in realistic review scenarios. To bridge this gap, we introduce MCR-Bench, the first defect state-aware benchmark designed for realistic multi-round code review. MCR-Bench covers five commonly-used programming languages and consists of 2,269 real-world multi-round code review tasks, each of which is annotated with fine-grained defect information and cross-round state labels. Each task in MCR-Bench is equipped with fine-grained defect metadata (e.g., description, type, severity) alongside dynamic state annotations, capturing the complete evolutionary trajectory of a defect throughout the multi-round process. We obtain several findings through extensive experiments on MCR-Bench with mainstream LLMs. (1) Limited overall capability: experiments reveal that mainstream LLMs exhibit limited overall performance in defect detection and defect lifecycle state tracking, with performance degrading significantly as the number of interaction rounds increases; (2) Defect-sensitive performance: LLMs' performance varies substantially across different defect types and severity levels, with semantically complex or low-salience defects being significantly more likely to be missed; (3) Underlying Failure Mechanisms: our in-depth error analysis dissects the distinct drivers of false positives and false negatives, revealing critical weaknesses such as cross-round temporal misalignment and inadequate long-range memory.