arXiv · 2504.21205
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
Abstract
This paper introduces SecRepoBench, a benchmark to evaluate code agents on secure code completion in real-world repositories. SecRepoBench has 318 code completion tasks in 27 C/C++ repositories, covering 15 CWEs. We evaluate 29 standalone LLMs and 15 code agents across 3 state-of-the-art agent frameworks using our benchmark. We find that state-of-the-art LLMs struggle with generating correct and secure code completions. However, code agents significantly outperform standalone LLMs. We show that SecRepoBench is more difficult than the prior state-of-the-art benchmark. Finally, our comprehensive analysis provides insights into potential directions for enhancing the ability of code agents to write correct and secure code in real-world repositories.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chihao Shen, Connor Dilgren, Purva Chiniya, Luke Griffith, Yu Ding, Yizheng Chen. 2025-04-29. SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories. https://arxiv.org/abs/2504.21205
Cite the original work for its findings. Save a collection to share your selection of sources.