arXiv · 2609.18022
VeriBugBench: An Empirically Grounded Framework for Constructing Verilog RTL Debugging Benchmarks
Abstract
RTL source-level debugging research requires benchmark artifacts that provide faulty designs together with precise change locations, executable test stimuli, and reproducible configurations. Available Verilog resources usually provide only a subset of these elements. We present VeriBugBench, a framework for constructing Verilog RTL debugging benchmarks through empirically grounded fault construction, LLM-based testbench enhancement, and execution-based retention. The mutation library maps recurring, multi-granularity repair patterns observed in RTL bug-fix histories to 19 executable inverse operators. For each project, an LLM generates a design-specific stimulus phase from the clean DUT and original testbench; the phase is composed with the original testbench for candidate execution. Applying the framework to 45 open-source projects yields VeriBugBench-v1.0, with 2,608 executable single-fault instances whose effects are observable at design outputs. Across the 45 projects, the assembled testbenches increase mean project-level fault observability from 36.01% to 39.54% and improve line coverage and execution-trace diversity on average. VeriBugBench provides versioned RTL variants, source-level ground truth, testbenches, and execution artifacts for evaluating RTL debugging methods.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xiankai Meng, Kejian Feng, Xinlin Zhao, Zhuo Zhang, Yan Lei, Xiaoguang Mao, Jiang Wu. 2026-09-18. VeriBugBench: An Empirically Grounded Framework for Constructing Verilog RTL Debugging Benchmarks. https://arxiv.org/abs/2609.18022
Cite the original work for its findings. Save a collection to share your selection of sources.