arXiv · 2502.00064
Evaluating Large Language Models in Vulnerability Detection Under Variable Context Windows
Abstract
This study examines the impact of tokenized Java code length on the accuracy and explicitness of ten major LLMs in vulnerability detection. Using chi-square tests and known ground truth, we found inconsistencies across models: some, like GPT-4, Mistral, and Mixtral, showed robustness, while others exhibited a significant link between tokenized length and performance. We recommend future LLM development focus on minimizing the influence of input length for better vulnerability detection. Additionally, preprocessing techniques that reduce token count while preserving code structure could enhance LLM accuracy and explicitness in these tasks.
Explore related subjects
Keep this discovery
Jie Lin, David Mohaisen. 2025-01-30. Evaluating Large Language Models in Vulnerability Detection Under Variable Context Windows. https://arxiv.org/abs/2502.00064
Cite the original work for its findings. Save a collection to share your selection of sources.