arXiv · 2504.00065
Assessing Code Understanding in LLMs
Abstract
We present an empirical evaluation of Large Language Models in code understanding associated with non-trivial, semantic-preserving program transformations such as copy propagation or constant folding. Our findings show that LLMs fail to judge semantic equivalence in approximately 41\% of cases when no context is provided and in 29\% when given a simple generic context. To improve accuracy, we advocate integrating LLMs with code-optimization tools to enhance training and facilitate more robust program understanding.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Cosimo Laneve, Alvise Spanò, Dalila Ressi, Sabina Rossi, Michele Bugliesi. 2025-03-31. Assessing Code Understanding in LLMs. https://arxiv.org/abs/2504.00065
Cite the original work for its findings. Save a collection to share your selection of sources.