arXiv · 2507.21130
INTEGRALBENCH: Benchmarking LLMs with Definite Integral Problems
Abstract
We present INTEGRALBENCH, a focused benchmark designed to evaluate Large Language Model (LLM) performance on definite integral problems. INTEGRALBENCH provides both symbolic and numerical ground truth solutions with manual difficulty annotations. Our evaluation of nine state-of-the-art LLMs reveals significant performance gaps and strong correlations between problem difficulty and model accuracy, establishing baseline metrics for this challenging domain. INTEGRALBENCH aims to advance automated mathematical reasoning by providing a rigorous evaluation framework specifically tailored for definite integral computation.
Explore related subjects
Keep this discovery
Bintao Tang, Xin Yang, Yuhao Wang, Zixuan Qiu, Zimo Ji, Wenyuan Jiang. 2025-07-22. INTEGRALBENCH: Benchmarking LLMs with Definite Integral Problems. https://arxiv.org/abs/2507.21130
Cite the original work for its findings. Save a collection to share your selection of sources.