arXiv · 2609.35461
AraDynFact: Dynamic Evaluation of Factual Knowledge in Arabic
Abstract
As Large Language Models (LLMs) continue to scale both in size and capabilities, their proficiency in the Arabic Language has seen significant advancement. However, a critical gap remains: the extent of their factual knowledge and cultural sensitivity to the diverse Arabic-speaking world remains largely underexplored. Current evaluation metrics often focus on translation or generic reasoning, failing to capture the rich historical, social, and regional nuances inherent to Arabic culture. In addition, most benchmarks rely on heavy work, with human intervention in some steps, making the evaluation of knowledge coverage expensive and slow. To address this deficiency, we introduce AraDynFact, a novel dynamic evaluation framework designed to rigorously assess the factual Arabic knowledge embedded in LLMs. Unlike static benchmarks, AraDynFact employs a dynamic approach to extract factual information and generate rich and answerable questions in a fast and automatic way. We apply AraDynFact to Arabic Wikipedia and audit the performance of several state-of-the-art models, ranging from Arabic-centric specialized LLMs to high-resource general purpose LLMs. In addition we found a high degree of correlation with existing, hand-crafted Arabic-centric benchmarks, confirming the potential of our dynamic approach.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ignacio Iacobacci, Faroq Altam, Zhaozhi Qian, Muhammad Alqurishi. 2026-09-28. AraDynFact: Dynamic Evaluation of Factual Knowledge in Arabic. https://arxiv.org/abs/2609.35461
Cite the original work for its findings. Save a collection to share your selection of sources.