arXiv · 2610.11598
Overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) Task
Abstract
In this paper, we provide an overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) task. Building on the success of the NTCIR-18 core task AEOLLM, we proposed AEOLLM-2 for NTCIR-19 to further investigate automatic evaluation methods for Large Language Models (LLMs), particularly in long-form text generation scenarios. In AEOLLM-2, we introduced a new subtask, Deep Research Evaluation, which focuses on the automatic evaluation of long-form deep research reports generated by LLMs. Participants developed evaluation methods to automatically assess the quality of these reports, and the performance of each method was measured by comparing its scores against human-annotated ground-truth labels. This year, we received 91 runs from 10 teams in total. This paper describes the background of the task, the dataset construction, the evaluation measures, the participants' methods, and the final evaluation results.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Junjie Chen, Yuxi Dong, Haitao Li, Yiqun Liu, Qingyao Ai. 2026-10-08. Overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) Task. https://arxiv.org/abs/2610.11598
Cite the original work for its findings. Save a collection to share your selection of sources.