arXiv ScienceSearch

arXiv · 2507.10579

Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors

Abstract

This shared task has aimed to assess pedagogical abilities of AI tutors powered by large language models (LLMs), focusing on evaluating the quality of tutor responses aimed at student's mistake remediation within educational dialogues. The task consisted of five tracks designed to automatically evaluate the AI tutor's performance across key dimensions of mistake identification, precise location of the mistake, providing guidance, and feedback actionability, grounded in learning science principles that define good and effective tutor responses, as well as the track focusing on detection of the tutor identity. The task attracted over 50 international teams across all tracks. The submitted models were evaluated against gold-standard human annotations, and the results, while promising, show that there is still significant room for improvement in this domain: the best results for the four pedagogical ability assessment tracks range between macro F1 scores of 58.34 (for providing guidance) and 71.81 (for mistake identification) on three-class problems, with the best F1 score in the tutor identification track reaching 96.98 on a 9-class task. In this paper, we overview the main findings of the shared task, discuss the approaches taken by the teams, and analyze their performance. All resources associated with this task are made publicly available to support future research in this critical domain.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ekaterina Kochmar, Kaushal Kumar Maurya, Kseniia Petukhova, KV Aditya Srivatsa, Anaïs Tack, Justin Vasselli. 2025-07-11. Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors. https://arxiv.org/abs/2507.10579

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Algorithmic Shortlisting in Participatory Budgeting

Participatory budgeting is a democratic innovation that allows citizens to propose and vote on public investment projects. To help organizers manage large volumes of submissions, we design and test privacy-preserving methods for algorithmic shortlisting. These algorithms predict which projects are likely to be funded using only project features and anonymous historical voting data. We demonstrate the limitations of a naive approach that uses a large language model to rank projects based on past success and propose a vote-based pipeline that enables state-of-the-art LLMs to perform on par with classical machine learning. Our findings indicate that user preferences in participatory budgeting are stable enough to allow algorithmic shortlisting to approximate an initial selection of projects effectively.

cs.CY

Human Resilience in the AI Era -- What Machines Can't Replace

AI is changing work and decision making faster than many institutions can adapt their operating practices. We argue that this adaptation gap makes human resilience a core capability for the AI era. We define resilience as the capacity to absorb disruption while preserving effective action and human agency around core purposes. The framework operates at three interacting levels. Psychological resilience keeps a person goal-directed under stress. Social resilience makes trusted support and correction available across a group. Organizational resilience turns detected problems into learning and recovery. We connect established resilience and technostress research with direct AI-in-the-loop experiments. General resilience is trainable, while AI-specific causal evidence is still emerging. Direct AI studies show that assistance can raise productivity and spread expertise. Other experiments show improved expressed empathy and more calibrated reliance. We translate these findings into a practical agenda for AI education, workplace design, governance, and evaluation. The central proposal is socio-technical: structural safeguards define the operating boundary, while resilient people and institutions provide adaptive capacity when conditions change.

cs.CY

The MEVIR Framework: A Virtue-Informed Moral-Epistemic Model of Human Trust Decisions

The 21st-century information landscape presents an unprecedented challenge: how do individuals make sound trust decisions amid complexity, polarization, and misinformation? Traditional rational-agent models fail to capture human trust formation, which involves a complex synthesis of reason, character, and pre-rational intuition. This report introduces the Moral-Epistemic VIRtue informed (MEVIR) framework, a comprehensive descriptive model integrating three theoretical perspectives: (1) a procedural model describing evidence-gathering and reasoning chains; (2) Linda Zagzebski's virtue epistemology, characterizing intellectual disposition and character-driven processes; and (3) Extended Moral Foundations Theory (EMFT), explaining rapid, automatic moral intuitions that anchor reasoning. Central to the framework are ontological concepts - Truth Bearers, Truth Makers, and Ontological Unpacking-revealing that disagreements often stem from fundamental differences in what counts as admissible reality. MEVIR reframes cognitive biases as systematic failures in applying epistemic virtues and demonstrates how different moral foundations lead agents to construct separate, internally coherent "trust lattices". Through case studies on vaccination mandates and climate policy, the framework shows that political polarization represents deeper divergence in moral priors, epistemic authorities, and evaluative heuristics. The report analyzes how propaganda, psychological operations, and echo chambers exploit the MEVIR process. The framework provides foundation for a Decision Support System to augment metacognition, helping individuals identify biases and practice epistemic virtues. The report concludes by acknowledging limitations and proposing longitudinal studies for future research.

cs.CY