arXiv Science⌕ Search

arXiv · 2610.12125

The learner who does not learn: when optimizing a pedagogical metric degrades LLM tutoring

Abstract

It is assumed that a natural way to improve the pedagogical quality of large language model tutors is to define a metric of instructional performance and fine-tune the model against it. To test this strategy, we designed a metric of pedagogical adaptivity that scores each instructional decision in a learning sequence against the conditions of the learning situation, which is the standard used for automated pedagogical scoring. We audited a frontier tutor across 2,000 learner scenarios, corrected its weakest cases by fine-tuning an open-weights proxy, and asked 31 trained educators to rate the pedagogical alignment of the outputs blind, before and after correction. The metric increased from +0.05 to +0.42 for the corrected cases, while the expert ratings decreased from 4.46 to 3.03, with the unmodified controls remaining unchanged and a base-proxy control ruling out the change of model. The tutor performed worse because any metric that scores decisions independently and averages them is maximized by repeating the single best decision, and the fine-tuned model collapsed to that exact optimum in every case, in and out of sample. Educators identified the repetition, which such metrics cannot represent, and preserving the learner's trajectory in the score reduced, but did not reverse, the metric's verdict. Weight analysis traced the correction to the model's output projection, where it had memorized its training strings rather than learned to adapt. We conclude that measurement validity does not imply optimization validity, and we derive design principles for benchmarks that assess or train AI-tutor instruction.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Daniel Domínguez Figaredo, Rafael Fernández De la Cruz. 2026-10-08. The learner who does not learn: when optimizing a pedagogical metric degrades LLM tutoring. https://arxiv.org/abs/2610.12125

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Who Does Withholding Delay? A Welfare Model of Open-Weight AI Release Under Asymmetric Proliferation

Withholding a dual-use AI model delays only the actors that lack other routes to a comparable capability. If sophisticated adversaries obtain substitutes faster than distributed defenders, restriction can delay defenders more than the adversaries it targets. We compare controlled access, a defender-first window followed by public release, safeguarded open weights, and minimally restricted open weights in a discounted welfare model with actor-specific substitute acquisition. Under exponential acquisition, restriction gives adversaries a positive discounted access advantage exactly when they substitute faster than defenders, and, with equal usefulness, immediate release adds more expected capability at a fixed horizon to the slower-substituting group. Neither result implies that release is preferable, because opportunistic misuse, defensive reach, safeguard friction, and irreversible losses can reverse the ranking. In a linear benchmark, broad release overtakes control above a unique adversary-substitution threshold whenever such a threshold exists, and we derive the probability that selected defenders deploy before both adversary substitution and public release. In a nonlinear implementation, each of the four policies is optimal somewhere in the parameter space. Three nested 2,048-point designs over thirteen inputs show that policy shares depend strongly on the chosen parameter bounds. Release records and cybersecurity reports illustrate the quantities a release review would need to measure and are kept separate from the calibration.

cs.CY↗

What Personal Information Improves LLM-Based Next-Location Prediction?

Large language models (LLMs) are increasingly used for individual next-location prediction, with personal information easily added to prompts alongside mobility history. Yet the incremental predictive value of such information remains unclear. Using linked sociodemographic records and mobility traces from 5,000 Shenzhen residents, this study separates model responsiveness from predictive value. GPT-5 is the primary model, with GPT-5.5 and Claude Opus 4.6 used for replication. In 1,000 paired prediction instances, models rank 100 candidate destinations with and without age, gender, occupation and income while all other inputs are held fixed. Behavioural history raises top-1 accuracy from 5.6% to 18.5% as prior history increases from zero to six days. By contrast, sociodemographic attributes produce no detectable overall gain, although replacing correct attributes with those of another person reduces accuracy by 5.4 percentage points. Candidate construction also matters, removing distance raises accuracy by 7.7 points under proximity sampling but lowers it by 22.3 points under popularity sampling, with the reversal reproduced across all three LLMs. These findings identify behavioural history as the clearest source of incremental value and show that personal information should be evaluated under matched, explicitly specified conditions before its privacy and governance costs are justified.

cs.CY↗

Political polarization and mental wellbeing: asymmetric evidence for bidirectionality

A growing body of evidence suggests that political polarization and mental health and wellbeing influence each other. Yet the two directions, how polarization shapes mental wellbeing and how mental wellbeing shapes polarization, remain siloed across disciplines. We review both literatures to evaluate the strength and alignment of evidence for bidirectional effects. We find that the literature varies widely in constructs, measures, and levels of analysis, and that the evidence for the two pathways is asymmetric: links from polarization to mental wellbeing are more direct, grounded in well-developed theoretical frameworks and social-relational mechanisms. In contrast, the reverse pathway is substantially less direct, with studies often examining cognitive and socio-emotional factors, or political outcomes adjacent to polarization rather than polarization itself. We identify clear gaps in the literature highlighting the need for future studies with conceptual clarity, unified frameworks, validated measures, and multilevel designs examining both pathways within the same empirical and theoretical setting.

cs.CY↗