arXiv ScienceSearch

arXiv · 2609.14789

Evaluating AI Tutoring at the Speed of Innovation: Practitioner-Led Micro-Randomised Trials of an AI Tutoring Platform in GCSE Science

Abstract

Artificial intelligence (AI) systems in education are developing on timescales that sit uneasily with conventional evaluation. By the time a large-scale trial has been designed, delivered, analysed and published, the technology under study may have changed materially. This creates a temporal problem for evidence-informed education: the need for timely evidence can encourage reliance on weak observational or usage data, while conventional rigorous evaluation may produce evidence too slowly to guide rapidly evolving practice. We examine teacher-led micro-randomised controlled trials (micro-RCTs) as one response to this problem. The empirical case is a four-week multisite individually randomised evaluation of Medly, an AI-powered tutoring platform, in GCSE Biology, Chemistry and Physics in English secondary schools. Of 929 students completing baseline assessment, 644 completed post-testing. In the primary ITT analysis, students allocated to Medly achieved higher post-test attainment than students undertaking business-as-usual self-directed revision (Hedges' g = 0.33, 95% CI 0.18 to 0.48). Positive estimates were observed in Physics (g = 0.31), Chemistry (g = 0.32) and Biology (g = 0.52), with no evidence of differential impact by disadvantage status. Greater platform engagement was associated with higher attainment, but these post-randomisation analyses are treated as exploratory rather than causal. Attrition was substantial (30.7%), outcome measures were curriculum-aligned rather than standardised, and process evaluation response was limited. We therefore interpret the findings as preliminary. We argue that the value of micro-RCTs for educational AI lies not in replacing definitive evaluation with small studies, but in enabling a rapid, cumulative evaluation architecture in which randomised estimates can be generated, replicated and updated as technologies and their implementation evolve.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wayne Harrison, Rahil Khowaja, Emma Dobson, Germaine Uwimpuhwe, Steve Higgins. 2026-09-13. Evaluating AI Tutoring at the Speed of Innovation: Practitioner-Led Micro-Randomised Trials of an AI Tutoring Platform in GCSE Science. https://arxiv.org/abs/2609.14789

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

ADSEL: Adaptive Dual Self-Expression Learning for EEG Feature Selection via Incomplete Multi-Dimensional Emotion Labels

EEG based multi-dimension emotion recognition has attracted substantial research interest in affective computing. However, the high dimensionality of EEG features, coupled with limited sample sizes, frequently leads to classifier overfitting and high computational complexity. Feature selection constitutes a critical strategy for mitigating these challenges. However, most existing EEG feature selection methods assume complete multi-dimensional emotion labels. In practice, open acquisition environment and the inherent subjectivity of emotion perception often result in incomplete label data, which can compromise model generalization. Additionally, existing feature selection methods for handling incomplete multi-dimensional labels primarily focus on correlations among various dimensions during label recovery, neglecting the correlation between samples in the label space and their interaction with various dimensions. To address these issues, we propose a novel incomplete multi-dimensional emotion feature selection framework integrating Adaptive Dual Self-Expression Learning (ADSEL) with least squares regression. ADSEL could establish a bidirectional pathway between sample-level and dimension-level self-expression learning processes within the label space. It could facilitate the cross-sharing of learned information between these processes, enabling the simultaneous exploitation of effective information across both samples and dimensions for label reconstruction. Consequently, ADSEL could enhance label recovery accuracy and effectively identifies the optimal EEG feature subset for multi-dimensional emotion recognition. ADSEL was evaluated against fourteen state-of-the-art feature selection methods on three public EEG datasets with multi-dimensional emotion labels. Experimental results demonstrate that ADSEL could achieve superior performance under conditions of partial label absence.

cs.HC

A Human-AI Collaborative Workflow for Mathematical Discovery: A Case Study in Grover-Compatible Riemannian Optimization

We investigate how large language models can be used as research tools in scientific computing while preserving mathematical rigor. We propose a human-in-the-loop workflow for interactive theorem proving and discovery with LLMs. Human experts retain control over problem formulation and assumptions, while the model searches for proofs or contradictions, proposes candidate properties and theorems, and helps construct structures and parameters that satisfy explicit constraints, supported by numerical experiments and simple verification checks. Experts treat these outputs as raw material, further refine them, and organize the results into precise statements and rigorous proofs. We instantiate this workflow in a main case study on the connection between manifold optimization and Grover's quantum search algorithm, where the pipeline identifies invariant subspaces and explores Grover-compatible retractions. The main case study uses the corresponding Grover-compatible convergence analysis, including an $O(\sqrt{N} \log(1/\varepsilon))$ PL-based bound established in the companion mathematical work, to illustrate the refinement stage of the workflow. Prompt records and reusable templates for implementing the workflow are provided. We further include a multi-oracle case study, document representative failed and corrected routes arising from this setting, and provide a structured failure-mode analysis.

cs.HC

Learning Password Best Practices Through In-Task Instruction

Users often make security- and privacy-relevant decisions without a clear understanding of the rules that govern safe behavior. We introduce pedagogical friction, a design approach that inserts brief, instructional interactions at the moment of action. We evaluate this approach in the context of password creation, a familiar task with clear quality criteria. We conducted a randomized study with 128 participants across four interface conditions that varied the depth and interactivity of guidance. We assessed three outcomes: (1) rule compliance in a subsequent password task without guidance, (2) accuracy on survey questions tied to password rules, and (3) behavior-knowledge alignment, which captures whether participants who correctly followed a rule also recognized it on the survey. Across the guided conditions, participants corrected most rule violations in the follow-up task and showed high behavior-knowledge alignment. Survey results suggested clearer advantages for some rule types, especially symbol related questions. These results position pedagogical friction as a lightweight intervention for security- and privacy-critical interfaces.

cs.HC