arXiv ScienceSearch

arXiv · 2608.07373

A bottom-up taxonomy of student discourse with a Socratic AI physics tutor

Abstract

Large language model (LLM) tutors are being deployed in introductory physics courses at a scale that produces transcript corpora far larger than traditional qualitative coding can absorb. A central question for physics education research (PER) is empirical and prior to any claim about effectiveness: what do students actually say to these tutors? We address this question for one Socratic AI tutor deployed in an introductory calculus-based mechanics course by building a bottom-up taxonomy of student discourse. Each student turn is assigned an emergent free-text label by an LLM coder using the surrounding conversational context; near-paraphrase labels are then consolidated into a smaller set of discourse categories using a similarity-based grouping procedure. The procedure is validated against a stratified human-coded sample. The resulting taxonomy of 357 categories is strikingly concentrated: the top 25 categories cover roughly half of all student turns, and two thematic bands: equation-handling and meta-procedural requests together dominate the head of the distribution. The substantive contribution is the taxonomy itself: a description of the discourse PER researchers can expect to encounter when students work with an AI tutor of this design, including a striking prevalence of meta-procedural turns in which students cede strategic control to the tutor

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Syed Furqan Abbas Hashmi, N. Sanjay Rebello. 2026-08-07. A bottom-up taxonomy of student discourse with a Socratic AI physics tutor. https://arxiv.org/abs/2608.07373

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Data-driven modeling in the introductory physics laboratory: Scaling analysis and data collapse in the specific heat of water experiment

In introductory physics laboratories, a central instructional goal is to help students construct and evaluate mathematical models from empirical data rather than applying given formulas. We present a data-driven redesign of the classic specific heat of water experiment that emphasizes scaling analysis and data collapse as tools for model construction. The activity combines structured experimental and analytical guidance with instructor-mediated questioning, while thermodynamic theory is deliberately postponed. Students collect temperature-time data under various experimental conditions, producing multiple data sets that initially appear unrelated. Through successive rescaling, students reduce the dimensionality of the variable space and achieve data collapse onto a single master curve, from which they formulate an empirical model relating energy input, mass, and temperature change. The analysis highlights a limitation of multiplicative scaling: the additive contribution of the calorimeter cannot be eliminated, leading to a structural non-identifiability of the subsystem contributions. To clarify the domain of validity of the model, a thermodynamic description is introduced a posteriori as a boundary-setting framework for interpreting the empirical model. In this sense, the central contribution of this work is to use data-driven modeling both to construct models and to reveal their intrinsic limitations. The experiment provides an accessible example of how scaling, data collapse, and theoretical reasoning can be integrated in an introductory laboratory.

physics.ed-ph

Missing Data on Physics Exams: Demographic Patterns, Course-Level Predictions, and Implications for Equity

In a previous quantitative retrospective study we showed that different demographic groups of students leave different numbers of problems blank on physics exams, leading to inequities in course outcomes. In that work we argued that there were good reasons to treat these blanks as missing data, rather than indicators of a lack of understanding. In this paper, we refine this analysis and show more detailed breakdowns of uncollected test item responses by race/ethnicity and first generation college student status, coming to the same conclusion: test item responses are uncollected for students with different ethnic and racial backgrounds at different rates, and these patterns are not exclusive to low-performing students. We also correct an error from our previous work, finding here that there is no significant gender difference in uncollected test item responses. Finally, we provide a more robust analysis of course level data illustrating that blanks are a variable controlled at the course level rather than the student level, providing more evidence for the use of a course deficit model (rather than a student deficit model) when examining equity disparities, and also suggesting that there are plausible means for instructors to minimize uncollected test item responses, and therefore reduce or even eliminate the bias associated with this missing data. We provide a couple suggestions for faculty who want to minimize the impacts of blanks.

physics.ed-ph

A Framework for Characterizing Learning Contributions Across the Initial Achievement Spectrum

Conceptual assessments are widely used in physics education research to evaluate changes in student understanding, yet class average measures can obscure how those changes are distributed across students with different levels of initial achievement. We introduce a Learning Contribution Framework that characterizes this distribution through the Learning Contribution Curve (LCC) and Learning Contribution Profile (LCP). For a general contribution measure G, the LCC represents cumulative contribution across students ranked by initial achievement, whereas the LCP describes the local mean contribution relative to the population mean of the individual G values. We apply the framework to the individual Hake normalized gain and develop a statistical model linking LCC and LCP behavior to the joint structure of pretest and posttest scores. Under the central assumption that the conditional mean of the subsequent score is linear in the initial score, we identify a score structure parameter $β$ and a critical score structure parameter $β_c$. Whether $β$ is greater than, less than, or equal to $β_c$ determines whether the expected LCP increases, decreases, or remains constant across initial achievement. Given the pretest distribution, the model further yields analytical predictions for the complete LCP and LCC that closely reproduce simulation results. Application to classroom concept-assessment data illustrates how the framework reveals local and cumulative patterns of learning contribution that are not evident from an overall class average revealed by the traditional Hake's normalized gain.

physics.ed-ph