arXiv ScienceSearch

arXiv subjects

Tong Wan

Publications and source records attributed to Tong Wan.

8 recordsLinked to original sources

Scalable Generation and Validation of Isomorphic Physics Problems with GenAI

Traditional synchronous STEM assessments face growing challenges including accessibility barriers, security concerns from resource-sharing platforms, and limited comparability across institutions. We present a framework for generating and evaluating large-scale isomorphic physics problem banks using Generative AI to enable asynchronous, multi-attempt assessments. Isomorphic problems test identical concepts through varied surface features and contexts, providing richer variation than conventional parameterized questions while maintaining consistent difficulty. Our generation framework employs prompt chaining and tool use to achieve precise control over structural variations (numeric values, spatial relations) alongside diverse contextual variations. For pre-deployment validation, we evaluate generated items using 17 open-source language models (LMs) (0.6B-32B) and compare against actual student performance (N>200) across three midterm exams. Results show that 73% of deployed banks achieve statistically homogeneous difficulty, and LMs pattern correlate strongly with student performance (Pearson's $\rho$ up to 0.594). Additionally, LMs successfully identify problematic variants, such as ambiguous problem texts. Model scale also proves critical for effective validation, where extremely small (<4B) and large (>14B) models exhibit floor and ceiling effects respectively, making mid-sized models optimal for detecting difficulty outliers.

cs.CY

Emergent Explicit Regulation in College Students Collaborative Scientific Inquiry Learning, Framework and A Case Study

Small group activities have been widely adopted in college level science courses. As students participate in these activities, it is important to consider how group members collectively regulate their activity and complete group task. Regulation in a group often involves adaptive responsivity from group members when they notice and deal with a challenge. The theoretical framework of socially shared regulation emphasizes group members collaboratively regulating within the group but does not focus on portraying how the shared regulation is developed in the moment. Currently, the field lacks a framework characterizing the momentary development of a regulatory action in a group. Our study addresses this gap. In our video data, incoming college students were enrolled in a summer program designed to promote students metacognitive skills to be incorporated in their study of science. We have observed various moments in which the students spontaneously made a move to regulate the hands-on, inquiry activity in completing their tasks and achieving group goals. We developed a framework called Emergent Explicit Regulation to characterize those moments. The EER framework captures students in the moment regulatory moves to respond to a challenge, articulating how those moves emerge, in what ways they are explicit and regulatory. In this paper, we first introduce the EER framework and situate the EER framework in the context of collaborative scientific inquiry learning. We then present a case study where we applied the EER framework to identify typical EER instances in one small group when the students completed the task of building a model to represent the climate of the Earth atmosphere. They worked collaboratively, faced and handled various challenges, completed the group task, and demonstrated multiple EERs in different psychological areas and in the inquiry practices designed in the activity.

physics.ed-ph

InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by grounding responses with retrieved information. As an emerging paradigm, Agentic RAG further enhances this process by introducing autonomous LLM agents into the information seeking process. However, existing benchmarks fall short in evaluating such systems, as they are confined to a static retrieval environment with a fixed, limited corpus} and simple queries that fail to elicit agentic behavior. Moreover, their evaluation protocols assess information seeking effectiveness by pre-defined gold sets of documents, making them unsuitable for the open-ended and dynamic nature of real-world web environments. To bridge this gap, we present InfoDeepSeek, a new benchmark with challenging questions designed for assessing agentic information seeking in real-world, dynamic web environments. We propose a systematic methodology for constructing challenging queries satisfying the criteria of determinacy, difficulty, and diversity. Based on this, we develop the first evaluation framework tailored to dynamic agentic information seeking, including fine-grained metrics about the accuracy, utility, and compactness of information seeking outcomes. Through extensive experiments across LLMs, search engines, and question types, InfoDeepSeek reveals nuanced agent behaviors and offers actionable insights for future research.

cs.IR

Using Large Language Models to Assign Partial Credit to Students' Explanations of Problem-Solving Process: Grade at Human Level Accuracy with Grading Confidence Index and Personalized Student-facing Feedback

This study examines the feasibility and potential advantages of using large language models, in particular GPT-4o, to perform partial credit grading of large numbers of student written responses to introductory level physics problems. Students were instructed to write down verbal explanations of their reasoning process when solving one conceptual and two numerical calculation problems on in class exams. The explanations were then graded according to a 3-item rubric with each item grades as binary (1 or 0). We first demonstrate that machine grading using GPT-4o with no examples nor reference answer can reliably agree with human graders on 70%-80% of all cases, which is equal to or higher than the level at which two human graders agree with each other. Two methods are essential for achieving this level of accuracy: 1. Adding explanation language to each rubric item that targets the errors of initial machine grading. 2. Running the grading process 5 times and taking the most frequent outcome. Next, we show that the variation in outcomes across 5 machine grading attempts as measured by the Shannon Entropy can serve as a grading confidence index, allowing a human instructor to identify ~40% of all potentially incorrect gradings by reviewing just 10 - 15% of all responses. Finally, we show that it is straightforward to use GPT-4o to write clear explanations of the partial credit grading outcomes. Those explanations can be used as feedback for students, which will allow students to understand their grades and raise different opinions when necessary. Almost all feedback messages generated were rated 3 or above on a 5-point scale by two experienced instructors. The entire grading and feedback generating process cost roughly $5 per 100 student answers, which shows immense promise for automating labor-intensive grading process by a combination of machine grading with human input and supervision.

physics.ed-ph

Achieving Human Level Partial Credit Grading of Written Responses to Physics Conceptual Question using GPT-3.5 with Only Prompt Engineering

Large language modules (LLMs) have great potential for auto-grading student written responses to physics problems due to their capacity to process and generate natural language. In this explorative study, we use a prompt engineering technique, which we name "scaffolded chain of thought (COT)", to instruct GPT-3.5 to grade student written responses to a physics conceptual question. Compared to common COT prompting, scaffolded COT prompts GPT-3.5 to explicitly compare student responses to a detailed, well-explained rubric before generating the grading outcome. We show that when compared to human raters, the grading accuracy of GPT-3.5 using scaffolded COT is 20% - 30% higher than conventional COT. The level of agreement between AI and human raters can reach 70% - 80%, comparable to the level between two human raters. This shows promise that an LLM-based AI grader can achieve human-level grading accuracy on a physics conceptual problem using prompt engineering techniques alone.

physics.ed-ph

Using skateboarding to develop a culturally relevant tutorial on static equilibrium

Culturally relevant pedagogy (CRP), initially developed by Ladson-Billings, is an instructional framework for supporting diverse learners by drawing on their cultural backgrounds and experiences. In line with the CRP framework, we developed a tutorial on static equilibrium using skateboarding, a popular activity on university campuses, as a culturally relevant context. To help students refine their conceptions about static equilibrium documented in the physics education research (PER) literature, we used the elicit-confront-resolve (ECR) strategy to develop the tutorial. In this paper, we provide a detailed account of how we operationalized the ECR strategy in designing the sequences of questions in the tutorial. Additionally, we present anecdotal evidence to show that this research-based culturally relevant tutorial appears to effectively engage students and motivate their interest in learning physics.

physics.ed-ph

Characterizing Discourse Group Roles in Inquiry-based University Science Labs

Group work is commonly adopted in university science laboratories. However, student small-group discourse in university science labs is rarely investigated. We aim to bridge the gap in the literature by characterizing student discourse group roles in inquiry-based science labs. The instructional context for the study was a summer program hosted at a private research university in the eastern United States. The program was designed as a bridge program for matriculating students who were first generation and/or deaf or hard-of-hearing (DHH). Accommodations such as interpreters and technology were provided for DHH students. We analyzed 19 students' discourse moves in five lab activities from the video recordings, resulting in a total of 48 student-lab units. We developed codes to describe student discourse moves: asking a question, proposing an idea, participating in discussion, chatting off-task, and talking with instructor. Through a cluster analysis using the 48 student-lab units on quantified discourse moves, we identified four discourse styles, High on-task high social, High on-task low social, Low on-task low social, and Low on-task high social. The results show that individual students tend to demonstrate varying discourse styles in different lab activities; students' discourse styles within the same groups tend to be aligned with their group members. By examining group members' discourse styles in mixed-gender groups, we did not observe a difference in engagement level between female and male students. DHH students in mixed hearing ability groups, however, were observed to have a lower level of engagement compared to their non-DHH group members. We discuss possible factors that may have contributed to the observations for genders and students with different hearing abilities. We also provide suggestions for promoting equitable small-group discourse in university science labs.

physics.ed-ph

Exploring Generative AI assisted feedback writing for students' written responses to a physics conceptual question with prompt engineering and few-shot learning

Instructor's feedback plays a critical role in students' development of conceptual understanding and reasoning skills. However, grading student written responses and providing personalized feedback can take a substantial amount of time. In this study, we explore using GPT-3.5 to write feedback to student written responses to conceptual questions with prompt engineering and few-shot learning techniques. In stage one, we used a small portion (n=20) of the student responses on one conceptual question to iteratively train GPT. Four of the responses paired with human-written feedback were included in the prompt as examples for GPT. We tasked GPT to generate feedback to the other 16 responses, and we refined the prompt after several iterations. In stage two, we gave four student researchers the 16 responses as well as two versions of feedback, one written by the authors and the other by GPT. Students were asked to rate the correctness and usefulness of each feedback, and to indicate which one was generated by GPT. The results showed that students tended to rate the feedback by human and GPT equally on correctness, but they all rated the feedback by GPT as more useful. Additionally, the successful rates of identifying GPT's feedback were low, ranging from 0.1 to 0.6. In stage three, we tasked GPT to generate feedback to the rest of the student responses (n=65). The feedback was rated by four instructors based on the extent of modification needed if they were to give the feedback to students. All the instructors rated approximately 70% of the feedback statements needing only minor or no modification. This study demonstrated the feasibility of using Generative AI as an assistant to generating feedback for student written responses with only a relatively small number of examples. An AI assistance can be one of the solutions to substantially reduce time spent on grading student written responses.

physics.ed-ph