arXiv ScienceSearch

arXiv · 2506.19662

Multimodal large language models and physics visual tasks: comparative analysis of performance and costs

Abstract

Multimodal large language models (MLLMs) capable of processing both text and visual inputs are increasingly being explored for uses in physics education, such as tutoring, formative assessment, and grading. This study evaluates a range of publicly available MLLMs on a set of standardized, image-based physics research-based conceptual assessments (concept inventories). We benchmark 15 models from three major providers (Anthropic, Google, and OpenAI) across 102 physics items, focusing on two main questions: (1) How well do these models perform on conceptual physics tasks involving visual representations? and (2) What are the financial costs associated with their use? The results show high variability in both performance and cost. The performance of the tested models ranges from 81.5% to as low as 21%. We also found that expensive models do not always outperform cheaper ones and that, depending on the demands of the context, cheaper models may be sufficiently capable for some tasks. This is especially relevant in contexts where financial resources are limited or for large-scale educational implementation of MLLMs. By providing these analyses, our aim is to inform teachers, institutions, and other educational stakeholders so that they can make evidence-based decisions about the selection of models for use in AI-supported physics education.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Giulia Polverini, Bor Gregorcic. 2025-09-11. Multimodal large language models and physics visual tasks: comparative analysis of performance and costs. https://doi.org/10.1088/1361-6404%2Fae03f8

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Force Concept Inventory Across Continents: Testing Q-Matrix Transferability and Cross-Cultural Differences in Mechanics Reasoning

The Force Concept Inventory (FCI) is one of the most widely used research-based assessments in physics education, yet the assumption that its underlying cognitive structure is transferable across educational contexts remains largely untested. This study investigates the transferability of FCI Q-matrices using the Generalized Deterministic Inputs, Noisy "And" Gate (G-DINA) cognitive diagnostic applied to two large cohorts: students from the Learning About STEM Student Outcomes (LASSO) online system database in the United States (N = 4,750) and introductory physics students at the University of Johannesburg, South Africa (N = 1,016). Rather than treating the analysis as a local model calibration exercise, we frame the problem as one of cross-context cognitive invariance. Differential Item Functioning (DIF) analyses revealed substantial cross-cultural differences, with 14 of 30 items exhibiting high DIF after controlling for latent skill mastery. These differences were concentrated in force dynamics and contact-force reasoning and remained invariant under alternative Q-matrix specifications. The findings suggest that observed differences reflect genuine variations in students' conceptual reasoning rather than psychometric artifacts, highlighting the importance of validating Q-matrix structures before deploying cognitive diagnostic and adaptive assessments across diverse non-local educational settings.

physics.ed-ph

Inference uncertainty about an aircraft crash

Problem-based learning benefits from situations taken from real life, which usually stimulate student interest. In this paper, we examine the shooting down of the Rwandan president's aircraft on April 6th, 1994. We discuss methods to infer information about the location from which the missile was launched, its trajectory and type, where the aircraft was struck and its trajectory during the fall. To this end, we developed a physics-based analysis based on expert reports, witness statements, and other publicly available information, as interpreted by our calculations. The analysis is designed to be understandable using undergraduate-level physics. The uncertainty of each result is discussed and propagated to ensure a proper assessment of the hypotheses and a traceability of their consequences. Such approach encourages the students to exercise their critical mind and teaches inference methods that are routinely used in physics research. In addition, it illustrates the importance and limits of scientific expertise during a judiciary process.

physics.ed-ph

Skepticism vs. Convenience: Physics Students' Perceptions and Use of Large Language Models Before and After Instruction

The recent emergence of large language models (LLMs) has produced research focusing on the ability of these tools to solve physics problems, evaluate student work, or otherwise impact the problem-solving process of students. However, studies exploring how physics students perceive LLMs (in terms of capabilities, educational impacts, usage, and role in problem solving) remain limited. This study evaluates the first-year physics students' perceptions toward LLMs and further explores how these perceptions change after practicing problem-solving with and without LLMs and engaging in a reflective lesson on the functioning and educational impacts of LLMs. We find that student opinions toward LLMs vary, with generally favorable perceptions of their capabilities but greater skepticism regarding their value for learning. Despite this skepticism, a majority of students self-report regularly using LLMs to obtain help, commonly reporting deadlines and convenience as motivating factors. Following the lesson, students expressed greater skepticism toward LLMs in several areas, with 88% of students believing that LLMs can leave them with a false sense of confidence about their understanding, up from 58% before the lesson.

physics.ed-ph