arXiv ScienceSearch

arXiv · 2509.18020

ClassMind: Scaling Classroom Observation and Instructional Feedback with Multimodal AI

Abstract

Classroom observation -- one of the most effective methods for teacher development -- remains limited due to high costs and a shortage of expert coaches. We present ClassMind, an AI-driven classroom observation system that integrates generative AI and multimodal learning to analyze classroom artifacts (e.g., class recordings) and deliver timely, personalized feedback aligned with pedagogical practices. At its core is AVA-Align, an agent framework that analyzes long classroom video recordings to generate temporally precise, best-practice-aligned feedback to support teacher reflection and improvement. Our three-phase study involved participatory co-design with educators, development of a full-stack system, and field testing with teachers at different stages of practice. Teachers highlighted the system's usefulness, ease of use, and novelty, while also raising concerns about privacy and the role of human judgment, motivating deeper exploration of future human--AI coaching partnerships. This work illustrates how multimodal AI can scale expert coaching and advance teacher development.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ao Qu, Yuxi Wen, Jiayi Zhang, Yunge Wen, Yibo Zhao, Alok Prakash, Andrés F. Salazar-Gómez, Paul Pu Liang, Jinhua Zhao. 2025-09-22. ClassMind: Scaling Classroom Observation and Instructional Feedback with Multimodal AI. https://arxiv.org/abs/2509.18020

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Comparing Human Oversight Strategies for Computer-Use Agents

LLM-powered computer-use agents (CUAs) shift users from direct manipulation to supervising an unfolding action sequence, yet existing oversight research mostly evaluates static decisions rather than real-time oversight across an agent trajectory. We compared four oversight strategies with 48 participants across 192 live web sessions. Quantitative results show that oversight strategy affected exposure to problematic actions, but not users' ability to intervene. Behavioral analysis shows why: control often returned too early or too late, creating a timing mismatch within the action sequence. Even when control arrived at the right moment, users judged whether the agent acted correctly, not whether the action was safe. Front-loading oversight into a plan did not solve this either: plans constrained unplanned actions but left unlisted safeguards unaddressed, while upfront approval inhibited runtime scrutiny. These findings show that oversight depends less on maximizing control than on aligning authority, timing, and attention with decision-critical moments.

cs.HC

Regimes of Scale in AI Meteorology

HCI work has explored the effective integration of AI/ML tools across application domains from healthcare to finance to transportation. We add to this literature with an analysis of AI/ML tools in meteorology, a domain that already uses big data and massive physics-based models. Drawing from 18 interviews with forecasters and meteorologists with varied connections to AI/ML weather modeling, we trace tensions in AI/ML weather application arising from what we call regimes of scale, different ways that AI/ML and meteorological systems make observations, data, and models scale. Rather than seeing AI/ML as a domain-agnostic tool, we argue that AI/ML methods were born from specific platform and internet infrastructures, and so they can struggle to integrate with very different (in this case meteorological) ways of organizing data pipelines.

cs.HC

"Am I Just Dumb?": Applicability, Action and Verification in Consumer IoT Security Advice

Public campaigns urge people to update their Internet of Things (IoT) devices and change default passwords. What happens when people try? We gave 28 participants in the Netherlands two pieces of government-issued advice and asked them to try applying each to three of six bestselling IoT devices (168 sessions). We located no manufacturer-set password shared across units, the kind the advice describes; the only device-level credential located was unique to its unit. Fewer than half the update sessions established firmware status. Told that a setting might not apply, no participant concluded it did not: they treated whatever related setting the interface offered as the target, and located the difficulty in themselves rather than in the advice or device. Generic advice asks people to judge what only manufacturers can state and only devices can report. Campaigns must be coordinated with device design, or replaced by secure defaults that remove the task.

cs.HC