arXiv ScienceSearch

arXiv · 2511.00105

Artificial Intelligence in Elementary STEM Education: A Systematic Review of Current Applications and Future Challenges

Abstract

Artificial intelligence (AI) is transforming elementary STEM education, yet evidence remains fragmented. This systematic review synthesizes 258 studies (2020-2025) examining AI applications across eight categories: intelligent tutoring systems (45% of studies), learning analytics (18%), automated assessment (12%), computer vision (8%), educational robotics (7%), multimodal sensing (6%), AI-enhanced extended reality (XR) (4%), and adaptive content generation. The analysis shows that most studies focus on upper elementary grades (65%) and mathematics (38%), with limited cross-disciplinary STEM integration (15%). While conversational AI demonstrates moderate effectiveness (d = 0.45-0.70 where reported), only 34% of studies include standardized effect sizes. Eight major gaps limit real-world impact: fragmented ecosystems, developmental inappropriateness, infrastructure barriers, lack of privacy frameworks, weak STEM integration, equity disparities, teacher marginalization, and narrow assessment scopes. Geographic distribution is also uneven, with 90% of studies originating from North America, East Asia, and Europe. Future directions call for interoperable architectures that support authentic STEM integration, grade-appropriate design, privacy-preserving analytics, and teacher-centered implementations that enhance rather than replace human expertise.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Majid Memari, Krista Ruggles. 2025-11-06. Artificial Intelligence in Elementary STEM Education: A Systematic Review of Current Applications and Future Challenges. https://arxiv.org/abs/2511.00105

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Measuring Human Contribution in AI-Assisted Content Generation

With the growing prevalence of generative artificial intelligence (AI), an increasing amount of content is no longer exclusively generated by humans but by generative AI models with human guidance. This shift presents notable challenges for the delineation of originality due to the varying degrees of human contribution in AI-assisted works. This study raises the research question of measuring human contribution in AI-assisted content generation and introduces a framework to address this question that is grounded in information theory. By calculating mutual information between human input and AI-assisted output relative to self-information of AI-assisted output, we quantify the proportional information contribution of humans in content generation. Our experimental results demonstrate that the proposed measure effectively discriminates between varying degrees of human contribution across multiple creative domains. We hope that this work lays a foundation for measuring human contributions in AI-assisted content generation in the era of generative AI.

cs.CY

Do Agents Repair When Challenged -- or Just Reply? Challenge, Repair, and Public Correction in a Deployed Agent Forum

As large language model (LLM) agents enter public forums, a key question is whether those forums sustain challenge, repair, and public correction, or merely produce norm-like language. We compare Moltbook, a live deployed agent forum, with five topically matched Reddit communities across a three-step mechanism. Relative to Reddit, Moltbook discussions are roughly ten times less threaded, leaving far fewer chances for challenge and response. When challenges do occur, the original author almost never returns (1.2\% vs.\ 40.9\% on Reddit), multi-turn continuation is nearly absent ($<$0.1\% vs.\ 38.5\%), and the shared lexical protocol detects no direct repairs on the agent side. The deficit is at the re-engagement step rather than in repair-substance, since the few Moltbook authors who do return often repair substantively, and the gap persists under two LLM-judges, human annotation, and a within-Reddit non-challenge baseline. Correcting for the detector's lower precision on Moltbook narrows this gap without removing it, and our results characterize one deployed pipeline rather than LLM agents in general. Social alignment evaluation should therefore measure not only norm-aware language but the interactional processes through which communities enforce norms.

cs.CY

Dataset repurposing and disruptive AI research

Technological advancements are enabling increasingly systematic and large-scale data collection across all areas of science, driving scientific innovation. In particular, AI research exemplifies this trend, having advanced rapidly through the assembly of massive datasets used to train and evaluate machine learning models. However, the escalating demand for data, the difficulty of creating high-quality datasets, and the exhaustion of easily accessible data sources in AI research raise important questions about how to maximize the value of existing datasets through recombination and repurposing. Here, we draw on two theoretical frameworks---recombinational novelty and transformational creativity---to examine the practice of data repurposing and its scientific impact. Focusing on AI, we analyze scientific outcomes associated with data repurposing across more than 10,000 machine learning papers. First, we find that although most repurposed datasets do not achieve broad visibility in the short term, data repurposing is associated with greater disruption. Second, when repurposed data is adopted by subsequent research, the repurposing paper is associated with higher disruption and increased citation impact. Third, repurposing teams tend to be more experienced, more institutionally prestigious, and involve academic--industry collaboration. However, team characteristics poorly predict which repurposed datasets will be adopted by the community. These findings suggest that data repurposing may be an important approach to scientific discovery, and that its successful adoption is more common among larger teams and collaborations spanning academia and industry.

cs.CY