arXiv ScienceSearch

arXiv subjects

Nizam Kadir

Publications and source records attributed to Nizam Kadir.

8 recordsLinked to original sources

Beyond Prompt-to-App: Accountable Translation in Teacher-Facing Agentic Authoring

Natural-language app builders let domain experts create software, but their pipelines transform professional intent across compilation, generation, checking, and approval. We report a bounded trace study of a teacher-facing agentic authoring system. Evidence comprises six eligible build attempts across three accounts; a separate corpus of 37 workshop units from 23 display names contextualizes commitments without person-level linkage. Compiled specifications added governance requirements, while downstream representations sometimes normalized case-specific learning relations. Two drafts met a stored package/security threshold despite analyzer reservations and unresolved correspondence to their briefs; four attempts in one account produced no usable payload, and repair messages did not translate internal terms into domain-legible revisions. We develop accountable translation as an analytic framework for making consequential changes attributable, inspectable, scoped in validation, and contestable. It extends HCI accounts of traceability and end-user debugging by locating professional authority and repair rights across heterogeneous technical and organizational handoffs.

cs.HC

From Misconceptions to Evidence: What Science Teachers Make Visible When Co-Designing Agentic Learning Apps

Science educators increasingly encounter AI tools that generate content, yet disciplinary teaching depends on eliciting learners' models, diagnosing misconceptions, interpreting evidence, and preserving professional judgment. This study asks how science teachers translate such epistemic work into specifications for agentic learning applications. It contributes to the conference theme, "Innovating Pedagogies, Inspiring Minds: Transforming Science Learning," and the Teachers' Professional Learning strand by examining app co-design as a form of pedagogical reasoning. We conducted a bounded qualitative cross-case analysis of four de-identified artifacts produced in a teacher professional-learning workshop: an experimental-design diagnostic, a Kinetic Particle Theory dialogue guide, a chemistry prior-knowledge checker, and a physics application/scaffolding tool. Each artifact was coded for the disciplinary problem, learner interaction, evidence made visible, teacher authority, and safeguard. All four connected a science-learning problem to an interaction and pedagogically interpretable evidence: misconceptions and gaps, explanations-in-progress, class-level readiness patterns, or investigation performance. However, only two made teacher control or evaluation explicit, and only two named a safeguard. The proposals therefore positioned AI less as an answer generator than as an elicitor, scaffold, and evidence-return mechanism, while leaving decision rights and protections unevenly specified. We argue that teacher professional learning should treat AI app ideation as epistemic specification work. A five-question design protocol--problem, learner interaction, evidence, teacher authority, and safeguard--can help teachers transform science-learning needs into accountable human-AI arrangements before building or adopting a tool.

cs.HC

RankCert: When Can Simulated Learners Safely Select an AI Tutor? Robust Decision Certification Under Structural Uncertainty

Simulation-based tutor selection can be unstable when predictively adequate learner models imply different policy rankings. RankCert certifies one of eight equal-budget tutoring policies only when model-averaged utility, probability-best, posterior regret, cross-domain rank, family coverage, and leave-one-domain-out and leave-one-visible-family-out averages support the same candidate; otherwise it abstains. We evaluated RankCert in 1,280 frozen held-out settings spanning five rotating held-out oracle families, 64 scenarios per family, and four cohort sizes. Calibration used a licensed, de-identified EdNet-KT1 derivative with 5,000 learners and 590,056 retained responses; all five family representatives passed the frozen adequacy gate. Minimum-domain mean pairwise top-1 agreement was 0.272917 (95% CI [0.253646, 0.293229]), showing substantial structural disagreement. Cohort-noise variance decreased from n = 30 to n = 300, while the structural family share remained nonzero. RankCert reduced total held-out decision loss relative to full-coverage point selection by 0.006605 normalized-outcome units (95% CI [0.004859, 0.008407]). At comparable coverage, however, it did not reduce selective risk relative to a confidence-gated point certificate (difference -0.000213; 95% CI [-0.003238, 0.002384]; Holm p = 0.929654). Certification occurred in 3.75% of settings and only in stable scenarios; RankCert abstained in every ambiguous, misspecified, and structural-conflict setting. "Safe" denotes only benchmark-scoped decision certification under the declared utility and uncertainty set; no human-learning, causal, deployment-effectiveness, or general-safety claim is made.

cs.AI

Auditable Release Control for Pedagogical Leakage in LLM Tutors

Large language model tutors can be correct and helpful yet disclose an answer or decisive reasoning before that disclosure is authorized. We formalize this state- and action-dependent failure as pedagogical leakage and introduce an authorization-aware complete-mediation boundary. A selector emits one of five disclosure contracts, trusted policy gates privileged modes, and a renderer proposes language. A single release function applies inspectable checks, optional cumulative verification, and action-specific fallback; replayable traces separate selection, generation, verification, and enforcement failures. Matched component attribution exposes a safety-utility frontier. On 599 fixed Gemini 3.5 proposals, strict mediation reduces blinded three-model panel-majority leakage flags from 181 to 0 (paired problem-cluster difference -30.22 points, 95% CI [-35.00,-25.72]), while replacing 581 responses and lowering helpfulness. Checker-triggered fallback alone yields 11 majority flags; adding the semantic verifier yields 14 and no reliable marginal gain. A global A1 scaffold yields 0 majority and 54 any-judge flags, outperforming fitted Q on automatic safety and utility. In an externally timestamped replication over 40 unseen problem clusters and 480 attack sequences, high-assurance release reduces majority flags from 42 to 8 (-7.08 points, 95% CI [-13.13,-2.29]); seven failures persist, one is introduced, and mean helpfulness falls by .192. These results establish an auditable release boundary and failure attribution under declared contracts, not universal semantic safety or learning gains.

cs.CR

EduPluginBench: Executable Assurance for AI-Generated Educational Plugins

Code-generation models can produce executable components, but compilation and functional tests do not establish compliance with least privilege, telemetry consent, provenance, privileged-write authority, lifecycle constraints, or bounded failure. We introduce EduPluginBench, an executable benchmark and staged admission method for generated plugins in governed software ecosystems. Across 1,440 activation-checked first-order mutants from 30 specifications, P0-P4 increased release-blocking-defect recall by 74.7 percentage points (specification-clustered 95% CI 73.4-75.8) over P0-P2, with no observed rejection among 120 clean references (95% Wilson upper bound 3.1%). A frozen transfer study of 600 unmodified generations from two current coding models found that 300/600 parsed, but none passed P0 or achieved P0-P4 conformance (95% upper bound 0.64%); downstream assurance estimands were undefined. An independently labelled Moodle study retained 16 vulnerable/fixed pairs; the frozen generic PHP detector found no vulnerable revisions. These negative transfer results prevent controlled contract consistency from being read as independent real-defect effectiveness. An earlier 540-generation diagnostic found that post-hoc bounded repair yielded 112 P0 passes, all nonconforming, with recall increasing from 13.4% to 100%. The artifact retains protocols, public-source provenance, raw generations, row-level decisions, audits, analysis code, and reproduction instructions.

cs.SE

Critsly: An Artefact-Aware AI Critique Teammate for Design Education and Project-Based Learning

Critique is central to design education and project-based learning, yet high-quality critique is often scarce, uneven, hard to document, and disconnected from evolving artefacts. We present Critsly, an artefact-aware AI critique workspace that turns AI from a detached feedback tool into a critique teammate in the learner's board context. Critsly combines a visual design canvas with structured AI-supported reflection, multi-perspective critique, action planning, optional peer/jury settings, and educator evidence traces. Unlike chatbot feedback tools that rely on isolated text prompts, Critsly grounds critique in a structured board state containing design intentions, board elements, annotations, links, and prior critique history. Reflecture, Critsly's guided reflection flow, works with Six Thinking Hats-inspired personas, board synthesis, generated action plans, exportable critique records, and educator evidence views in one workflow. The demo follows a learner from design intention to board-aware critique, persona-based evaluation, action-plan generation, and educator-facing evidence. Critsly contributes a working example of AI-supported critique that is more frequent, structured, inspectable, and actionable for learners and educators.

cs.CY

From Tools to Teacher-Built Teammates: No-Code Pedagogical Plugin Authoring with LearnAdapt Agentic Studio and PedOS 1.1 Lumina

Teachers and researchers need to adapt educational AI to local goals, but most systems remain difficult to customize or study without coding expertise. We present LearnAdapt Agentic Studio on PedOS 1.1 Lumina, a no-code authoring and governed runtime environment for educational AI plugins. A non-coder describes a desired learning interaction in plain English; the system prepares a previewable plugin artifact, runs safety checks, and supports submission for review. PedOS then deploys approved plugins into a directory for installation. Crucially, telemetry is strictly gated to authenticated users running approved plugins. The demo shows the complete lifecycle from prompt to governed evidence capture, shifting from fixed tools to teacher-built teammates.

cs.CY

From Untamed Black Box to Interpretable Pedagogical Orchestration: The Ensemble of Specialized LLMs Architecture for Adaptive Tutoring

Monolithic Large Language Models (LLMs) used in educational dialogue often behave as "black boxes," where pedagogical decisions are implicit and difficult to audit, frequently violating instructional constraints by providing answers too early. We introduce the Ensemble of Specialized LLMS (ES-LLMS) architecture that separates decision-making from wording. Pedagogical actions are selected by a deterministic rules-based orchestrator coordinating specialized agents covering tutoring, assessment, feedback, scaffolding, motivation and ethics-guided by an interpretable Bayesian Knowledge Tracing (BKT) student model. An LLM renderer surface-realizes the chosen action in natural language. This design emphasizes reliability and controllability: constraints such as "attempt-before-hint" and hint caps are enforced as explicit rules, and the system logs per-turn agent traces and constraint checks. Validation of pedagogical quality via human expert reviewers (N=6) and a multi-LLM-as-Judge panel (six state-of-the-art models) showed that ES-LLMs were preferred in 91.7% and 79.2% of cases, respectively. The architecture significantly outperformed monolithic baselines across all seven dimensions, particularly in Scaffolding & Guidance, and Trust & Explainability. Furthermore, a Monte Carlo simulation (N=2,400) exposed a "Mastery Gain Paradox," where monolithic tutors inflated short-term performance through over-assistance. In contrast, ES-LLMs achieved 100% adherence to pedagogical constraints (e.g., attempt-before-hint) and a 3.3x increase in hint efficiency. Operationally, ES-LLMs reduced costs by 54% and latency by 22% by utilizing stateless prompts. We conclude that structural decoupling is essential for transforming stochastic models into trustworthy, verifiable and resource-efficient pedagogical agents.

cs.CY