arXiv ScienceSearch

arXiv subjects

Chunyi Zhao

Publications and source records attributed to Chunyi Zhao.

2 recordsLinked to original sources

Reliability, validity, and diagnostic evidence for multi-model LLM short-answer scoring

Large language models (LLMs) are increasingly used or proposed for educational scoring, but single-model and single-run evaluations provide limited evidence for assessment use. Short-answer scoring requires evidence about reliability, validity, severity, diagnostic value, and failure cases. This study evaluated repeated multi-model OCG-PRES guided LLM scoring for short-answer assessment. The analysis used 996 SciEntsBank responses. GPT, DeepSeek, and Qianwen each scored every response across three independent runs using five OCG-PRES dimensions: concept coverage, relation accuracy, reasoning completeness, contradiction control, and domain relevance. Scores were evaluated against official binary and five-category labels and compared with non-LLM baselines based on answer length, Jaccard keyword overlap, TF-IDF cosine similarity, and a combined traditional logistic model. Repeated-run reliability was high for all models, with ICC(3,k) = .977 for GPT, .992 for DeepSeek, and .981 for Qianwen. DeepSeek was the most stable across runs. GPT showed the strongest official-label alignment by AUC (.909), while Qianwen was stricter, with higher precision but lower recall under the fixed threshold = 3.0 rule. OCG-PRES scores followed expected diagnostic patterns across five official categories and outperformed all non-LLM baselines in AUC and F1. Repeated multi-model OCG-PRES scoring provides reliability, validity, and diagnostic evidence for LLM-assisted short-answer scoring. The findings support cautious, evidence-based use as a scoring support tool rather than a replacement for human judgement.

cs.CL

Why Organizational Rules Fail AI: O-I-B-A-R and the Externalization of Decision Boundaries

AI systems increasingly enter organizations through policies, procedures, playbooks, prompts, and other explicit representations of work. Yet formal descriptions often differ from situated practice, and captured know-what can omit the contextual know-how experts use when judgments are uncertain. We argue that a recurring class of organizational AI failures arises partly from a knowledge representation problem at the sociotechnical interface: the AI receives the procedure, while the organization operates on the procedure plus negative boundaries, runtime judgments, responsibility assignments, and learning history. We introduce O-I-B-A-R (OPEN, IS, BUT, ACTION, RESULT), a scaffold for externalizing these missing decision boundaries. IS records when a judgment holds. BUT records a concrete failure containing information beyond the logical negation of IS. Comparable success and failure cases are decomposed toward a minimally sufficient changing variable, which becomes a value-bearing decision dimension. A suspension represents the state in which the dimension is known but its current value is unresolved, specifying what must be measured, asked, retrieved, or escalated to a human. RESULT confirms a boundary, shifts a threshold, or exposes a new dimension. Incidents can generate new dimensions, unresolved values can define human-AI handoffs, and feedback can expand the decision space. We also identify a sociotechnical tension: durable and attributable failure histories can suppress the candor on which useful boundary knowledge depends. Externalization must therefore be designed as an organizational intervention with real costs and incentives.

cs.CY