arXiv · 2610.04430
MERCI Cards: An LLM Evaluation and Deployment Framework for High-Stakes Domains
Abstract
As LLMs are increasingly deployed in high-stakes professional workflows, engineers and researchers require principled protocols to systematically track, monitor, and improve model performance across deployment cycles. We present a mathematical framework for iterative LLM evaluation and deployment, and demonstrate its application to AI systems used in criminal justice. Our framework formalizes LLM integration in high-stakes, high-risk, and resource-constrained domains across model selection, rubric design, evaluations and deployment via a weighted multi-objective optimization. We demonstrate that MERCI Cards can guide improvements of the system across deployment iterations, direct developer attention toward under-performing areas, and focus user attention on validation and error-correction in the LLM's outputs.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Aparna Komarla, Annalisa Szymanski. 2026-10-03. MERCI Cards: An LLM Evaluation and Deployment Framework for High-Stakes Domains. https://arxiv.org/abs/2610.04430
Cite the original work for its findings. Save a collection to share your selection of sources.