arXiv · 2609.32160
Typed Decision Models: An Early Evidence Audit and Evaluation Checklist
Abstract
Typed decision models (TDMs) return probability distributions over caller-defined options without generating text. TypeSafe released Jev, a commercial typed decision model, on 15 September 2026, and a small body of evaluation and replication work appeared within days. We review 28 papers posted between 19 and 24 September and relate their findings to earlier work on label-probability classification, constrained decoding, reranking, calibration, and model cascades. In this early literature, the typed readout itself has not shown an independent accuracy advantage over comparable label-probability readouts. Jev's clearest gains are in latency and cost, while accuracy gaps remain on harder tasks. In practical deployments, confidence is often used to decide when to defer to a stronger model or a human. We use recurring weaknesses in these studies to derive a 14-item evaluation checklist for future TDM work. Because the evidence covers only the first nine days after the release of one hosted model, the review should be read as an early evidence map rather than a settled assessment of the model class.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lijuan Tang, Yuemeng Zheng. 2026-09-26. Typed Decision Models: An Early Evidence Audit and Evaluation Checklist. https://arxiv.org/abs/2609.32160
Cite the original work for its findings. Save a collection to share your selection of sources.