arXiv · 2609.33874
JET: Justification Evaluation in Transformer
Abstract
JET uses pretrained language and vision-language models to select among a finite set of answers without additional training. It evaluates candidate likelihoods directly and shares computation across candidates. Experiments on desktop CPUs and consumer GPUs assess decision accuracy and execution cost. Qwen3.6-35B-A3B achieves 87.48% accuracy on the full MMLU test set and 3.69 requests per second on a separately timed MMLU subset. The accuracy-throughput comparison covers model, hardware, and reasoning choices, with Jev as an external reference. Controlled execution experiments show 2.18-2.23-fold speedups from prefix reuse and cache management, and a 30.8% reduction in process time from input preparation optimizations, with unchanged outputs. Optional reasoning has a task-dependent accuracy-throughput trade-off. These results support local decision inference from existing models.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shenghao Ding. 2026-09-27. JET: Justification Evaluation in Transformer. https://arxiv.org/abs/2609.33874
Cite the original work for its findings. Save a collection to share your selection of sources.