Pinning Decisions Before Failure: Executable Records of Underspecified Choices in AI-Assisted Code Generation
A natural-language requirement leaves questions open, and a model asked to implement it settles them silently: across 600 generated test suites from three models, 42.7% contain no test that distinguishes the competing readings. An acceptance example written before implementing is the usual remedy, but an example both readings satisfy resolves nothing. We insert two steps into that practice: enumerate the requirement's underspecified points by name, then for each construct two throwaway implementations differing only in that point and keep a candidate input only if executing both shows they disagree. The result is recorded as a decision pin: the named point, the confirmed input, and the value the person chose between the two exhibited results. The same record then constrains generation and decides compliance by execution. On a benchmark of 40 tasks with paired reference implementations and two models, a separating input is obtained for 92.5% of decision points and a pin identifying the intended decision for 85-90%. As checks on 703 independently generated implementations, pins agree with the benchmark's classification on 96-97%, with disagreements on three tasks, one where both classifiers erred. As generation constraints, pins are honoured at the pinned input in all 210 generations and are never worse than a prose rule on held-out inputs in 38 cells, though no better than prose stating the same scope. Every compliance failure under a prose rule came from the model deciding the rule's scope itself; one such case silently overturned another recorded decision, was attributed to a single rule by leave-one-out on both models, and was missed by text-level reconciliation but caught by re-running the recorded input. The setting yields too few such conflicts to evaluate a regression step, and we say why.