Separating Memory and Workflow Effects in Predicting Individual Answers
Language agents choose what to remember about a person and how to use that memory. We separate these choices when predicting a person's unseen answer to an interview question. On 1,768 tasks from 188 people, a concrete memory from a verified interview prefix outscores a trait description by 0.0158 (95% interval [0.0044, 0.0271]). Crossing both memories with one-shot generation and three-answer fusion, fusion lowers concrete-memory scores by 0.0123 ([-0.0189, -0.0056]); prompted and trained selectors do not detectably beat a random candidate, and one call on the longer, unrewritten record outscores every memory condition. At matched context budgets, OwnWords, one call on the person's BM25-ranked sentences, outperforms the written memory on 500 people outside the benchmark (+0.0127, [+0.0037, +0.0217]; an earlier held-out test was inconclusive) and across four budgets on 300 people (mean +0.0218, [+0.0138, +0.0298]), but does not detectably outperform recency truncation; these comparisons do not isolate verbatim wording. On a survey benchmark it predicts ordinal answers more closely than the memory but is not more accurate on exact choices (less accurate in one of two screener samples). Interview scores use a model-based content rubric without human ratings, and all benchmark and some confirmation participants were seen during development.