arXiv · 2610.02347
From Behavior to Provenance: Attributing Tabular Foundation Models to Synthetic Pretraining Data
Abstract
Training-data attribution aims to identify which training examples shape model behavior, yet validating such claims is difficult because causal training influence is rarely observable. We argue that controlled synthetic pretraining makes attribution experimentally testable. Using O'PRIOR, a provenance-rich synthetic task generator for tabular foundation models, we construct a testbed in which every pretraining task carries explicit lineage over structural mechanisms, missingness, confounding, shortcuts, and distribution shift. We combine behavior-conditioned attribution with counterfactual retraining and provenance-aware interventions to test both task-level faithfulness and mechanism-level consistency. On held-out real tasks, removing the top-attributed 5% of synthetic tasks decreases mean ROC-AUC by 0.013, compared with 0.002$\pm$0.004 under random removal, while removing bottom-attributed tasks improves performance by 0.003. Within shortcut-provenance tasks, targeted removal yields an effect of 0.043 versus 0.016 for matched random removal. Provenance discrimination is more modest by ranking AUROC (0.55-0.62), despite substantial top-k enrichment, revealing that provenance association and interventional faithfulness need not coincide. Our results establish synthetic provenance as a controlled setting for verifiable contributive attribution
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mohamed Bouadi, Nassim Bouarour, Shivam Dubey, Aditya Tanna, Vinay Kumar Sankarapu. 2026-10-01. From Behavior to Provenance: Attributing Tabular Foundation Models to Synthetic Pretraining Data. https://arxiv.org/abs/2610.02347
Cite the original work for its findings. Save a collection to share your selection of sources.