arXiv · 2609.13202
Do Tabular Foundation Models Still Need Feature Engineering?
Abstract
Feature engineering has long been a cornerstone of tabular machine learning. Tabular foundation models (TFMs) are pretrained on a wide range of tabular datasets and applied via in-context learning. Their rise raises a natural question: does manual feature construction still matter as these models become more capable? To answer this, we perform a controlled study across several versions of two major TFM families, testing a wide range of existing feature engineering techniques on benchmark datasets from TabArena. We find a consistent pattern: feature engineering gains are concentrated in earlier model generations and become negligible for the strongest models. These results suggest that stronger TFMs depend less on explicitly engineered input representations. In a complementary experiment, however, adding in-context information from related datasets still improves performance. Our findings indicate a shift in the source of performance gains for stronger TFMs: re-representing existing inputs becomes less effective, while providing additional task-relevant context remains beneficial.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yifan WU, Pinjun Dong, Jiran Tao, Binyan Jiang. 2026-08-14. Do Tabular Foundation Models Still Need Feature Engineering?. https://arxiv.org/abs/2609.13202
Cite the original work for its findings. Save a collection to share your selection of sources.