arXiv · 2501.14717
What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects
Abstract
Table modeling has progressed for decades. In this work, we revisit this trajectory and highlight emerging challenges in the LLM era, particularly the paradox of choice: the difficulty of attributing performance gains amid diverse base models and training sets in the context of table instruction tuning. We replicate four table LLMs by instruction-tuning three foundation models on four existing datasets, yielding 12 models. We then evaluate these models across 16 table benchmarks. Our study is the first to quantitatively disentangle the effects of training data and base model selection, revealing that base model choice plays a more dominant role than the training data itself. Generalization and reasoning remain challenging, inviting future effort on table modeling. Based on our findings, we share our thoughts on the future directions for table modeling.
Explore related subjects
Keep this discovery
Naihao Deng, Sheng Zhang, Henghui Zhu, Shuaichen Chang, Jiani Zhang, Alexander Hanbo Li, Chung-Wei Hang, Hideo Kobayashi, Yiqun Hu, Patrick Ng. 2025-01-24. What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects. https://arxiv.org/abs/2501.14717
Cite the original work for its findings. Save a collection to share your selection of sources.