arXiv · 2609.37255
Loss-Guided Pretraining Data Selection for Time-Series Foundation Models
Abstract
Time series foundation models (TSFMs) are pretrained on heterogeneous collections containing billions of observations, yet their training windows are typically sampled without estimating whether they provide useful learning signal. We introduce a static data-selection framework that scores each window with a reference forecaster and retains an intermediate interval within every source dataset. Specifically, we connect forecasting loss to optimization difficulty by showing that normalized squared loss controls the per-sample gradient norm under a local Jacobian condition. We then define a reference loss score and apply dataset-stratified selection to preserve the diversity of samples. Across various TSFM architectures, our method outperforms random selection by an absolute margin and even improves both relative MASE and CRPS over full-data pretraining by retaining fewer candidate pretraining windows. Further analyses show strong cross-scale and cross-architecture score correlations, indicating that a small reference model can often select data for larger targets, provided that the reference and target share compatible difficulty orderings.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yike Li, Shaoxu Song, Jianmin Wang. 2026-09-29. Loss-Guided Pretraining Data Selection for Time-Series Foundation Models. https://arxiv.org/abs/2609.37255
Cite the original work for its findings. Save a collection to share your selection of sources.