Learning Transferable Sensor Models via Language-Informed Pretraining
Multimodal language models have demonstrated strong semantic understanding and reasoning over physiological and behavioral signals captured from diverse healthcare sensors. However, existing sensor-language models can only process a fixed number of sensor channels at a fixed sampling rate, making them difficult to adapt to new or different sensor configurations. This inflexibility prevents models from generalizing across heterogeneous health sensing ecosystems, where data spans high-frequency wearable streams to sparse, day-scale clinical signals. To bridge this gap, we introduce SLIP (Sensor Language-Informed Pretraining), an open-source framework for learning language-aligned representations that generalize across diverse sensor setups. SLIP integrates contrastive alignment with sensor-conditioned captioning, facilitating both discriminative understanding and generative reasoning. By repurposing a pretrained decoder-only language model via cross-attention and introducing a flexible patch embedder, SLIP transfers to new sensor configurations at inference without retraining, regardless of their native sampling rate or input length. Across 11 datasets, SLIP demonstrates superior performance in retrieval, signal captioning, and question answering. It achieves a 77.14% average linear-probing accuracy, a 3.56% improvement over the strongest baseline (SensorLM). Beyond classification, SLIP supports open-vocabulary sensor captioning and question answering without any architectural modifications. Our work offers broad implications for the development of sensor-language models, utilizing cross-modal supervision to ensure more robust generalizability in health sensing applications.