arXiv · 2503.24157
LLM4FS: Leveraging Large Language Models for Feature Selection
Abstract
Recent advances in large language models (LLMs) have provided new opportunities for decision-making, particularly in the task of automated feature selection. In this paper, we first comprehensively evaluate LLM-based feature selection methods, covering the state-of-the-art DeepSeek-R1, GPT-o3-mini, and GPT-4.5. Then, we propose a new hybrid strategy called LLM4FS that integrates LLMs with traditional data-driven methods. Specifically, input data samples into LLMs, and directly call traditional data-driven techniques such as random forest and forward sequential selection. Notably, our analysis reveals that the hybrid strategy leverages the contextual understanding of LLMs and the high statistical reliability of traditional data-driven methods to achieve excellent feature selection performance, even surpassing LLMs and traditional data-driven methods. Finally, we point out the limitations of its application in decision-making. Our code is available at https://github.com/xianchaoxiu/LLM4FS.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jianhao Li, Xianchao Xiu. 2025-03-31. LLM4FS: Leveraging Large Language Models for Feature Selection. https://arxiv.org/abs/2503.24157
Cite the original work for its findings. Save a collection to share your selection of sources.