arXiv · 2109.08399
Cross-Leverage Scores for Selecting Subsets of Explanatory Variables
Abstract
In a standard regression problem, we have a set of explanatory variables whose effect on some response vector is modeled. For wide binary data, such as genetic marker data, we often have two limitations. First, we have more parameters than observations. Second, main effects are not the main focus; instead the primary aim is to uncover interactions between the binary variables that effect the response. Methods such as logic regression are able to find combinations of the explanatory variables that capture higher-order relationships in the response. However, the number of explanatory variables these methods can handle is highly limited. To address these two limitations we need to reduce the number of variables prior to computationally demanding analyses. In this paper, we demonstrate the usefulness of using so-called cross-leverage scores as a means of sampling subsets of explanatory variables while retaining the valuable interactions.
Explore related subjects
Keep this discovery
Katharina Parry, Leo N. Geppert, Alexander Munteanu, Katja Ickstadt. 2021-09-17. Cross-Leverage Scores for Selecting Subsets of Explanatory Variables. https://arxiv.org/abs/2109.08399
Cite the original work for its findings. Save a collection to share your selection of sources.