arXiv · 0907.3241
Statistical mechanics of sparse generalization and model selection
Abstract
One of the crucial tasks in many inference problems is the extraction of sparse information out of a given number of high-dimensional measurements. In machine learning, this is frequently achieved using, as a penality term, the $L_p$ norm of the model parameters, with $p\leq 1$ for efficient dilution. Here we propose a statistical-mechanics analysis of the problem in the setting of perceptron memorization and generalization. Using a replica approach, we are able to evaluate the relative performance of naive dilution (obtained by learning without dilution, following by applying a threshold to the model parameters), $L_1$ dilution (which is frequently used in convex optimization) and $L_0$ dilution (which is optimal but computationally hard to implement). Whereas both $L_p$ diluted approaches clearly outperform the naive approach, we find a small region where $L_0$ works almost perfectly and strongly outperforms the simpler to implement $L_1$ dilution.
Explore related subjects
Keep this discovery
Alejandro Lage-Castellanos, Andrea Pagnani, Martin Weigt. 2009-07-18. Statistical mechanics of sparse generalization and model selection. https://doi.org/10.1088/1742-5468/2009/10/p10009
Cite the original work for its findings. Save a collection to share your selection of sources.