UAI Conference 2018 Conference Paper
Structured nonlinear variable selection
- Magda Gregorova
- Alexandros Kalousis
- Stéphane Marchand-Maillet
In the above, R(f ) is a suitable penalty typically based on some prior assumption about the function space (e. g. smoothness), and τ > 0 is a suitable regularization hyper-parameter. The principal assumption we consider in this paper is that the function f is sparse with respect to the original input space X, that is it depends only on l d input variables. We investigate structured sparsity methods for variable selection in regression problems where the target depends nonlinearly on the inputs. We focus on general nonlinear functions not limiting a priori the function space to additive models. We propose two new regularizers based on partial derivatives as nonlinear equivalents of group lasso and elastic net. We formulate the problem within the framework of learning in reproducing kernel Hilbert spaces and show how the variational problem can be reformulated into a more practical finite dimensional equivalent. We develop a new algorithm derived from the ADMM principles that relies solely on closed forms of the proximal operators. We explore the empirical properties of our new algorithm for Nonlinear Variable Selection based on Derivatives (NVSD) on a set of experiments and confirm favourable properties of our structured-sparsity models and the algorithm in terms of both prediction and variable selection accuracy. 1 Learning with variable selection is a well-established and rather well-explored problem in the case of linear Pd models f (x) = a xa wa, e. g. Hastie et al. (2015). The main ideas from linear models have been Pdsuccessfully transferred to additive models f (x) = a fa (xa ), e. g. Ravikumar et al. (2007), Bach (2009), Koltchinskii and Yuan (2010), and Yin et al. (2012), Pd or to additive models with interactions f (x) = a fa (xa ) + Pd a<b fa, b (xa, xb ), e. g Lin and Zhang (2006) and Tyagi et al. (2016). However, sparse modelling of general non-linear functions is more intricate. A promising stream of works focuses on the use of non-linear (conditional) crosscovariance operators arising from embedding probability measures into Hilbert function spaces, e. g. Yamada et al. (2014) and Chen et al. (2017).