Arrow Research search

Author name cluster

Stéphane Marchand-Maillet

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

4 papers
1 author row

Possible papers

4

UAI Conference 2018 Conference Paper

Structured nonlinear variable selection

  • Magda Gregorova
  • Alexandros Kalousis
  • Stéphane Marchand-Maillet

In the above, R(f ) is a suitable penalty typically based on some prior assumption about the function space (e. g. smoothness), and τ > 0 is a suitable regularization hyper-parameter. The principal assumption we consider in this paper is that the function f is sparse with respect to the original input space X, that is it depends only on l d input variables. We investigate structured sparsity methods for variable selection in regression problems where the target depends nonlinearly on the inputs. We focus on general nonlinear functions not limiting a priori the function space to additive models. We propose two new regularizers based on partial derivatives as nonlinear equivalents of group lasso and elastic net. We formulate the problem within the framework of learning in reproducing kernel Hilbert spaces and show how the variational problem can be reformulated into a more practical finite dimensional equivalent. We develop a new algorithm derived from the ADMM principles that relies solely on closed forms of the proximal operators. We explore the empirical properties of our new algorithm for Nonlinear Variable Selection based on Derivatives (NVSD) on a set of experiments and confirm favourable properties of our structured-sparsity models and the algorithm in terms of both prediction and variable selection accuracy. 1 Learning with variable selection is a well-established and rather well-explored problem in the case of linear Pd models f (x) = a xa wa, e. g. Hastie et al. (2015). The main ideas from linear models have been Pdsuccessfully transferred to additive models f (x) = a fa (xa ), e. g. Ravikumar et al. (2007), Bach (2009), Koltchinskii and Yuan (2010), and Yin et al. (2012), Pd or to additive models with interactions f (x) = a fa (xa ) + Pd a<b fa, b (xa, xb ), e. g Lin and Zhang (2006) and Tyagi et al. (2016). However, sparse modelling of general non-linear functions is more intricate. A promising stream of works focuses on the use of non-linear (conditional) crosscovariance operators arising from embedding probability measures into Hilbert function spaces, e. g. Yamada et al. (2014) and Chen et al. (2017).

ICML Conference 2015 Conference Paper

Information Geometry and Minimum Description Length Networks

  • Ke Sun 0001
  • Jun Wang 0017
  • Alexandros Kalousis
  • Stéphane Marchand-Maillet

We study parametric unsupervised mixture learning. We measure the loss of intrinsic information from the observations to complex mixture models, and then to simple mixture models. We present a geometric picture, where all these representations are regarded as free points in the space of probability distributions. Based on minimum description length, we derive a simple geometric principle to learn all these models together. We present a new learning machine with theories, algorithms, and simulations.

ICML Conference 2014 Conference Paper

An Information Geometry of Statistical Manifold Learning

  • Ke Sun 0001
  • Stéphane Marchand-Maillet

Manifold learning seeks low-dimensional representations of high-dimensional data. The main tactics have been exploring the geometry in an input data space and an output embedding space. We develop a manifold learning theory in a hypothesis space consisting of models. A model means a specific instance of a collection of points, e. g. , the input data collectively or the output embedding collectively. The semi-Riemannian metric of this hypothesis space is uniquely derived in closed form based on the information geometry of probability distributions. There, manifold learning is interpreted as a trajectory of intermediate models. The volume of a continuous region reveals an amount of information. It can be measured to define model complexity and embedding quality. This provides deep unified perspectives of manifold learning theory.

ICML Conference 2014 Conference Paper

Two-Stage Metric Learning

  • Jun Wang 0017
  • Ke Sun 0001
  • Fei Sha
  • Stéphane Marchand-Maillet
  • Alexandros Kalousis

In this paper, we present a novel two-stage metric learning algorithm. We first map each learning instance to a probability distribution by computing its similarities to a set of fixed anchor points. Then, we define the distance in the input data space as the Fisher information distance on the associated statistical manifold. This induces in the input data space a new family of distance metric which presents unique properties. Unlike kernelized metric learning, we do not require the similarity measure to be positive semi-definite. Moreover, it can also be interpreted as a local metric learning algorithm with well defined distance approximation. We evaluate its performance on a number of datasets. It outperforms significantly other metric learning methods and SVM.

v2026.09.13