Arrow Research search

Author name cluster

Erkki Oja

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

10 papers
2 author rows

Possible papers

10

JMLR Journal 2016 Journal Article

Low-Rank Doubly Stochastic Matrix Decomposition for Cluster Analysis

  • Zhirong Yang
  • Jukka Corander
  • Erkki Oja

Cluster analysis by nonnegative low-rank approximations has experienced a remarkable progress in the past decade. However, the majority of such approximation approaches are still restricted to nonnegative matrix factorization (NMF) and suffer from the following two drawbacks: 1) they are unable to produce balanced partitions for large-scale manifold data which are common in real-world clustering tasks; 2) most existing NMF-type clustering methods cannot automatically determine the number of clusters. We propose a new low-rank learning method to address these two problems, which is beyond matrix factorization. Our method approximately decomposes a sparse input similarity in a normalized way and its objective can be used to learn both cluster assignments and the number of clusters. For efficient optimization, we use a relaxed formulation based on Data- Cluster-Data random walk, which is also shown to be equivalent to low-rank factorization of the doubly-stochastically normalized cluster incidence matrix. The probabilistic cluster assignments can thus be learned with a multiplicative majorization-minimization algorithm. Experimental results show that the new method is more accurate both in terms of clustering large-scale manifold data sets and of selecting the number of clusters. [abs] [ pdf ][ bib ] &copy JMLR 2016. ( edit, beta )

NeurIPS Conference 2012 Conference Paper

Clustering by Nonnegative Matrix Factorization Using Graph Random Walk

  • Zhirong Yang
  • Tele Hao
  • Onur Dikmen
  • Xi Chen
  • Erkki Oja

Nonnegative Matrix Factorization (NMF) is a promising relaxation technique for clustering analysis. However, conventional NMF methods that directly approximate the pairwise similarities using the least square error often yield mediocre performance for data in curved manifolds because they can capture only the immediate similarities between data samples. Here we propose a new NMF clustering method which replaces the approximated matrix with its smoothed version using random walk. Our method can thus accommodate farther relationships between data samples. Furthermore, we introduce a novel regularization in the proposed objective function in order to improve over spectral clustering. The new learning objective is optimized by a multiplicative Majorization-Minimization algorithm with a scalable implementation for learning the factorizing matrix. Extensive experimental results on real-world datasets show that our method has strong performance in terms of cluster purity.

NeurIPS Conference 2009 Conference Paper

Heavy-Tailed Symmetric Stochastic Neighbor Embedding

  • Zhirong Yang
  • Irwin King
  • Zenglin Xu
  • Erkki Oja

Stochastic Neighbor Embedding (SNE) has shown to be quite promising for data visualization. Currently, the most popular implementation, t-SNE, is restricted to a particular Student t-distribution as its embedding distribution. Moreover, it uses a gradient descent algorithm that may require users to tune parameters such as the learning step size, momentum, etc. , in finding its optimum. In this paper, we propose the Heavy-tailed Symmetric Stochastic Neighbor Embedding (HSSNE) method, which is a generalization of the t-SNE to accommodate various heavy-tailed embedding similarity functions. With this generalization, we are presented with two difficulties. The first is how to select the best embedding similarity among all heavy-tailed functions and the second is how to optimize the objective function once the heave-tailed function has been selected. Our contributions then are: (1) we point out that various heavy-tailed embedding similarities can be characterized by their negative score functions. Based on this finding, we present a parameterized subset of similarity functions for choosing the best tail-heaviness for HSSNE; (2) we present a fixed-point optimization algorithm that can be applied to all heavy-tailed functions and does not require the user to set any parameters; and (3) we present two empirical studies, one for unsupervised visualization showing that our optimization algorithm runs as fast and as good as the best known t-SNE implementation and the other for semi-supervised visualization showing quantitative superiority using the homogeneity measure as well as qualitative advantage in cluster separation over t-SNE.

TCS Journal 2002 Journal Article

Unsupervised learning in neural computation

  • Erkki Oja

In this article, we consider unsupervised learning from the point of view of applying neural computation on signal and data analysis problems. The article is an introductory survey, concentrating on the main principles and categories of unsupervised learning. In neural computation, there are two classical categories for unsupervised learning methods and models: first, extensions of principal component analysis and factor analysis, and second, learning vector coding or clustering methods that are based on competitive learning. These are covered in this article. The more recent trend in unsupervised learning is to consider this problem in the framework of probabilistic generative models. If it is possible to build and estimate a model that explains the data in terms of some latent variables, key insights may be obtained into the true nature and structure of the data. This approach is also briefly reviewed.

NeurIPS Conference 1998 Conference Paper

Sparse Code Shrinkage: Denoising by Nonlinear Maximum Likelihood Estimation

  • Aapo Hyvärinen
  • Patrik Hoyer
  • Erkki Oja

Sparse coding is a method for finding a representation of data in which each of the components of the representation is only rarely significantly active. Such a representation is closely related to re(cid: 173) dundancy reduction and independent component analysis, and has some neurophysiological plausibility. In this paper, we show how sparse coding can be used for denoising. Using maximum likelihood estimation of nongaussian variables corrupted by gaussian noise, we show how to apply a shrinkage nonlinearity on the components of sparse coding so as to reduce noise. Furthermore, we show how to choose the optimal sparse coding basis for denoising. Our method is closely related to the method of wavelet shrinkage, but has the important benefit over wavelet methods that both the features and the shrinkage parameters are estimated directly from the data.

NeurIPS Conference 1997 Conference Paper

Independent Component Analysis for Identification of Artifacts in Magnetoencephalographic Recordings

  • Ricardo Vigário
  • Veikko Jousmäki
  • Matti Hämäläinen
  • Riitta Hari
  • Erkki Oja

We have studied the application of an independent component analysis (ICA) approach to the identification and possible removal of artifacts from a magnetoencephalographic (MEG) recording. This statistical tech(cid: 173) nique separates components according to the kurtosis of their amplitude distributions over time, thus distinguishing between strictly periodical signals, and regularly and irregularly occurring signals. Many artifacts belong to the last category. In order to assess the effectiveness of the method, controlled artifacts were produced, which included saccadic eye movements and blinks, increased muscular tension due to biting and the presence of a digital watch inside the magnetically shielded room. The results demonstrate the capability of the method to identify and clearly isolate the produced artifacts.

NeurIPS Conference 1997 Conference Paper

S-Map: A Network with a Simple Self-Organization Algorithm for Generative Topographic Mappings

  • Kimmo Kiviluoto
  • Erkki Oja

The S-Map is a network with a simple learning algorithm that com(cid: 173) bines the self-organization capability of the Self-Organizing Map (SOM) and the probabilistic interpretability of the Generative To(cid: 173) pographic Mapping (GTM). The simulations suggest that the S(cid: 173) Map algorithm has a stronger tendency to self-organize from ran(cid: 173) dom initial configuration than the GTM. The S-Map algorithm can be further simplified to employ pure Hebbian learning, with(cid: 173) out changing the qualitative behaviour of the network.

NeurIPS Conference 1996 Conference Paper

One-unit Learning Rules for Independent Component Analysis

  • Aapo Hyvärinen
  • Erkki Oja

Neural one-unit learning rules for the problem of Independent Com(cid: 173) ponent Analysis (ICA) and blind source separation are introduced. In these new algorithms, every ICA neuron develops into a sepa(cid: 173) rator that finds one of the independent components. The learning rules use very simple constrained Hebbianjanti-Hebbian learning in which decorrelating feedback may be added. To speed up the convergence of these stochastic gradient descent rules, a novel com(cid: 173) putationally efficient fixed-point algorithm is introduced.

v2026.09.13