Arrow Research search

Author name cluster

Aapo Hyvärinen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

44 papers
2 author rows

Possible papers

44

ICML Conference 2025 Conference Paper

Density Ratio Estimation with Conditional Probability Paths

  • Hanlin Yu
  • Arto Klami
  • Aapo Hyvärinen
  • Anna Korba
  • Omar Chehab

Density ratio estimation in high dimensions can be reframed as integrating a certain quantity, the time score, over probability paths which interpolate between the two densities. In practice, the time score has to be estimated based on samples from the two densities. However, existing methods for this problem remain computationally expensive and can yield inaccurate estimates. Inspired by recent advances in generative modeling, we introduce a novel framework for time score estimation, based on a conditioning variable. Choosing the conditioning variable judiciously enables a closed-form objective function. We demonstrate that, compared to previous approaches, our approach results in faster learning of the time score and competitive or better estimation accuracies of the density ratio on challenging tasks. Furthermore, we establish theoretical guarantees on the error of the estimated density ratio.

ICML Conference 2024 Conference Paper

Causal Representation Learning Made Identifiable by Grouping of Observational Variables

  • Hiroshi Morioka
  • Aapo Hyvärinen

A topic of great current interest is Causal Representation Learning (CRL), whose goal is to learn a causal model for hidden features in a data-driven manner. Unfortunately, CRL is severely ill-posed since it is a combination of the two notoriously ill-posed problems of representation learning and causal discovery. Yet, finding practical identifiability conditions that guarantee a unique solution is crucial for its practical applicability. Most approaches so far have been based on assumptions on the latent causal mechanisms, such as temporal causality, or existence of supervision or interventions; these can be too restrictive in actual applications. Here, we show identifiability based on novel, weak constraints, which requires no temporal structure, intervention, nor weak supervision. The approach is based on assuming the observational mixing exhibits a suitable grouping of the observational variables. We also propose a novel self-supervised estimation framework consistent with the model, prove its statistical consistency, and experimentally show its superior CRL performances compared to the state-of-the-art baselines. We further demonstrate its robustness against latent confounders and causal cycles.

YNIMG Journal 2023 Journal Article

Unsupervised representation learning of spontaneous MEG data with nonlinear ICA

  • Yongjie Zhu
  • Tiina Parviainen
  • Erkka Heinilä
  • Lauri Parkkonen
  • Aapo Hyvärinen

Resting-state magnetoencephalography (MEG) data show complex but structured spatiotemporal patterns. However, the neurophysiological basis of these signal patterns is not fully known and the underlying signal sources are mixed in MEG measurements. Here, we developed a method based on the nonlinear independent component analysis (ICA), a generative model trainable with unsupervised learning, to learn representations from resting-state MEG data. After being trained with a large dataset from the Cam-CAN repository, the model has learned to represent and generate patterns of spontaneous cortical activity using latent nonlinear components, which reflects principal cortical patterns with specific spectral modes. When applied to the downstream classification task of audio-visual MEG, the nonlinear ICA model achieves competitive performance with deep neural networks despite limited access to labels. We further validate the generalizability of the model across different datasets by applying it to an independent neurofeedback dataset for decoding the subject's attentional states, providing a real-time feature extraction and decoding mindfulness and thought-inducing tasks with an accuracy of around 70% at the individual level, which is much higher than obtained by linear ICA or other baseline methods. Our results demonstrate that nonlinear ICA is a valuable addition to existing tools, particularly suited for unsupervised representation learning of spontaneous MEG activity which can then be applied to specific goals or tasks when labelled data are scarce.

UAI Conference 2022 Conference Paper

Binary independent component analysis: a non-stationarity-based approach

  • Antti Hyttinen
  • Vitória Barin Pacela
  • Aapo Hyvärinen

We consider independent component analysis of binary data. While fundamental in practice, this case has been much less developed than ICA for continuous data. We start by assuming a linear mixing model in a continuous-valued latent space, followed by a binary observation model. Importantly, we assume that the sources are non-stationary; this is necessary since any non-Gaussianity would essentially be destroyed by the binarization. Interestingly, the model allows for closed-form likelihood by employing the cumulative distribution function of the multivariate Gaussian distribution. In stark contrast to the continuous-valued case, we prove non-identifiability of the model with few observed variables; our empirical results imply identifiability when the number of observed variables is higher. We present a practical method for binary ICA that uses only pairwise marginals, which are faster to compute than the full multivariate likelihood. Experiments give insight into the requirements for the number of observed variables, segments, and latent sources that allow the model to be estimated.

YNIMG Journal 2022 Journal Article

Dynamics of retinotopic spatial attention revealed by multifocal MEG

  • Ilmari Kurki
  • Aapo Hyvärinen
  • Linda Henriksson

Visual focal attention is both fast and spatially localized, making it challenging to investigate using human neuroimaging paradigms. Here, we used a new multivariate multifocal mapping method with magnetoencephalography (MEG) to study how focal attention in visual space changes stimulus-evoked responses across the visual field. The observer's task was to detect a color change in the target location, or at the central fixation. Simultaneously, 24 regions in visual space were stimulated in parallel using an orthogonal, multifocal mapping stimulus sequence. First, we used univariate analysis to estimate stimulus-evoked responses in each channel. Then we applied multivariate pattern analysis to look for attentional effects on the responses. We found that attention to a target location causes two spatially and temporally separate effects. Initially, attentional modulation is brief, observed at around 60-130 ms post stimulus, and modulates responses not only at the target location but also in adjacent regions. A later modulation was observed from around 200 ms, which was specific to the location of the attentional target. The results support the idea that focal attention employs several processing stages and suggest that early attentional modulation is less spatially specific than late.

UAI Conference 2022 Conference Paper

The optimal noise in noise-contrastive learning is not what you think

  • Omar Chehab
  • Alexandre Gramfort
  • Aapo Hyvärinen

Learning a parametric model of a data distribution is a well-known statistical problem that has seen renewed interest as it is brought to scale in deep learning. Framing the problem as a self-supervised task, where data samples are discriminated from noise samples, is at the core of state-of-the-art methods, beginning with Noise-Contrastive Estimation (NCE). Yet, such contrastive learning requires a good noise distribution, which is hard to specify; domain-specific heuristics are therefore widely used. While a comprehensive theory is missing, it is widely assumed that the optimal noise should in practice be made equal to the data, both in distribution and proportion. This setting underlies Generative Adversarial Networks (GANs) in particular. Here, we empirically and theoretically challenge this assumption on the optimal noise. We show that deviating from this assumption can actually lead to better statistical estimators, in terms of asymptotic variance. In particular, the optimal noise distribution is different from the data’s and even from a different family.

YNIMG Journal 2020 Journal Article

Brain activity reflects the predictability of word sequences in listened continuous speech

  • Miika Koskinen
  • Mikko Kurimo
  • Joachim Gross
  • Aapo Hyvärinen
  • Riitta Hari

Natural speech builds on contextual relations that can prompt predictions of upcoming utterances. To study the neural underpinnings of such predictive processing we asked 10 healthy adults to listen to a 1-h-long audiobook while their magnetoencephalographic (MEG) brain activity was recorded. We correlated the MEG signals with acoustic speech envelope, as well as with estimates of Bayesian word probability with and without the contextual word sequence (N-gram and Unigram, respectively), with a focus on time-lags. The MEG signals of auditory and sensorimotor cortices were strongly coupled to the speech envelope at the rates of syllables (4-8 ​Hz) and of prosody and intonation (0.5-2 ​Hz). The probability structure of word sequences, independently of the acoustical features, affected the ≤ 2-Hz signals extensively in auditory and rolandic regions, in precuneus, occipital cortices, and lateral and medial frontal regions. Fine-grained temporal progression patterns occurred across brain regions 100-1000 ​ms after word onsets. Although the acoustic effects were observed in both hemispheres, the contextual influences were statistically significantly lateralized to the left hemisphere. These results serve as a brain signature of the predictability of word sequences in listened continuous speech, confirming and extending previous results to demonstrate that deeply-learned knowledge and recent contextual information are employed dynamically and in a left-hemisphere-dominant manner in predicting the forthcoming words in natural speech.

UAI Conference 2020 Conference Paper

Hidden Markov Nonlinear ICA: Unsupervised Learning from Nonstationary Time Series

  • Hermanni Hälvä
  • Aapo Hyvärinen

Recent advances in nonlinear Independent Component Analysis (ICA) provide a principled framework for unsupervised feature learning and disentanglement. The central idea in such works is that the latent components are assumed to be independent conditional on some observed auxiliary variables, such as the time-segment index. This requires manual segmentation of data into non-stationary segments which is computationally expensive, inaccurate and often impossible. These models are thus not fully unsupervised. We remedy these limitations by combining nonlinear ICA with a Hidden Markov Model, resulting in a model where a latent state acts in place of the observed segment-index. We prove identifiability of the proposed model for a general mixing nonlinearity, such as a neural network. We also show how maximum likelihood estimation of the model can be done using the expectation-maximization algorithm. Thus, we achieve a new nonlinear ICA framework which is unsupervised, more efficient, as well as able to model underlying temporal dynamics.

YNIMG Journal 2020 Journal Article

Nonlinear ICA of fMRI reveals primitive temporal structures linked to rest, task, and behavioral traits

  • Hiroshi Morioka
  • Vince Calhoun
  • Aapo Hyvärinen

Accumulating evidence from whole brain functional magnetic resonance imaging (fMRI) suggests that the human brain at rest is functionally organized in a spatially and temporally constrained manner. However, because of their complexity, the fundamental mechanisms underlying time-varying functional networks are still not well understood. Here, we develop a novel nonlinear feature extraction framework called local space-contrastive learning (LSCL), which extracts distinctive nonlinear temporal structure hidden in time series, by training a deep temporal convolutional neural network in an unsupervised, data-driven manner. We demonstrate that LSCL identifies certain distinctive local temporal structures, referred to as temporal primitives, which repeatedly appear at different time points and spatial locations, reflecting dynamic resting-state networks. We also show that these temporal primitives are also present in task-evoked spatiotemporal responses. We further show that the temporal primitives capture unique aspects of behavioral traits such as fluid intelligence and working memory. These results highlight the importance of capturing transient spatiotemporal dynamics within fMRI data and suggest that such temporal primitives may capture fundamental information underlying both spontaneous and task-induced fMRI dynamics.

UAI Conference 2020 Conference Paper

Robust contrastive learning and nonlinear ICA in the presence of outliers

  • Hiroaki Sasaki
  • Takashi Takenouchi
  • Ricardo Pio Monti
  • Aapo Hyvärinen

Nonlinear independent component analysis (ICA) is a general framework for unsupervised representation learning, and aimed at recovering the latent variables in data. Recent practical methods perform nonlinear ICA by solving classification problems based on logistic regression. However, it is well-known that logistic regression is vulnerable to outliers, and thus the performance can be strongly weakened by outliers. In this paper, we first theoretically analyze nonlinear ICA models in the presence of outliers. Our analysis implies that estimation in nonlinear ICA can be seriously hampered when outliers exist on the tails of the (noncontaminated) target density, which happens in a typical case of contamination by outliers. We develop two robust nonlinear ICA methods based on the $\gamma$-divergence, which is a robust alternative to the KL-divergence in logistic regression. The proposed methods are theoretically shown to have desired robustness properties in the context of nonlinear ICA. We also experimentally demonstrate that the proposed methods are very robust and outperform existing methods in the presence of outliers. Finally, the proposed method is applied to ICA-based causal discovery and shown to find a plausible causal relationship on fMRI data.

UAI Conference 2019 Conference Paper

Causal Discovery with General Non-Linear Relationships using Non-Linear ICA

  • Ricardo Pio Monti
  • Kun Zhang 0001
  • Aapo Hyvärinen

We consider the problem of inferring causal relationships between two or more passively observed variables. While the problem of such causal discovery has been extensively studied especially in the bivariate setting, the majority of current methods assume a linear causal relationship, and the few methods which consider non-linear dependencies usually make the assumption of additive noise. Here, we propose a framework through which we can perform causal discovery in the presence of general non-linear relationships. The proposed method is based on recent progress in non-linear independent component analysis and exploits the non-stationarity of observations in order to recover the underlying sources or latent disturbances. We show rigorously that in the case of bivariate causal discovery, such non-linear ICA can be used to infer the causal direction via a series of independence tests. We further propose an alternative measure of causal direction based on asymptotic approximations to the likelihood ratio, as well as an extension to multivariate causal discovery. We demonstrate the capabilities of the proposed method via a series of simulation studies and conclude with an application to neuroimaging data.

JMLR Journal 2019 Journal Article

Neural Empirical Bayes

  • Saeed Saremi
  • Aapo Hyvärinen

We unify kernel density estimation and empirical Bayes and address a set of problems in unsupervised machine learning with a geometric interpretation of those methods, rooted in the concentration of measure phenomenon. Kernel density is viewed symbolically as $X\rightharpoonup Y$ where the random variable $X$ is smoothed to $Y= X+N(0,\sigma^2 I_d)$, and empirical Bayes is the machinery to denoise in a least-squares sense, which we express as $X \leftharpoondown Y$. A learning objective is derived by combining these two, symbolically captured by $X \rightleftharpoons Y$. Crucially, instead of using the original nonparametric estimators, we parametrize the energy function with a neural network denoted by $\phi$; at optimality, $\nabla \phi \approx -\nabla \log f$ where $f$ is the density of $Y$. The optimization problem is abstracted as interactions of high-dimensional spheres which emerge due to the concentration of isotropic Gaussians. We introduce two algorithmic frameworks based on this machinery: (i) a “walk-jump” sampling scheme that combines Langevin MCMC (walks) and empirical Bayes (jumps), and (ii) a probabilistic framework for associative memory, called NEBULA, defined a la Hopfield by the gradient flow of the learned energy to a set of attractors. We finish the paper by reporting the emergence of very rich “creative memories” as attractors of NEBULA for highly-overlapping spheres. [abs] [ pdf ][ bib ] &copy JMLR 2019. ( edit, beta )

UAI Conference 2018 Conference Paper

A unified probabilistic model for learning latent factors and their connectivities from high-dimensional data

  • Ricardo Pio Monti
  • Aapo Hyvärinen

Connectivity estimation is challenging in the context of high-dimensional data. A useful preprocessing step is to group variables into clusters, however, it is not always clear how to do so from the perspective of connectivity estimation. Another practical challenge is that we may have data from multiple related classes (e. g. , multiple subjects or conditions) and wish to incorporate constraints on the similarities across classes. We propose a probabilistic model which simultaneously performs both a grouping of variables (i. e. , detecting community structure) and estimation of connectivities between the groups which correspond to latent variables. The model is essentially a factor analysis model where the factors are allowed to have arbitrary correlations, while the factor loading matrix is constrained to express a community structure. The model can be applied on multiple classes so that the connectivities can be different between the classes, while the community structure is the same for all classes. We propose an efficient estimation algorithm based on score matching, and prove the identifiability of the model. Finally, we present an extension to directed (causal) connectivities over latent variables. Simulations and experiments on fMRI data validate the practical utility of the method.

JMLR Journal 2018 Journal Article

Mode-Seeking Clustering and Density Ridge Estimation via Direct Estimation of Density-Derivative-Ratios

  • Hiroaki Sasaki
  • Takafumi Kanamori
  • Aapo Hyvärinen
  • Gang Niu
  • Masashi Sugiyama

Modes and ridges of the probability density function behind observed data are useful geometric features. Mode-seeking clustering assigns cluster labels by associating data samples with the nearest modes, and estimation of density ridges enables us to find lower-dimensional structures hidden in data. A key technical challenge both in mode-seeking clustering and density ridge estimation is accurate estimation of the ratios of the first- and second-order density derivatives to the density. A naive approach takes a three-step approach of first estimating the data density, then computing its derivatives, and finally taking their ratios. However, this three-step approach can be unreliable because a good density estimator does not necessarily mean a good density derivative estimator, and division by the estimated density could significantly magnify the estimation error. To cope with these problems, we propose a novel estimator for the density-derivative-ratios. The proposed estimator does not involve density estimation, but rather directly approximates the ratios of density derivatives of any order. Moreover, we establish a convergence rate of the proposed estimator. Based on the proposed estimator, novel methods both for mode-seeking clustering and density ridge estimation are developed, and the respective convergence rates to the mode and ridge of the underlying density are also established. Finally, we experimentally demonstrate that the developed methods significantly outperform existing methods, particularly for relatively high-dimensional data. [abs] [ pdf ][ bib ] &copy JMLR 2018. ( edit, beta )

JMLR Journal 2017 Journal Article

Density Estimation in Infinite Dimensional Exponential Families

  • Bharath Sriperumbudur
  • Kenji Fukumizu
  • Arthur Gretton
  • Aapo Hyvärinen
  • Revant Kumar

In this paper, we consider an infinite dimensional exponential family $\mathcal{P}$ of probability densities, which are parametrized by functions in a reproducing kernel Hilbert space $\mathcal{H}$, and show it to be quite rich in the sense that a broad class of densities on $\mathbb{R}^d$ can be approximated arbitrarily well in Kullback-Leibler (KL) divergence by elements in $\mathcal{P}$. Motivated by this approximation property, the paper addresses the question of estimating an unknown density $p_0$ through an element in $\mathcal{P}$. Standard techniques like maximum likelihood estimation (MLE) or pseudo MLE (based on the method of sieves), which are based on minimizing the KL divergence between $p_0$ and $\mathcal{P}$, do not yield practically useful estimators because of their inability to efficiently handle the log-partition function. We propose an estimator $\hat{p}_n$ based on minimizing the Fisher divergence, $J(p_0\Vert p)$ between $p_0$ and $p\in \mathcal{P}$, which involves solving a simple finite-dimensional linear system. When $p_0\in\mathcal{P}$, we show that the proposed estimator is consistent, and provide a convergence rate of $n^{-\min\left\{\frac{2}{3},\frac{2\beta+1}{2\beta+2}\right\}}$ in Fisher divergence under the smoothness assumption that $\log p_0\in\mathcal{R}(C^\beta)$ for some $\beta\ge 0$, where $C$ is a certain Hilbert-Schmidt operator on $\mathcal{H}$ and $\mathcal{R}(C^\beta)$ denotes the image of $C^\beta$. We also investigate the misspecified case of $p_0\notin\mathcal{P}$ and show that $J(p_0\Vert\hat{p}_n)\rightarrow \inf_{p\in\mathcal{P}}J(p_0\Vert p)$ as $n\rightarrow \infty$, and provide a rate for this convergence under a similar smoothness condition as above. Through numerical simulations we demonstrate that the proposed estimator outperforms the non- parametric kernel density estimator, and that the advantage of the proposed estimator grows as $d$ increases. [abs] [ pdf ][ bib ] &copy JMLR 2017. ( edit, beta )

ICML Conference 2017 Conference Paper

SPLICE: Fully Tractable Hierarchical Extension of ICA with Pooling

  • Junichiro Hirayama
  • Aapo Hyvärinen
  • Motoaki Kawanabe

We present a novel probabilistic framework for a hierarchical extension of independent component analysis (ICA), with a particular motivation in neuroscientific data analysis and modeling. The framework incorporates a general subspace pooling with linear ICA-like layers stacked recursively. Unlike related previous models, our generative model is fully tractable: both the likelihood and the posterior estimates of latent variables can readily be computed with analytically simple formulae. The model is particularly simple in the case of complex-valued data since the pooling can be reduced to taking the modulus of complex numbers. Experiments on electroencephalography (EEG) and natural images demonstrate the validity of the method.

YNIMG Journal 2014 Journal Article

Group-level spatial independent component analysis of Fourier envelopes of resting-state MEG data

  • Pavan Ramkumar
  • Lauri Parkkonen
  • Aapo Hyvärinen

We developed a data-driven method to spatiotemporally and spectrally characterize the dynamics of brain oscillations in resting-state magnetoencephalography (MEG) data. The method, called envelope spatial Fourier independent component analysis (eSFICA), maximizes the spatial and spectral sparseness of Fourier energies of a cortically constrained source current estimate. We compared this method using a simulated data set against 5 other variants of independent component analysis and found that eSFICA performed on par with its temporal variant, eTFICA, and better than other ICA variants, in characterizing dynamics at time scales of the order of minutes. We then applied eSFICA to real MEG data obtained from 9 subjects during rest. The method identified several networks showing within- and cross-frequency inter-areal functional connectivity profiles which resemble previously reported resting-state networks, such as the bilateral sensorimotor network at ~20Hz, the lateral and medial parieto-occipital sources at ~10Hz, a subset of the default-mode network at ~8 and ~15Hz, and lateralized temporal lobe sources at ~8Hz. Finally, we interpreted the estimated networks as spatiospectral filters and applied the filters to obtain the dynamics during a natural stimulus sequence presented to the same 9 subjects. We observed occipital alpha modulation to visual stimuli, bilateral rolandic mu modulation to tactile stimuli and video clips of hands, and the temporal lobe network modulation to speech stimuli, but no modulation of the sources in the default-mode network. We conclude that (1) the proposed method robustly detects inter-areal cross-frequency networks at long time scales, (2) the functional relevance of the resting-state networks can be probed by applying the obtained spatiospectral filters to data from measurements with controlled external stimulation.

YNIMG Journal 2014 Journal Article

Group-PCA for very large fMRI datasets

  • Stephen M. Smith
  • Aapo Hyvärinen
  • Gaël Varoquaux
  • Karla L. Miller
  • Christian F. Beckmann

Increasingly-large datasets (for example, the resting-state fMRI data from the Human Connectome Project) are demanding analyses that are problematic because of the sheer scale of the aggregate data. We present two approaches for applying group-level PCA; both give a close approximation to the output of PCA applied to full concatenation of all individual datasets, while having very low memory requirements regardless of the number of datasets being combined. Across a range of realistic simulations, we find that in most situations, both methods are more accurate than current popular approaches for analysis of multi-subject resting-state fMRI studies. The group-PCA output can be used to feed into a range of further analyses that are then rendered practical, such as the estimation of group-averaged voxelwise connectivity, group-level parcellation, and group-ICA.

YNIMG Journal 2013 Journal Article

Decoding magnetoencephalographic rhythmic activity using spectrospatial information

  • Jukka-Pekka Kauppi
  • Lauri Parkkonen
  • Riitta Hari
  • Aapo Hyvärinen

We propose a new data-driven decoding method called Spectral Linear Discriminant Analysis (Spectral LDA) for the analysis of magnetoencephalography (MEG). The method allows investigation of changes in rhythmic neural activity as a result of different stimuli and tasks. The introduced classification model only assumes that each “brain state” can be characterized as a combination of neural sources, each of which shows rhythmic activity at one or several frequency bands. Furthermore, the model allows the oscillation frequencies to be different for each such state. We present decoding results from 9 subjects in a four-category classification problem defined by an experiment involving randomly alternating epochs of auditory, visual and tactile stimuli interspersed with rest periods. The performance of Spectral LDA was very competitive compared with four alternative classifiers based on different assumptions concerning the organization of rhythmic brain activity. In addition, the spectral and spatial patterns extracted automatically on the basis of trained classifiers showed that Spectral LDA offers a novel and interesting way of analyzing spectrospatial oscillatory neural activity across the brain. All the presented classification methods and visualization tools are freely available as a Matlab toolbox.

JMLR Journal 2013 Journal Article

Pairwise Likelihood Ratios for Estimation of Non-Gaussian Structural Equation Models

  • Aapo Hyvärinen
  • Stephen M. Smith

We present new measures of the causal direction, or direction of effect, between two non-Gaussian random variables. They are based on the likelihood ratio under the linear non-Gaussian acyclic model (LiNGAM). We also develop simple first-order approximations of the likelihood ratio and analyze them based on related cumulant-based measures, which can be shown to find the correct causal directions. We show how to apply these measures to estimate LiNGAM for more than two variables, and even in the case of more variables than observations. We further extend the method to cyclic and nonlinear models. The proposed framework is statistically at least as good as existing ones in the cases of few data points or noisy data, and it is computationally and conceptually very simple. Results on simulated fMRI data indicate that the method may be useful in neuroimaging where the number of time points is typically quite small. [abs] [ pdf ][ bib ] &copy JMLR 2013. ( edit, beta )

JMLR Journal 2012 Journal Article

Noise-Contrastive Estimation of Unnormalized Statistical Models, with Applications to Natural Image Statistics

  • Michael U. Gutmann
  • Aapo Hyvärinen

We consider the task of estimating, from observed data, a probabilistic model that is parameterized by a finite number of parameters. In particular, we are considering the situation where the model probability density function is unnormalized. That is, the model is only specified up to the partition function. The partition function normalizes a model so that it integrates to one for any choice of the parameters. However, it is often impossible to obtain it in closed form. Gibbs distributions, Markov and multi-layer networks are examples of models where analytical normalization is often impossible. Maximum likelihood estimation can then not be used without resorting to numerical approximations which are often computationally expensive. We propose here a new objective function for the estimation of both normalized and unnormalized models. The basic idea is to perform nonlinear logistic regression to discriminate between the observed data and some artificially generated noise. With this approach, the normalizing partition function can be estimated like any other parameter. We prove that the new estimation method leads to a consistent (convergent) estimator of the parameters. For large noise sample sizes, the new estimator is furthermore shown to behave like the maximum likelihood estimator. In the estimation of unnormalized models, there is a trade-off between statistical and computational performance. We show that the new method strikes a competitive trade-off in comparison to other estimation methods for unnormalized models. As an application to real data, we estimate novel two-layer models of natural image statistics with spline nonlinearities. [abs] [ pdf ][ bib ] &copy JMLR 2012. ( edit, beta )

JMLR Journal 2011 Journal Article

DirectLiNGAM: A Direct Method for Learning a Linear Non-Gaussian Structural Equation Model

  • Shohei Shimizu
  • Takanori Inazumi
  • Yasuhiro Sogawa
  • Aapo Hyvärinen
  • Yoshinobu Kawahara
  • Takashi Washio
  • Patrik O. Hoyer
  • Kenneth Bollen

Structural equation models and Bayesian networks have been widely used to analyze causal relations between continuous variables. In such frameworks, linear acyclic models are typically used to model the data-generating process of variables. Recently, it was shown that use of non-Gaussianity identifies the full structure of a linear acyclic model, that is, a causal ordering of variables and their connection strengths, without using any prior knowledge on the network structure, which is not the case with conventional methods. However, existing estimation methods are based on iterative search algorithms and may not converge to a correct solution in a finite number of steps. In this paper, we propose a new direct method to estimate a causal ordering and connection strengths based on non-Gaussianity. In contrast to the previous methods, our algorithm requires no algorithmic parameters and is guaranteed to converge to the right solution within a small fixed number of steps if the data strictly follows the model, that is, if all the model assumptions are met and the sample size is infinite. [abs] [ pdf ][ bib ] &copy JMLR 2011. ( edit, beta )

NeurIPS Conference 2011 Conference Paper

Structural equations and divisive normalization for energy-dependent component analysis

  • Jun-ichiro Hirayama
  • Aapo Hyvärinen

Components estimated by independent component analysis and related methods are typically not independent in real data. A very common form of nonlinear dependency between the components is correlations in their variances or ener- gies. Here, we propose a principled probabilistic model to model the energy- correlations between the latent variables. Our two-stage model includes a linear mixing of latent signals into the observed ones like in ICA. The main new fea- ture is a model of the energy-correlations based on the structural equation model (SEM), in particular, a Linear Non-Gaussian SEM. The SEM is closely related to divisive normalization which effectively reduces energy correlation. Our new two- stage model enables estimation of both the linear mixing and the interactions re- lated to energy-correlations, without resorting to approximations of the likelihood function or other non-principled approaches. We demonstrate the applicability of our method with synthetic dataset, natural images and brain signals.

YNIMG Journal 2011 Journal Article

Testing the ICA mixing matrix based on inter-subject or inter-session consistency

  • Aapo Hyvärinen

Independent component analysis (ICA) is increasingly used for analyzing brain imaging data. ICA typically gives a large number of components, many of which may be just random, due to insufficient sample size, violations of the model, or algorithmic problems. Few methods are available for computing the statistical significance (reliability) of the components. We propose to approach this problem by performing ICA separately on a number of subjects, and finding components which are sufficiently consistent (similar) over subjects. Similarity is defined here as the similarity of the mixing coefficients, which usually correspond to spatial patterns in EEG and MEG. The threshold of what is “sufficient” is rigorously defined by a null hypothesis under which the independent components are random orthogonal components in the whitened space. Components which are consistent in different subjects are found by clustering under the constraint that a cluster can only contain one source from each subject, and by constraining the number of the false positives based on the null hypothesis. Instead of different subjects, the method can also be applied on different recording sessions from a single subject. The testing method is particularly applicable to EEG and MEG analysis.

UAI Conference 2010 Conference Paper

A Family of Computationally E cient and Simple Estimators for Unnormalized Statistical Models

  • Miika Pihlaja
  • Michael Gutmann
  • Aapo Hyvärinen

We introduce a new family of estimators for unnormalized statistical models. Our family of estimators is parameterized by two nonlinear functions and uses a single sample from an auxiliary distribution, generalizing Maximum Likelihood Monte Carlo estimation of Geyer and Thompson (1992). The family is such that we can estimate the partition function like any other parameter in the model. The estimation is done by optimizing an algebraically simple, well defined objective function, which allows for the use of dedicated optimization methods. We establish consistency of the estimator family and give an expression for the asymptotic covariance matrix, which enables us to further analyze the influence of the nonlinearities and the auxiliary density on estimation performance. Some estimators in our family are particularly stable for a wide range of auxiliary densities. Interestingly, a specific choice of the nonlinearity establishes a connection between density estimation and classification by nonlinear logistic regression. Finally, the optimal amount of auxiliary samples relative to the given amount of the data is considered from the perspective of computational efficiency.

JMLR Journal 2010 Journal Article

Estimation of a Structural Vector Autoregression Model Using Non-Gaussianity

  • Aapo Hyvärinen
  • Kun Zhang
  • Shohei Shimizu
  • Patrik O. Hoyer

Analysis of causal effects between continuous-valued variables typically uses either autoregressive models or structural equation models with instantaneous effects. Estimation of Gaussian, linear structural equation models poses serious identifiability problems, which is why it was recently proposed to use non-Gaussian models. Here, we show how to combine the non-Gaussian instantaneous model with autoregressive models. This is effectively what is called a structural vector autoregression (SVAR) model, and thus our work contributes to the long-standing problem of how to estimate SVAR's. We show that such a non-Gaussian model is identifiable without prior knowledge of network structure. We propose computationally efficient methods for estimating the model, as well as methods to assess the significance of the causal influences. The model is successfully applied on financial and brain imaging data. [abs] [ pdf ][ bib ] &copy JMLR 2010. ( edit, beta )

YNIMG Journal 2010 Journal Article

Independent component analysis of short-time Fourier transforms for spontaneous EEG/MEG analysis

  • Aapo Hyvärinen
  • Pavan Ramkumar
  • Lauri Parkkonen
  • Riitta Hari

Analysis of spontaneous EEG/MEG needs unsupervised learning methods. While independent component analysis (ICA) has been successfully applied on spontaneous fMRI, it seems to be too sensitive to technical artifacts in EEG/MEG. We propose to apply ICA on short-time Fourier transforms of EEG/MEG signals, in order to find more “interesting” sources than with time-domain ICA, and to more meaningfully sort the obtained components. The method is especially useful for finding sources of rhythmic activity. Furthermore, we propose to use a complex mixing matrix to model sources which are spatially extended and have different phases in different EEG/MEG channels. Simulations with artificial data and experiments on resting-state MEG demonstrate the utility of the method.

UAI Conference 2010 Conference Paper

Source Separation and Higher-Order Causal Analysis of MEG and EEG

  • Kun Zhang 0001
  • Aapo Hyvärinen

Separation of the sources and analysis of their connectivity have been an important topic in EEG/MEG analysis. To solve this problem in an automatic manner, we propose a twolayer model, in which the sources are conditionally uncorrelated from each other, but not independent; the dependence is caused by the causality in their time-varying variances (envelopes). The model is identified in two steps. We first propose a new source separation technique which takes into account the autocorrelations (which may be time-varying) and time-varying variances of the sources. The causality in the envelopes is then discovered by exploiting a special kind of multivariate GARCH (generalized autoregressive conditional heteroscedasticity) model. The resulting causal diagram gives the effective connectivity between the separated sources; in our experimental results on MEG data, sources with similar functions are grouped together, with negative influences between groups, and the groups are connected via some interesting sources.

UAI Conference 2009 Conference Paper

A direct method for estimating a causal ordering in a linear non-Gaussian acyclic model

  • Shohei Shimizu
  • Aapo Hyvärinen
  • Yoshinobu Kawahara

Structural equation models and Bayesian networks have been widely used to analyze causal relations between continuous variables. In such frameworks, linear acyclic models are typically used to model the datagenerating process of variables. Recently, it was shown that use of non-Gaussianity identifies a causal ordering of variables in a linear acyclic model without using any prior knowledge on the network structure, which is not the case with conventional methods. However, existing estimation methods are based on iterative search algorithms and may not converge to a correct solution in a finite number of steps. In this paper, we propose a new direct method to estimate a causal ordering based on non-Gaussianity. In contrast to the previous methods, our algorithm requires no algorithmic parameters and is guaranteed to converge to the right solution within a small fixed number of steps if the data strictly follows the model.

UAI Conference 2009 Conference Paper

On the Identifiability of the Post-Nonlinear Causal Model

  • Kun Zhang 0001
  • Aapo Hyvärinen

By taking into account the nonlinear effect of the cause, the inner noise effect, and the measurement distortion effect in the observed variables, the post-nonlinear (PNL) causal model has demonstrated its excellent performance in distinguishing the cause from effect. However, its identifiability has not been properly addressed, and how to apply it in the case of more than two variables is also a problem. In this paper, we conduct a systematic investigation on its identifiability in the two-variable case. We show that this model is identifiable in most cases; by enumerating all possible situations in which the model is not identifiable, we provide sufficient conditions for its identifiability. Simulations are given to support the theoretical results. Moreover, in the case of more than two variables, we show that the whole causal structure can be found by applying the PNL causal model to each structure in the Markov equivalent class and testing if the disturbance is independent of the direct causes for each variable. In this way the exhaustive search over all possible causal structures is avoided.

UAI Conference 2008 Conference Paper

Causal discovery of linear acyclic models with arbitrary distributions

  • Patrik O. Hoyer
  • Aapo Hyvärinen
  • Richard Scheines
  • Peter Spirtes
  • Joseph D. Ramsey
  • Gustavo Lacerda
  • Shohei Shimizu

An important task in data analysis is the discovery of causal relationships between observed variables. For continuous-valued data, linear acyclic causal models are commonly used to model the data-generating process, and the inference of such models is a wellstudied problem. However, existing methods have significant limitations. Methods based on conditional independencies (Spirtes et al. 1993; Pearl 2000) cannot distinguish between independence-equivalent models, whereas approaches purely based on Independent Component Analysis (Shimizu et al. 2006) are inapplicable to data which is partially Gaussian. In this paper, we generalize and combine the two approaches, to yield a method able to learn the model structure in many cases for which the previous methods provide answers that are either incorrect or are not as informative as possible. We give exact graphical conditions for when two distinct models represent the same family of distributions, and empirically demonstrate the power of our method through thorough simulations.

ICML Conference 2008 Conference Paper

Causal modelling combining instantaneous and lagged effects: an identifiable model based on non-Gaussianity

  • Aapo Hyvärinen
  • Shohei Shimizu
  • Patrik O. Hoyer

Causal analysis of continuous-valued variables typically uses either autoregressive models or linear Gaussian Bayesian networks with instantaneous effects. Estimation of Gaussian Bayesian networks poses serious identifiability problems, which is why it was recently proposed to use non-Gaussian models. Here, we show how to combine the non-Gaussian instantaneous model with autoregressive models. We show that such a non-Gaussian model is identifiable without prior knowledge of network structure, and we propose an estimation method shown to be consistent. This approach also points out how neglecting instantaneous effects can lead to completely wrong estimates of the autoregressive coefficients.

JMLR Journal 2006 Journal Article

A Linear Non-Gaussian Acyclic Model for Causal Discovery

  • Shohei Shimizu
  • Patrik O. Hoyer
  • Aapo Hyvärinen
  • Antti Kerminen

In recent years, several methods have been proposed for the discovery of causal structure from non-experimental data. Such methods make various assumptions on the data generating process to facilitate its identification from purely observational data. Continuing this line of research, we show how to discover the complete causal structure of continuous-valued data, under the assumptions that (a) the data generating process is linear, (b) there are no unobserved confounders, and (c) disturbance variables have non-Gaussian distributions of non-zero variances. The solution relies on the use of the statistical method known as independent component analysis, and does not require any pre-specified time-ordering of the variables. We provide a complete Matlab package for performing this LiNGAM analysis (short for Linear Non-Gaussian Acyclic Model), and demonstrate the effectiveness of the method using artificially generated data and real-world data. [abs] [ pdf ][ bib ] &copy JMLR 2006. ( edit, beta )

NeurIPS Conference 2006 Conference Paper

Emergence of conjunctive visual features by quadratic independent component analysis

  • J. t. Lindgren
  • Aapo Hyvärinen

In previous studies, quadratic modelling of natural images has resulted in cell models that react strongly to edges and bars. Here we apply quadratic Independent Component Analysis to natural image patches, and show that up to a small approximation error, the estimated components are computing conjunctions of two linear features. These conjunctive features appear to represent not only edges and bars, but also inherently two-dimensional stimuli, such as corners. In addition, we show that for many of the components, the underlying linear features have essentially V1 simple cell receptive field characteristics. Our results indicate that the development of the V2 cells preferring angles and corners may be partly explainable by the principle of unsupervised sparse coding of natural images.

UAI Conference 2005 Conference Paper

Discovery of Non-gaussian Linear Causal Models using ICA

  • Shohei Shimizu
  • Aapo Hyvärinen
  • Yutaka Kano
  • Patrik O. Hoyer

In recent years, several methods have been proposed for the discovery of causal structure from non-experimental data (Spirtes et al. 2000; Pearl 2000). Such methods make various assumptions on the data generating process to facilitate its identification from purely observational data. Continuing this line of research, we show how to discover the complete causal structure of continuous-valued data, under the assumptions that (a) the data generating process is linear, (b) there are no unobserved confounders, and (c) disturbance variables have non-gaussian distributions of non-zero variances. The solution relies on the use of the statistical method known as independent component analysis (ICA), and does not require any pre-specified time-ordering of the variables. We provide a complete Matlab package for performing this LiNGAM analysis (short for Linear Non-Gaussian Acyclic Model), and demonstrate the effectiveness of the method using artificially generated data.

JMLR Journal 2005 Journal Article

Estimation of Non-Normalized Statistical Models by Score Matching

  • Aapo Hyvärinen

One often wants to estimate statistical models where the probability density function is known only up to a multiplicative normalization constant. Typically, one then has to resort to Markov Chain Monte Carlo methods, or approximations of the normalization constant. Here, we propose that such models can be estimated by minimizing the expected squared distance between the gradient of the log-density given by the model and the gradient of the log-density of the observed data. While the estimation of the gradient of log-density function is, in principle, a very difficult non-parametric problem, we prove a surprising result that gives a simple formula for this objective function. The density function of the observed data does not appear in this formula, which simplifies to a sample average of a sum of some derivatives of the log-density given by the model. The validity of the method is demonstrated on multivariate Gaussian and independent component analysis models, and by estimating an overcomplete filter set for natural image data. [abs] [ pdf ][ bib ] &copy JMLR 2005. ( edit, beta )

YNIMG Journal 2004 Journal Article

Validating the independent components of neuroimaging time series via clustering and visualization

  • Johan Himberg
  • Aapo Hyvärinen
  • Fabrizio Esposito

Recently, independent component analysis (ICA) has been widely used in the analysis of brain imaging data. An important problem with most ICA algorithms is, however, that they are stochastic; that is, their results may be somewhat different in different runs of the algorithm. Thus, the outputs of a single run of an ICA algorithm should be interpreted with some reserve, and further analysis of the algorithmic reliability of the components is needed. Moreover, as with any statistical method, the results are affected by the random sampling of the data, and some analysis of the statistical significance or reliability should be done as well. Here we present a method for assessing both the algorithmic and statistical reliability of estimated independent components. The method is based on running the ICA algorithm many times with slightly different conditions and visualizing the clustering structure of the obtained components in the signal space. In experiments with magnetoencephalographic (MEG) and functional magnetic resonance imaging (fMRI) data, the method was able to show that expected components are reliable; furthermore, it pointed out components whose interpretation was not obvious but whose reliability should incite the experimenter to investigate the underlying technical or physical phenomena. The method is implemented in a software package called Icasso.

YNIMG Journal 2003 Journal Article

Independent component analysis of nondeterministic fMRI signal sources

  • Vesa Kiviniemi
  • Juha-Heikki Kantola
  • Jukka Jauhiainen
  • Aapo Hyvärinen
  • Osmo Tervonen

Neuronal activation can be separated from other signal sources of functional magnetic resonance imaging (fMRI) data by using independent component analysis (ICA). Without deliberate neuronal activity of the brain cortex, the fMRI signal is a stochastic sum of various physiological and artifact related signal sources. The ability of spatial-domain ICA to separate spontaneous physiological signal sources was evaluated in 15 anesthetized children known to present prominent vasomotor fluctuations in the functional cortices. ICA separated multiple clustered signal sources in the primary sensory areas in all of the subjects. The spatial distribution and frequency spectra of the signal sources correspond to the known properties of 0. 03-Hz very-low-frequency vasomotor waves in fMRI data. In addition, ICA was able to separate major artery and sagittal sinus related signal sources in each subject. The characteristics of the blood vessel related signal sources were different from the parenchyma sources. ICA analysis of fMRI can be used for both assessing the statistical independence of brain signals and segmenting nondeterministic signal sources for further analysis.

NeurIPS Conference 2002 Conference Paper

Interpreting Neural Response Variability as Monte Carlo Sampling of the Posterior

  • Patrik Hoyer
  • Aapo Hyvärinen

The responses of cortical sensory neurons are notoriously variable, with the number of spikes evoked by identical stimuli varying significantly from trial to trial. This variability is most often interpreted as ‘noise’, purely detrimental to the sensory system. In this paper, we propose an al- ternative view in which the variability is related to the uncertainty, about world parameters, which is inherent in the sensory stimulus. Specifi- cally, the responses of a population of neurons are interpreted as stochas- tic samples from the posterior distribution in a latent variable model. In addition to giving theoretical arguments supporting such a representa- tional scheme, we provide simulations suggesting how some aspects of response variability might be understood in this framework.

NeurIPS Conference 2002 Conference Paper

Temporal Coherence, Natural Image Sequences, and the Visual Cortex

  • Jarmo Hurri
  • Aapo Hyvärinen

We show that two important properties of the primary visual cortex emerge when the principle of temporal coherence is applied to natural image sequences. The properties are simple-cell-like receptive fields and complex-cell-like pooling of simple cell outputs, which emerge when we apply two different approaches to temporal coherence. In the first approach we extract receptive fields whose outputs are as temporally co- herent as possible. This approach yields simple-cell-like receptive fields (oriented, localized, multiscale). Thus, temporal coherence is an alterna- tive to sparse coding in modeling the emergence of simple cell receptive fields. The second approach is based on a two-layer statistical generative model of natural image sequences. In addition to modeling the temporal coherence of individual simple cells, this model includes inter-cell tem- poral dependencies. Estimation of this model from natural data yields both simple-cell-like receptive fields, and complex-cell-like pooling of simple cell outputs. In this completely unsupervised learning, both lay- ers of the generative model are estimated simultaneously from scratch. This is a significant improvement on earlier statistical models of early vision, where only one layer has been learned, and others have been fixed a priori.

NeurIPS Conference 1999 Conference Paper

Emergence of Topography and Complex Cell Properties from Natural Images using Extensions of ICA

  • Aapo Hyvärinen
  • Patrik Hoyer

Independent component analysis of natural images leads to emer(cid: 173) gence of simple cell properties, Le. linear filters that resemble wavelets or Gabor functions. In this paper, we extend ICA to explain further properties of VI cells. First, we decompose natural images into independent subspaces instead of scalar components. This model leads to emergence of phase and shift invariant fea(cid: 173) tures, similar to those in VI complex cells. Second, we define a topography between the linear components obtained by ICA. The topographic distance between two components is defined by their higher-order correlations, so that two components are close to each other in the topography if they are strongly dependent on each other. This leads to simultaneous emergence of both topography and invariances similar to complex cell properties.

NeurIPS Conference 1998 Conference Paper

Sparse Code Shrinkage: Denoising by Nonlinear Maximum Likelihood Estimation

  • Aapo Hyvärinen
  • Patrik Hoyer
  • Erkki Oja

Sparse coding is a method for finding a representation of data in which each of the components of the representation is only rarely significantly active. Such a representation is closely related to re(cid: 173) dundancy reduction and independent component analysis, and has some neurophysiological plausibility. In this paper, we show how sparse coding can be used for denoising. Using maximum likelihood estimation of nongaussian variables corrupted by gaussian noise, we show how to apply a shrinkage nonlinearity on the components of sparse coding so as to reduce noise. Furthermore, we show how to choose the optimal sparse coding basis for denoising. Our method is closely related to the method of wavelet shrinkage, but has the important benefit over wavelet methods that both the features and the shrinkage parameters are estimated directly from the data.

NeurIPS Conference 1997 Conference Paper

New Approximations of Differential Entropy for Independent Component Analysis and Projection Pursuit

  • Aapo Hyvärinen

We derive a first-order approximation of the density of maximum entropy for a continuous 1-D random variable, given a number of simple constraints. This results in a density expansion which is somewhat similar to the classical polynomial density expansions by Gram-Charlier and Edgeworth. Using this approximation of density, an approximation of 1-D differential entropy is derived. The approximation of entropy is both more exact and more ro(cid: 173) bust against outliers than the classical approximation based on the polynomial density expansions, without being computationally more expensive. The approximation has applications, for example, in independent component analysis and projection pursuit.

NeurIPS Conference 1996 Conference Paper

One-unit Learning Rules for Independent Component Analysis

  • Aapo Hyvärinen
  • Erkki Oja

Neural one-unit learning rules for the problem of Independent Com(cid: 173) ponent Analysis (ICA) and blind source separation are introduced. In these new algorithms, every ICA neuron develops into a sepa(cid: 173) rator that finds one of the independent components. The learning rules use very simple constrained Hebbianjanti-Hebbian learning in which decorrelating feedback may be added. To speed up the convergence of these stochastic gradient descent rules, a novel com(cid: 173) putationally efficient fixed-point algorithm is introduced.

v2026.09.13