Arrow Research search

Author name cluster

Guillermo Sapiro

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

44 papers
2 author rows

Possible papers

44

ICML Conference 2025 Conference Paper

Addressing Misspecification in Simulation-based Inference through Data-driven Calibration

  • Antoine Wehenkel
  • Juan L. Gamella
  • Ozan Sener
  • Jens Behrmann
  • Guillermo Sapiro
  • Jörn-Henrik Jacobsen
  • Marco Cuturi

Driven by steady progress in deep generative modeling, simulation-based inference (SBI) has emerged as the workhorse for inferring the parameters of stochastic simulators. However, recent work has demonstrated that model misspecification can harm SBI’s reliability, preventing its adoption in important applications where only misspecified simulators are available. This work introduces robust posterior estimation (RoPE), a framework that overcomes model misspecification with a small real-world calibration set of ground truth parameter measurements. We formalize the misspecification gap as the solution of an optimal transport (OT) problem between learned representations of real-world and simulated observations, allowing RoPE to learn a model of the misspecification without placing additional assumptions on its nature. RoPE shows how the calibration set and OT together offer a controllable balance between calibrated uncertainty and informative inference even under severely misspecified simulators. Results on four synthetic tasks and two real-world problems with ground-truth labels demonstrate that RoPE outperforms baselines and consistently returns informative and calibrated credible intervals.

TMLR Journal 2025 Journal Article

Autoregressive Models in Vision: A Survey

  • Jing Xiong
  • Gongye Liu
  • Lun Huang
  • Chengyue Wu
  • Taiqiang Wu
  • Yao Mu
  • Yuan Yao
  • Hui Shen

Autoregressive modeling has been a huge success in the field of natural language processing (NLP). Recently, autoregressive models have emerged as a significant area of focus in computer vision, where they excel in producing high-quality visual content. Autoregressive models in NLP typically operate on subword tokens. However, the representation strategy in computer vision can vary in different levels, i.e., pixel-level, token-level, or scale-level, reflecting the diverse and hierarchical nature of visual data compared to the sequential structure of language. This survey comprehensively examines the literature on autoregressive models applied to vision. To improve readability for researchers from diverse research backgrounds, we start with preliminary sequence representation and modeling in vision. Next, we divide the fundamental frameworks of visual autoregressive models into three general sub-categories, including pixel-based, token-based, and scale-based models based on the representation strategy. We then explore the interconnections between autoregressive models and other generative models. Furthermore, we present a multifaceted categorization of autoregressive models in computer vision, including image generation, video generation, 3D generation, and multimodal generation. We also elaborate on their applications in diverse domains, including emerging domains such as embodied AI and 3D medical AI, with about 250 related references. Finally, we highlight the current challenges to autoregressive models in vision with suggestions about potential research directions. We have also set up a Github repository to organize the papers included in this survey at: https://github.com/ChaofanTao/Autoregressive-Models-in-Vision-Survey.

TMLR Journal 2025 Journal Article

FoldDiff: Folding in Point Cloud Diffusion

  • Yuzhou Zhao
  • Juan Matias Di Martino
  • Amirhossein Farzam
  • Guillermo Sapiro

Diffusion denoising has emerged as a powerful approach for modeling data distributions, treating data as particles with their position and velocity modeled by a stochastic diffusion process. While this framework assumes data resides in a fixed vector spaces (e.g., images as pixel-ordered vectors), point clouds present unique challenges due to their unordered representation. Existing point cloud diffusion methods often rely on voxelization to address this issue, but this approach is computationally expensive, with cubically scaling complexity. In this work, we investigate the misalignment between point cloud irregularity and diffusion models, analyzing it through the lens of denoising implicit priors. First, we demonstrate how the unknown permutations inherent in point cloud structures disrupt denoising implicit priors. To address this, we then propose a novel folding-based approach that reorders point clouds into a permutation-invariant grid, enabling diffusion to be performed directly on the structured representation. This construction is exploited both globally and locally. Globally, \reviewcdmS{folded objects can represent point cloud objects} in a fixed vector space (like images), therefore it enables us to extend the work of denoising as implicit priors to point clouds. \reviewcdmS{Locally, the folded tokens are} efficient and novel token representations that can improve existing transformer-based point cloud diffusion models. Our experiments show that the proposed folding operation integrates effectively with both denoising implicit priors as well as advanced diffusion architectures, such as UNet and Diffusion Transformers (DiTs). Notably, DiT with \reviewcdmS{locally} folded tokens achieves competitive generative performance compared to state-of-the-art models while significantly reducing training and inference costs relative to voxelization-based methods.

ICLR Conference 2025 Conference Paper

SSOLE: Rethinking Orthogonal Low-rank Embedding for Self-Supervised Learning

  • Lun Huang
  • Qiang Qiu 0001
  • Guillermo Sapiro

Self-supervised learning (SSL) aims to learn meaningful representations from unlabeled data. Orthogonal Low-rank Embedding (OLE) shows promise for SSL by enhancing intra-class similarity in a low-rank subspace and promoting inter-class dissimilarity in a high-rank subspace, making it particularly suitable for multi-view learning tasks. However, directly applying OLE to SSL poses significant challenges: (1) the virtually infinite number of "classes" in SSL makes achieving the OLE objective impractical, leading to representational collapse; and (2) low-rank constraints may fail to distinguish between positively and negatively correlated features, further undermining learning. To address these issues, we propose SSOLE (Self-Supervised Orthogonal Low-rank Embedding), a novel framework that integrates OLE principles into SSL by (1) decoupling the low-rank and high-rank enforcement to align with SSL objectives; and (2) applying low-rank constraints to feature deviations from their mean, ensuring better alignment of positive pairs by accounting for the signs of cosine similarities. Our theoretical analysis and empirical results demonstrate that these adaptations are crucial to SSOLE’s effectiveness. Moreover, SSOLE achieves competitive performance across SSL benchmarks without relying on large batch sizes, memory banks, or dual-encoder architectures, making it an efficient and scalable solution for self-supervised tasks. Code is available at https://github.com/husthuaan/ssole.

ICML Conference 2024 Conference Paper

From Geometry to Causality- Ricci Curvature and the Reliability of Causal Inference on Networks

  • Amirhossein Farzam
  • Allen R. Tannenbaum
  • Guillermo Sapiro

Causal inference on networks faces challenges posed in part by violations of standard identification assumptions due to dependencies between treatment units. Although graph geometry fundamentally influences such dependencies, the potential of geometric tools for causal inference on networked treatment units is yet to be unlocked. Moreover, despite significant progress utilizing graph neural networks (GNNs) for causal inference on networks, methods for evaluating their achievable reliability without ground truth are lacking. In this work we establish for the first time a theoretical link between network geometry, the graph Ricci curvature in particular, and causal inference, formalizing the intrinsic challenges that negative curvature poses to estimating causal parameters. The Ricci curvature can then be used to assess the reliability of causal estimates in structured data, as we empirically demonstrate. Informed by this finding, we propose a method using the geometric Ricci flow to reduce causal effect estimation error in networked data, showcasing how this newfound connection between graph geometry and causal inference could improve GNN-based causal inference. Bridging graph geometry and causal inference, this paper opens the door to geometric techniques for improving causal estimation on networks.

TMLR Journal 2024 Journal Article

Generalizing Neural Additive Models via Statistical Multimodal Analysis

  • Young Kyung Kim
  • Juan Matias Di Martino
  • Guillermo Sapiro

Interpretable models are gaining increasing attention in the machine learning community, and significant progress is being made to develop simple, interpretable, yet powerful deep learning approaches. Generalized Additive Models (GAM) and Neural Additive Models (NAM) are prime examples. Despite these methods' great potential and popularity in critical applications, e.g., medical applications, they fail to generalize to distributions with more than one mode (multimodal\footnote{In this paper, multimodal refers to the context of distributions, wherein a distribution possesses more than one mode.}). The main reason behind this limitation is that these "all-fit-one" models collapse multiple relationships by being forced to fit the data unimodally. We address this critical limitation by proposing interpretable multimodal network frameworks capable of learning a Mixture of Neural Additive Models (MNAM). The proposed MNAM learns relationships between input features and outputs in a multimodal fashion and assigns a probability to each mode. The proposed method shares similarities with Mixture Density Networks (MDN) while keeping the interpretability that characterizes GAM and NAM. We demonstrate how the proposed MNAM balances between rich representations and interpretability with numerous empirical observations and pedagogical studies. We present and discuss different training alternatives and provided extensive practical evaluation to assess the proposed framework. The code is available at \href{https://github.com/youngkyungkim93/MNAM}{https://github.com/youngkyungkim93/MNAM}.

TMLR Journal 2023 Journal Article

Consistent Collaborative Filtering via Tensor Decomposition

  • Shiwen Zhao
  • Guillermo Sapiro

Collaborative filtering is the de facto standard for analyzing users’ activities and building recommendation systems for items. In this work we develop Sliced Anti-symmetric Decomposition (SAD), a new model for collaborative filtering based on implicit feedback. In contrast to traditional techniques where a latent representation of users (user vectors) and items (item vectors) are estimated, SAD introduces one additional latent vector to each item, using a novel three-way tensor view of user-item interactions. This new vector extends user-item preferences calculated by standard dot products to general inner products, producing interactions between items when evaluating their relative preferences. SAD reduces to state-of-the-art (SOTA) collaborative filtering models when the vector collapses to 1, while in this paper we allow its value to be estimated from data. Allowing the values of the new item vector to be different from 1 has profound implications. It suggests users may have nonlinear mental models when evaluating items, allowing the existence of cycles in pairwise comparisons. We demonstrate the efficiency of SAD in both simulated and real world datasets containing over 1M user-item interactions. By comparing with seven SOTA collaborative filtering models with implicit feedbacks, SAD produces the most consistent personalized preferences, in the meanwhile maintaining top-level of accuracy in personalized recommendations. We release the model and inference algorithms in a Python library https://github.com/apple/ml-sad.

TMLR Journal 2023 Journal Article

Robust Hybrid Learning With Expert Augmentation

  • Antoine Wehenkel
  • Jens Behrmann
  • Hsiang Hsu
  • Guillermo Sapiro
  • Gilles Louppe
  • Joern-Henrik Jacobsen

Hybrid modelling reduces the misspecification of expert models by combining them with machine learning (ML) components learned from data. Similarly to many ML algorithms, hybrid model performance guarantees are limited to the training distribution. Leveraging the insight that the expert model is usually valid even outside the training domain, we overcome this limitation by introducing a hybrid data augmentation strategy termed \textit{expert augmentation}. Based on a probabilistic formalization of hybrid modelling, we demonstrate that expert augmentation, which can be incorporated into existing hybrid systems, improves generalization. We empirically validate the expert augmentation on three controlled experiments modelling dynamical systems with ordinary and partial differential equations. Finally, we assess the potential real-world applicability of expert augmentation on a dataset of a real double pendulum.

JMLR Journal 2022 Journal Article

Scaling-Translation-Equivariant Networks with Decomposed Convolutional Filters

  • Wei Zhu
  • Qiang Qiu
  • Robert Calderbank
  • Guillermo Sapiro
  • Xiuyuan Cheng

Encoding the scale information explicitly into the representation learned by a convolutional neural network (CNN) is beneficial for many computer vision tasks especially when dealing with multiscale inputs. We study, in this paper, a scaling-translation-equivariant ($\mathcal{ST}$-equivariant) CNN with joint convolutions across the space and the scaling group, which is shown to be both sufficient and necessary to achieve equivariance for the regular representation of the scaling-translation group $\mathcal{ST}$. To reduce the model complexity and computational burden, we decompose the convolutional filters under two pre-fixed separable bases and truncate the expansion to low-frequency components. A further benefit of the truncated filter expansion is the improved deformation robustness of the equivariant representation, a property which is theoretically analyzed and empirically verified. Numerical experiments demonstrate that the proposed scaling-translation-equivariant network with decomposed convolutional filters (ScDCFNet) achieves significantly improved performance in multiscale image classification and better interpretability than regular CNNs at a reduced model size. [abs] [ pdf ][ bib ] &copy JMLR 2022. ( edit, beta )

ICML Conference 2021 Conference Paper

Blind Pareto Fairness and Subgroup Robustness

  • Natalia Martínez
  • Martín Bertrán
  • Afroditi Papadaki
  • Miguel Rodrigues 0001
  • Guillermo Sapiro

Much of the work in the field of group fairness addresses disparities between predefined groups based on protected features such as gender, age, and race, which need to be available at train, and often also at test, time. These approaches are static and retrospective, since algorithms designed to protect groups identified a priori cannot anticipate and protect the needs of different at-risk groups in the future. In this work we analyze the space of solutions for worst-case fairness beyond demographics, and propose Blind Pareto Fairness (BPF), a method that leverages no-regret dynamics to recover a fair minimax classifier that reduces worst-case risk of any potential subgroup of sufficient size, and guarantees that the remaining population receives the best possible level of service. BPF addresses fairness beyond demographics, that is, it does not rely on predefined notions of at-risk groups, neither at train nor at test time. Our experimental results show that the proposed framework improves worst-case risk in multiple standard datasets, while simultaneously providing better levels of service for the remaining population. The code is available at github. com/natalialmg/BlindParetoFairness

ICRA Conference 2021 Conference Paper

Cirrus: A Long-range Bi-pattern LiDAR Dataset

  • Ze Wang 0008
  • Sihao Ding 0002
  • Ying Li 0139
  • Jonas Fenn
  • Sohini Roychowdhury
  • Andreas Wallin
  • Lane Martin
  • Scott Ryvola

In this paper, we introduce Cirrus, a new long-range bi-pattern LiDAR public dataset for autonomous driving tasks such as 3D object detection, critical to highway driving and timely decision making. Our platform is equipped with a high-resolution video camera and a pair of LiDAR sensors with a 250-meter effective range, which is significantly longer than existing public datasets. We record paired point clouds simultaneously using both Gaussian and uniform scanning patterns. Point density varies significantly across such a long range, and different scanning patterns further diversify object representation in LiDAR. In Cirrus, eight categories of objects are exhaustively annotated in the LiDAR point clouds for the entire effective range. To illustrate the kind of studies supported by this new dataset, we introduce LiDAR model adaptation across different ranges, scanning patterns, and sensor devices. Promising results show the great potential of this new dataset to the robotics and computer vision communities.

NeurIPS Conference 2020 Conference Paper

A Dictionary Approach to Domain-Invariant Learning in Deep Networks

  • Ze Wang
  • Xiuyuan Cheng
  • Guillermo Sapiro
  • Qiang Qiu

In this paper, we consider domain-invariant deep learning by explicitly modeling domain shifts with only a small amount of domain-specific parameters in a Convolutional Neural Network (CNN). By exploiting the observation that a convolutional filter can be well approximated as a linear combination of a small set of dictionary atoms, we show for the first time, both empirically and theoretically, that domain shifts can be effectively handled by decomposing a convolutional layer into a domain-specific atom layer and a domain-shared coefficient layer, while both remain convolutional. An input channel will now first convolve spatially only with each respective domain-specific dictionary atom to ``absorb" domain variations, and then output channels are linearly combined using common decomposition coefficients trained to promote shared semantics across domains. We use toy examples, rigorous analysis, and real-world examples with diverse datasets and architectures, to show the proposed plug-in framework's effectiveness in cross and joint domain performance and domain adaptation. With the proposed architecture, we need only a small set of dictionary atoms to model each additional domain, which brings a negligible amount of additional parameters, typically a few hundred.

NeurIPS Conference 2020 Conference Paper

Instance-based Generalization in Reinforcement Learning

  • Martin Bertran
  • Natalia Martinez
  • Mariano Phielipp
  • Guillermo Sapiro

Agents trained via deep reinforcement learning (RL) routinely fail to generalize to unseen environments, even when these share the same underlying dynamics as the training levels. Understanding the generalization properties of RL is one of the challenges of modern machine learning. Towards this goal, we analyze policy learning in the context of Partially Observable Markov Decision Processes (POMDPs) and formalize the dynamics of training levels as instances. We prove that, independently of the exploration strategy, reusing instances introduces significant changes on the effective Markov dynamics the agent observes during training. Maximizing expected rewards impacts the learned belief state of the agent by inducing undesired instance-specific speed-running policies instead of generalizable ones, which are sub-optimal on the training set. We provide generalization bounds to the value gap in train and test environments based on the number of training instances, and use insights based on these to improve performance on unseen levels. We propose training a shared belief representation over an ensemble of specialized policies, from which we compute a consensus policy that is used for data collection, disallowing instance-specific exploitation. We experimentally validate our theory, observations, and the proposed computational solution over the CoinRun benchmark.

ICML Conference 2020 Conference Paper

Minimax Pareto Fairness: A Multi Objective Perspective

  • Natalia Martínez
  • Martín Bertrán
  • Guillermo Sapiro

In this work we formulate and formally characterize group fairness as a multi-objective optimization problem, where each sensitive group risk is a separate objective. We propose a fairness criterion where a classifier achieves minimax risk and is Pareto-efficient w. r. t. all groups, avoiding unnecessary harm, and can lead to the best zero-gap model if policy dictates so. We provide a simple optimization algorithm compatible with deep neural networks to satisfy these constraints. Since our method does not require test-time access to sensitive attributes, it can be applied to reduce worst-case classification errors between outcomes in unbalanced classification problems. We test the proposed methodology on real case-studies of predicting income, ICU patient mortality, skin lesions classification, and assessing credit risk, demonstrating how our framework compares favorably to other approaches.

ICLR Conference 2020 Conference Paper

Stochastic Conditional Generative Networks with Basis Decomposition

  • Ze Wang 0008
  • Xiuyuan Cheng
  • Guillermo Sapiro
  • Qiang Qiu 0001

While generative adversarial networks (GANs) have revolutionized machine learning, a number of open questions remain to fully understand them and exploit their power. One of these questions is how to efficiently achieve proper diversity and sampling of the multi-mode data space. To address this, we introduce BasisGAN, a stochastic conditional multi-mode image generator. By exploiting the observation that a convolutional filter can be well approximated as a linear combination of a small set of basis elements, we learn a plug-and-played basis generator to stochastically generate basis elements, with just a few hundred of parameters, to fully embed stochasticity into convolutional filters. By sampling basis elements instead of filters, we dramatically reduce the cost of modeling the parameter space with no sacrifice on either image diversity or fidelity. To illustrate this proposed plug-and-play framework, we construct variants of BasisGAN based on state-of-the-art conditional image generation networks, and train the networks by simply plugging in a basis generator, without additional auxiliary components, hyperparameters, or training objectives. The experimental success is complemented with theoretical results indicating how the perturbations introduced by the proposed sampling of basis elements can propagate to the appearance of generated images.

ICML Conference 2019 Conference Paper

Adversarially Learned Representations for Information Obfuscation and Inference

  • Martín Bertrán
  • Natalia Martínez
  • Afroditi Papadaki
  • Qiang Qiu 0001
  • Miguel Rodrigues 0001
  • Galen Reeves
  • Guillermo Sapiro

Data collection and sharing are pervasive aspects of modern society. This process can either be voluntary, as in the case of a person taking a facial image to unlock his/her phone, or incidental, such as traffic cameras collecting videos on pedestrians. An undesirable side effect of these processes is that shared data can carry information about attributes that users might consider as sensitive, even when such information is of limited use for the task. It is therefore desirable for both data collectors and users to design procedures that minimize sensitive information leakage. Balancing the competing objectives of providing meaningful individualized service levels and inference while obfuscating sensitive information is still an open problem. In this work, we take an information theoretic approach that is implemented as an unconstrained adversarial game between Deep Neural Networks in a principled, data-driven manner. This approach enables us to learn domain-preserving stochastic transformations that maintain performance on existing algorithms while minimizing sensitive information leakage.

ICML Conference 2018 Conference Paper

DCFNet: Deep Neural Network with Decomposed Convolutional Filters

  • Qiang Qiu 0001
  • Xiuyuan Cheng
  • A. Robert Calderbank
  • Guillermo Sapiro

Filters in a Convolutional Neural Network (CNN) contain model parameters learned from enormous amounts of data. In this paper, we suggest to decompose convolutional filters in CNN as a truncated expansion with pre-fixed bases, namely the Decomposed Convolutional Filters network (DCFNet), where the expansion coefficients remain learned from data. Such a structure not only reduces the number of trainable parameters and computation, but also imposes filter regularity by bases truncation. Through extensive experiments, we consistently observe that DCFNet maintains accuracy for image classification tasks with a significant reduction of model parameters, particularly with Fourier-Bessel (FB) bases, and even with random bases. Theoretically, we analyze the representation stability of DCFNet with respect to input variations, and prove representation stability under generic assumptions on the expansion coefficients. The analysis is consistent with the empirical observations.

YNIMG Journal 2018 Journal Article

Estimation of white matter fiber parameters from compressed multiresolution diffusion MRI using sparse Bayesian learning

  • Pramod Kumar Pisharady
  • Stamatios N. Sotiropoulos
  • Julio M. Duarte-Carvajalino
  • Guillermo Sapiro
  • Christophe Lenglet

We present a sparse Bayesian unmixing algorithm BusineX: Bayesian Unmixing for Sparse Inference-based Estimation of Fiber Crossings (X), for estimation of white matter fiber parameters from compressed (under-sampled) diffusion MRI (dMRI) data. BusineX combines compressive sensing with linear unmixing and introduces sparsity to the previously proposed multiresolution data fusion algorithm RubiX, resulting in a method for improved reconstruction, especially from data with lower number of diffusion gradients. We formulate the estimation of fiber parameters as a sparse signal recovery problem and propose a linear unmixing framework with sparse Bayesian learning for the recovery of sparse signals, the fiber orientations and volume fractions. The data is modeled using a parametric spherical deconvolution approach and represented using a dictionary created with the exponential decay components along different possible diffusion directions. Volume fractions of fibers along these directions define the dictionary weights. The proposed sparse inference, which is based on the dictionary representation, considers the sparsity of fiber populations and exploits the spatial redundancy in data representation, thereby facilitating inference from under-sampled q-space. The algorithm improves parameter estimation from dMRI through data-dependent local learning of hyperparameters, at each voxel and for each possible fiber orientation, that moderate the strength of priors governing the parameter variances. Experimental results on synthetic and in-vivo data show improved accuracy with a lower uncertainty in fiber parameter estimates. BusineX resolves a higher number of second and third fiber crossings. For under-sampled data, the algorithm is also shown to produce more reliable estimates.

NeurIPS Conference 2015 Conference Paper

Discriminative Robust Transformation Learning

  • Jiaji Huang
  • Qiang Qiu
  • Guillermo Sapiro
  • Robert Calderbank

This paper proposes a framework for learning features that are robust to data variation, which is particularly important when only a limited number of trainingsamples are available. The framework makes it possible to tradeoff the discriminative value of learned features against the generalization error of the learning algorithm. Robustness is achieved by encouraging the transform that maps data to features to be a local isometry. This geometric property is shown to improve (K, \epsilon)-robustness, thereby providing theoretical justification for reductions in generalization error observed in experiments. The proposed optimization frameworkis used to train standard learning algorithms such as deep neural networks. Experimental results obtained on benchmark datasets, such as labeled faces in the wild, demonstrate the value of being able to balance discrimination and robustness.

JMLR Journal 2015 Journal Article

Learning Transformations for Clustering and Classification

  • Qiang Qiu
  • Guillermo Sapiro

A low-rank transformation learning framework for subspace clustering and classification is proposed here. Many high- dimensional data, such as face images and motion sequences, approximately lie in a union of low-dimensional subspaces. The corresponding subspace clustering problem has been extensively studied in the literature to partition such high-dimensional data into clusters corresponding to their underlying low- dimensional subspaces. Low-dimensional intrinsic structures are often violated for real-world observations, as they can be corrupted by errors or deviate from ideal models. We propose to address this by learning a linear transformation on subspaces using nuclear norm as the modeling and optimization criteria. The learned linear transformation restores a low-rank structure for data from the same subspace, and, at the same time, forces a maximally separated structure for data from different subspaces. In this way, we reduce variations within the subspaces, and increase separation between the subspaces for a more robust subspace clustering. This proposed learned robust subspace clustering framework significantly enhances the performance of existing subspace clustering methods. Basic theoretical results presented here help to further support the underlying framework. To exploit the low-rank structures of the transformed subspaces, we further introduce a fast subspace clustering technique, which efficiently combines robust PCA with sparse modeling. When class labels are present at the training stage, we show this low-rank transformation framework also significantly enhances classification performance. Extensive experiments using public data sets are presented, showing that the proposed approach significantly outperforms state-of-the-art methods for subspace clustering and classification. The learned low cost transform is also applicable to other classification frameworks. [abs] [ pdf ][ bib ] &copy JMLR 2015. ( edit, beta )

YNIMG Journal 2014 Journal Article

Automatic clustering and population analysis of white matter tracts using maximum density paths

  • Gautam Prasad
  • Shantanu H. Joshi
  • Neda Jahanshad
  • Julio Villalon-Reina
  • Iman Aganj
  • Christophe Lenglet
  • Guillermo Sapiro
  • Katie L. McMahon

We introduce a framework for population analysis of white matter tracts based on diffusion-weighted images of the brain. The framework enables extraction of fibers from high angular resolution diffusion images (HARDI); clustering of the fibers based partly on prior knowledge from an atlas; representation of the fiber bundles compactly using a path following points of highest density (maximum density path; MDP); and registration of these paths together using geodesic curve matching to find local correspondences across a population. We demonstrate our method on 4-Tesla HARDI scans from 565 young adults to compute localized statistics across 50 white matter tracts based on fractional anisotropy (FA). Experimental results show increased sensitivity in the determination of genetic influences on principal fiber tracts compared to the tract-based spatial statistics (TBSS) method. Our results show that the MDP representation reveals important parts of the white matter structure and considerably reduces the dimensionality over comparable fiber matching approaches.

ICLR Conference 2014 Conference Paper

Learning Transformations for Classification Forests

  • Qiang Qiu 0001
  • Guillermo Sapiro

This work introduces a transformation-based learner model for classification forests. The weak learner at each split node plays a crucial role in a classification tree. We propose to optimize the splitting objective by learning a linear transformation on subspaces using nuclear norm as the optimization criteria. The learned linear transformation restores a low-rank structure for data from the same class, and, at the same time, maximizes the separation between different classes, thereby improving the performance of the split function. Theoretical and experimental results support the proposed framework.

JBHI Journal 2014 Journal Article

Semiautomatic Segmentation of Brain Subcortical Structures From High-Field MRI

  • JinYoung Kim
  • Christophe Lenglet
  • Yuval Duchin
  • Guillermo Sapiro
  • Noam Harel

Volumetric segmentation of subcortical structures, such as the basal ganglia and thalamus, is necessary for noninvasive diagnosis and neurosurgery planning. This is a challenging problem due in part to limited boundary information between structures, similar intensity profiles across the different structures, and low contrast data. This paper presents a semiautomatic segmentation system exploiting the superior image quality of ultrahigh field (7 T) MRI. The proposed approach utilizes the complementary edge information in the multiple structural MRI modalities. It combines optimally selected two modalities from susceptibility-weighted, T 2 -weighted, and diffusion MRI, and introduces a tailored new edge indicator function. In addition to this, we employ prior shape and configuration knowledge of the subcortical structures in order to guide the evolution of geometric active surfaces. Neighboring structures are segmented iteratively, constraining oversegmentation at their borders with a nonoverlapping penalty. Several experiments with data acquired on a 7 T MRI scanner demonstrate the feasibility and power of the approach for the segmentation of basal ganglia components critical for neurosurgery applications such as deep brain stimulation surgery.

ICLR Conference 2014 Conference Paper

Sparse similarity-preserving hashing

  • Jonathan Masci
  • Alex M. Bronstein
  • Michael M. Bronstein
  • Pablo Sprechmann
  • Guillermo Sapiro

In recent years, a lot of attention has been devoted to efficient nearest neighbor search by means of similarity-preserving hashing. One of the plights of existing hashing techniques is the intrinsic trade-off between performance and computational complexity: while longer hash codes allow for lower false positive rates, it is very difficult to increase the embedding dimensionality without incurring in very high false negatives rates or prohibiting computational costs. In this paper, we propose a way to overcome this limitation by enforcing the hash codes to be sparse. Sparse high-dimensional codes enjoy from the low false positive rates typical of long hashes, while keeping the false negative rates similar to those of a shorter dense hashing scheme with equal number of degrees of freedom. We use a tailored feed-forward neural network for the hashing function. Extensive experimental evaluation involving visual and multi-modal data shows the benefits of the proposed method.

YNIMG Journal 2013 Journal Article

Advances in diffusion MRI acquisition and processing in the Human Connectome Project

  • Stamatios N. Sotiropoulos
  • Saad Jbabdi
  • Junqian Xu
  • Jesper L. Andersson
  • Steen Moeller
  • Edward J. Auerbach
  • Matthew F. Glasser
  • Moises Hernandez

The Human Connectome Project (HCP) is a collaborative 5-year effort to map human brain connections and their variability in healthy adults. A consortium of HCP investigators will study a population of 1200 healthy adults using multiple imaging modalities, along with extensive behavioral and genetic data. In this overview, we focus on diffusion MRI (dMRI) and the structural connectivity aspect of the project. We present recent advances in acquisition and processing that allow us to obtain very high-quality in-vivo MRI data, whilst enabling scanning of a very large number of subjects. These advances result from 2years of intensive efforts in optimising many aspects of data acquisition and processing during the piloting phase of the project. The data quality and methods described here are representative of the datasets and processing pipelines that will be made freely available to the community at quarterly intervals, beginning in 2013.

IROS Conference 2013 Conference Paper

Locating occupants in preschool classrooms using a multiple RGB-D sensor system

  • Nicholas Walczak
  • Joshua Fasching
  • William D. Toczyski
  • Vassilios Morellas
  • Guillermo Sapiro
  • Nikolaos P. Papanikolopoulos

Presented are results demonstrating that, in developing a system with its first objective being the sustained detection of adults and young children as they move and interact in a normal preschool setting, the direct application of the straightforward RGB-D innovations presented here significantly outperforms even far more algorithmically advanced methods relying solely on images. The use of multiple RGB-D sensors by this project for depth-aware object localization economically resolves numerous issues regularly frustrating earlier vision-only detection and human surveillance methods, issues such as occlusions, illumination changes, unexpected postures, atypical morphologies, erratic or unanticipated motions, reflections, and misleading textures and colorations. This multiple RGB-D installation forms the front-end for a multi-step pipeline, the first portion of which seeks to isolate, in situ, 3D renderings of classroom occupants sufficient for a later analysis of their behaviors and interactions. Towards this end, a voxel-based approach to foreground/background separation and an effective adaptation of supervoxel clustering for 3D were developed, and 3D and image-only methods were tested and compared. The project's setting is highly challenging, but then so are its longer term goals: the automated detection of early childhood precursors, ofttimes very subtle, to a number of increasingly common developmental disorders.

YNIMG Journal 2013 Journal Article

Pushing spatial and temporal resolution for functional and diffusion MRI in the Human Connectome Project

  • Kamil Uğurbil
  • Junqian Xu
  • Edward J. Auerbach
  • Steen Moeller
  • An T. Vu
  • Julio M. Duarte-Carvajalino
  • Christophe Lenglet
  • Xiaoping Wu

The Human Connectome Project (HCP) relies primarily on three complementary magnetic resonance (MR) methods. These are: 1) resting state functional MR imaging (rfMRI) which uses correlations in the temporal fluctuations in an fMRI time series to deduce ‘functional connectivity’; 2) diffusion imaging (dMRI), which provides the input for tractography algorithms used for the reconstruction of the complex axonal fiber architecture; and 3) task based fMRI (tfMRI), which is employed to identify functional parcellation in the human brain in order to assist analyses of data obtained with the first two methods. We describe technical improvements and optimization of these methods as well as instrumental choices that impact speed of acquisition of fMRI and dMRI images at 3T, leading to whole brain coverage with 2mm isotropic resolution in 0. 7s for fMRI, and 1. 25mm isotropic resolution dMRI data for tractography analysis with three-fold reduction in total dMRI data acquisition time. Ongoing technical developments and optimization for acquisition of similar data at 7T magnetic field are also presented, targeting higher spatial resolution, enhanced specificity of functional imaging signals, mitigation of the inhomogeneous radio frequency (RF) fields, and reduced power deposition. Results demonstrate that overall, these approaches represent a significant advance in MR imaging of the human brain to investigate brain function and structure.

NeurIPS Conference 2013 Conference Paper

Robust Multimodal Graph Matching: Sparse Coding Meets Graph Matching

  • Marcelo Fiori
  • Pablo Sprechmann
  • Joshua Vogelstein
  • Pablo Muse
  • Guillermo Sapiro

Graph matching is a challenging problem with very important applications in a wide range of fields, from image and video analysis to biological and biomedical problems. We propose a robust graph matching algorithm inspired in sparsity-related techniques. We cast the problem, resembling group or collaborative sparsity formulations, as a non-smooth convex optimization problem that can be efficiently solved using augmented Lagrangian techniques. The method can deal with weighted or unweighted graphs, as well as multimodal data, where different graphs represent different types of data. The proposed approach is also naturally integrated with collaborative graph inference techniques, solving general network inference problems where the observed variables, possibly coming from different modalities, are not in correspondence. The algorithm is tested and compared with state-of-the-art graph matching techniques in both synthetic and real graphs. We also present results on multimodal graphs and applications to collaborative inference of brain connectivity from alignment-free functional magnetic resonance imaging (fMRI) data.

NeurIPS Conference 2013 Conference Paper

Supervised Sparse Analysis and Synthesis Operators

  • Pablo Sprechmann
  • Roee Litman
  • Tal Ben Yakar
  • Alexander Bronstein
  • Guillermo Sapiro

In this paper, we propose a new and computationally efficient framework for learning sparse models. We formulate a unified approach that contains as particular cases models promoting sparse synthesis and analysis type of priors, and mixtures thereof. The supervised training of the proposed model is formulated as a bilevel optimization problem, in which the operators are optimized to achieve the best possible performance on a specific task, e. g. , reconstruction or classification. By restricting the operators to be shift invariant, our approach can be thought as a way of learning analysis+synthesis sparsity-promoting convolutional operators. Leveraging recent ideas on fast trainable regressors designed to approximate exact sparse codes, we propose a way of constructing feed-forward neural networks capable of approximating the learned models at a fraction of the computational cost of exact solvers. In the shift-invariant case, this leads to a principled way of constructing task-specific convolutional networks. We illustrate the proposed models on several experiments in music analysis and image processing applications.

ICRA Conference 2012 Conference Paper

A multi-sensor visual tracking system for behavior monitoring of at-risk children

  • Ravishankar Sivalingam
  • Anoop Cherian
  • Joshua Fasching
  • Nicholas Walczak
  • Nathaniel D. Bird
  • Vassilios Morellas
  • Barbara Murphy
  • Kathryn Cullen

Clinical studies confirm that mental illnesses such as autism, Obsessive Compulsive Disorder (OCD), etc. show behavioral abnormalities even at very young ages; the early diagnosis of which can help steer effective treatments. Most often, the behavior of such at-risk children deviate in very subtle ways from that of a normal child; correct diagnosis of which requires prolonged and continuous monitoring of their activities by a clinician, which is a difficult and time intensive task. As a result, the development of automation tools for assisting in such monitoring activities will be an important step towards effective utilization of the diagnostic resources. In this paper, we approach the problem from a computer vision standpoint, and propose a novel system for the automatic monitoring of the behavior of children in their natural environment through the deployment of multiple non-invasive sensors (cameras and depth sensors). We provide details of our system, together with algorithms for the robust tracking of the activities of the children. Our experiments, conducted in the Shirley G. Moore Laboratory School, demonstrate the effectiveness of our methodology.

IROS Conference 2012 Conference Paper

Detecting risk-markers in children in a preschool classroom

  • Joshua Fasching
  • Nicholas Walczak
  • Ravishankar Sivalingam
  • Kathryn Cullen
  • Barbara Murphy
  • Guillermo Sapiro
  • Vassilios Morellas
  • Nikolaos P. Papanikolopoulos

Early intervention in mental disorders can dramatically increase an individual's quality of life. Additionally, when symptoms of mental illness appear in childhood or adolescence, they represent the later stages of a process that began years earlier. One goal of psychiatric research is to identify risk-markers: genetic, neural, behavioral and/or social deviations that indicate elevated risk of a particular mental disorder. Ideally, screening of risk-markers should occur in a community setting, and not a clinical setting which may be time-consuming and resource-intensive. Given this situation, a system for automatically detecting risk-markers in children would be highly valuable. In this paper, we describe such a system that has been installed at the Shirley G. Moore Lab School, a research pre-school at the University of Minnesota. This system consists of multiple RGB+D sensors and is able to detect children and adults in the classroom, tracking them as they move around the room. We use the tracking results to extract high-level information about the behavior and social interaction of children, that can then be used to screen for early signs of mental disorders.

NeurIPS Conference 2012 Conference Paper

Finding Exemplars from Pairwise Dissimilarities via Simultaneous Sparse Recovery

  • Ehsan Elhamifar
  • Guillermo Sapiro
  • René Vidal

Given pairwise dissimilarities between data points, we consider the problem of finding a subset of data points called representatives or exemplars that can efficiently describe the data collection. We formulate the problem as a row-sparsity regularized trace minimization problem which can be solved efficiently using convex programming. The solution of the proposed optimization program finds the representatives and the probability that each data point is associated to each one of the representatives. We obtain the range of the regularization parameter for which the solution of the proposed optimization program changes from selecting one representative to selecting all data points as the representatives. When data points are distributed around multiple clusters according to the dissimilarities, we show that the data in each cluster select only representatives from that cluster. Unlike metric-based methods, our algorithm does not require that the pairwise dissimilarities be metrics and can be applied to dissimilarities that are asymmetric or violate the triangle inequality. We demonstrate the effectiveness of the proposed algorithm on synthetic data as well as real-world datasets of images and text.

YNIMG Journal 2012 Journal Article

Hierarchical topological network analysis of anatomical human brain connectivity and differences related to sex and kinship

  • Julio M. Duarte-Carvajalino
  • Neda Jahanshad
  • Christophe Lenglet
  • Katie L. McMahon
  • Greig I. de Zubicaray
  • Nicholas G. Martin
  • Margaret J. Wright
  • Paul M. Thompson

Modern non-invasive brain imaging technologies, such as diffusion weighted magnetic resonance imaging (DWI), enable the mapping of neural fiber tracts in the white matter, providing a basis to reconstruct a detailed map of brain structural connectivity networks. Brain connectivity networks differ from random networks in their topology, which can be measured using small worldness, modularity, and high-degree nodes (hubs). Still, little is known about how individual differences in structural brain network properties relate to age, sex, or genetic differences. Recently, some groups have reported brain network biomarkers that enable differentiation among individuals, pairs of individuals, and groups of individuals. In addition to studying new topological features, here we provide a unifying general method to investigate topological brain networks and connectivity differences between individuals, pairs of individuals, and groups of individuals at several levels of the data hierarchy, while appropriately controlling false discovery rate (FDR) errors. We apply our new method to a large dataset of high quality brain connectivity networks obtained from High Angular Resolution Diffusion Imaging (HARDI) tractography in 303 young adult twins, siblings, and unrelated people. Our proposed approach can accurately classify brain connectivity networks based on sex (93% accuracy) and kinship (88. 5% accuracy). We find statistically significant differences associated with sex and kinship both in the brain connectivity networks and in derived topological metrics, such as the clustering coefficient and the communicability matrix.

NeurIPS Conference 2012 Conference Paper

Topology Constraints in Graphical Models

  • Marcelo Fiori
  • Pablo Musé
  • Guillermo Sapiro

Graphical models are a very useful tool to describe and understand natural phenomena, from gene expression to climate change and social interactions. The topological structure of these graphs/networks is a fundamental part of the analysis, and in many cases the main goal of the study. However, little work has been done on incorporating prior topological knowledge onto the estimation of the underlying graphical models from sample data. In this work we propose extensions to the basic joint regression model for network estimation, which explicitly incorporate graph-topological constraints into the corresponding optimization approach. The first proposed extension includes an eigenvector centrality constraint, thereby promoting this important prior topological property. The second developed extension promotes the formation of certain motifs, triangle-shaped ones in particular, which are known to exist for example in genetic regulatory networks. The presentation of the underlying formulations, which serve as examples of the introduction of topological constraints in network estimation, is complemented with examples in diverse datasets demonstrating the importance of incorporating such critical prior knowledge.

JMLR Journal 2010 Journal Article

Online Learning for Matrix Factorization and Sparse Coding

  • Julien Mairal
  • Francis Bach
  • Jean Ponce
  • Guillermo Sapiro

Sparse coding-that is, modelling data vectors as sparse linear combinations of basis elements-is widely used in machine learning, neuroscience, signal processing, and statistics. This paper focuses on the large-scale matrix factorization problem that consists of learning the basis set in order to adapt it to specific data. Variations of this problem include dictionary learning in signal processing, non-negative matrix factorization and sparse principal component analysis. In this paper, we propose to address these tasks with a new online optimization algorithm, based on stochastic approximations, which scales up gracefully to large data sets with millions of training samples, and extends naturally to various matrix factorization formulations, making it suitable for a wide range of learning problems. A proof of convergence is presented, along with experiments with natural images and genomic data demonstrating that it leads to state-of-the-art performance in terms of speed and optimization for both small and large data sets. [abs] [ pdf ][ bib ] &copy JMLR 2010. ( edit, beta )

NeurIPS Conference 2009 Conference Paper

Non-Parametric Bayesian Dictionary Learning for Sparse Image Representations

  • Mingyuan Zhou
  • Haojun Chen
  • Lu Ren
  • Guillermo Sapiro
  • Lawrence Carin
  • John Paisley

Non-parametric Bayesian techniques are considered for learning dictionaries for sparse image representations, with applications in denoising, inpainting and compressive sensing (CS). The beta process is employed as a prior for learning the dictionary, and this non-parametric method naturally infers an appropriate dictionary size. The Dirichlet process and a probit stick-breaking process are also considered to exploit structure within an image. The proposed method can learn a sparse dictionary in situ; training images may be exploited if available, but they are not required. Further, the noise variance need not be known, and can be non-stationary. Another virtue of the proposed method is that sequential inference can be readily employed, thereby allowing scaling to large images. Several example results are presented, using both Gibbs and variational Bayesian inference, with comparisons to other state-of-the-art approaches.

ICML Conference 2009 Conference Paper

Online dictionary learning for sparse coding

  • Julien Mairal
  • Francis R. Bach
  • Jean Ponce
  • Guillermo Sapiro

Sparse coding---that is, modelling data vectors as sparse linear combinations of basis elements---is widely used in machine learning, neuroscience, signal processing, and statistics. This paper focuses on learning the basis set, also called dictionary, to adapt it to specific data, an approach that has recently proven to be very effective for signal reconstruction and classification in the audio and image processing domains. This paper proposes a new online optimization algorithm for dictionary learning, based on stochastic approximations, which scales up gracefully to large datasets with millions of training samples. A proof of convergence is presented, along with experiments with natural images demonstrating that it leads to faster performance and better dictionaries than classical batch algorithms for both small and large datasets.

NeurIPS Conference 2008 Conference Paper

Supervised Dictionary Learning

  • Julien Mairal
  • Jean Ponce
  • Guillermo Sapiro
  • Andrew Zisserman
  • Francis Bach

It is now well established that sparse signal models are well suited to restoration tasks and can effectively be learned from audio, image, and video data. Recent research has been aimed at learning discriminative sparse models instead of purely reconstructive ones. This paper proposes a new step in that direction with a novel sparse representation for signals belonging to different classes in terms of a shared dictionary and multiple decision functions. It is shown that the linear variant of the model admits a simple probabilistic interpretation, and that its most general variant also admits a simple interpretation in terms of kernels. An optimization framework for learning all the components of the proposed model is presented, along with experiments on standard handwritten digit and texture classification tasks.

NeurIPS Conference 2006 Conference Paper

Stratification Learning: Detecting Mixed Density and Dimensionality in High Dimensional Point Clouds

  • Gloria Haro
  • Gregory Randall
  • Guillermo Sapiro

The study of point cloud data sampled from a stratification, a collection of manifolds with possible different dimensions, is pursued in this paper. We present a technique for simultaneously soft clustering and estimating the mixed dimensionality and density of such structures. The framework is based on a maximum likelihood estimation of a Poisson mixture model. The presentation of the approach is completed with artificial and real examples demonstrating the importance of extending manifold learning to stratification learning.

YNIMG Journal 2004 Journal Article

Implicit brain imaging

  • Facundo Mémoli
  • Guillermo Sapiro
  • Paul Thompson

We describe how implicit surface representations can be used to solve fundamental problems in brain imaging. This kind of representation is not only natural following the state-of-the-art segmentation algorithms reported in the literature to extract the different brain tissues, but it is also, as shown in this paper, the most appropriate one from the computational point of view. Examples are provided for finding constrained special curves on the cortex, such as sulcal beds, regularizing surface-based measures, such as cortical thickness, and for computing warping fields between surfaces such as the brain cortex. All these result from efficiently solving partial differential equations (PDEs) and variational problems on surfaces represented in implicit form. The implicit framework avoids the need to construct intermediate mappings between 3-D anatomical surfaces and parametric objects such planes or spheres, a complex step that introduces errors and is required by many other cortical processing approaches.

v2026.09.13