Arrow Research search

Author name cluster

Lyle H. Ungar

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
2 author rows

Possible papers

12

JMLR Journal 2015 Journal Article

Eigenwords: Spectral Word Embeddings

  • Paramveer S. Dhillon
  • Dean P. Foster
  • Lyle H. Ungar

Spectral learning algorithms have recently become popular in data-rich domains, driven in part by recent advances in large scale randomized SVD, and in spectral estimation of Hidden Markov Models. Extensions of these methods lead to statistical estimation algorithms which are not only fast, scalable, and useful on real data sets, but are also provably correct. Following this line of research, we propose four fast and scalable spectral algorithms for learning word embeddings -- low dimensional real vectors (called Eigenwords ) that capture the "meaning" of words from their context. All the proposed algorithms harness the multi-view nature of text data i.e. the left and right context of each word, are fast to train and have strong theoretical properties. Some of the variants also have lower sample complexity and hence higher statistical power for rare words. We provide theory which establishes relationships between these algorithms and optimality criteria for the estimates they provide. We also perform thorough qualitative and quantitative evaluation of Eigenwords showing that simple linear approaches give performance comparable to or superior than the state-of-the-art non-linear deep learning based methods. [abs] [ pdf ][ bib ] &copy JMLR 2015. ( edit, beta )

YNIMG Journal 2014 Journal Article

Subject-specific functional parcellation via Prior Based Eigenanatomy

  • Paramveer S. Dhillon
  • David A. Wolk
  • Sandhitsu R. Das
  • Lyle H. Ungar
  • James C. Gee
  • Brian B. Avants

We present a new framework for prior-constrained sparse decomposition of matrices derived from the neuroimaging data and apply this method to functional network analysis of a clinically relevant population. Matrix decomposition methods are powerful dimensionality reduction tools that have found widespread use in neuroimaging. However, the unconstrained nature of these totally data-driven techniques makes it difficult to interpret the results in a domain where network-specific hypotheses may exist. We propose a novel approach, Prior Based Eigenanatomy (p-Eigen), which seeks to identify a data-driven matrix decomposition but at the same time constrains the individual components by spatial anatomical priors (probabilistic ROIs). We formulate our novel solution in terms of prior-constrained ℓ 1 penalized (sparse) principal component analysis. p-Eigen starts with a common functional parcellation for all the subjects and refines it with subject-specific information. This enables modeling of the inter-subject variability in the functional parcel boundaries and allows us to construct subject-specific networks with reduced sensitivity to ROI placement. We show that while still maintaining correspondence across subjects, p-Eigen extracts biologically-relevant and patient-specific functional parcels that facilitate hypothesis-driven network analysis. We construct default mode network (DMN) connectivity graphs using p-Eigen refined ROIs and use them in a classification paradigm. Our results show that the functional connectivity graphs derived from p-Eigen significantly aid classification of mild cognitive impairment (MCI) as well as the prediction of scores in a Delayed Recall memory task when compared to graph metrics derived from 1) standard registration-based seed ROI definitions, 2) totally data-driven ROIs, 3) a model based on standard demographics plus hippocampal volume as covariates, and 4) Ward Clustering based data-driven ROIs. In summary, p-Eigen incarnates a new class of prior-constrained dimensionality reduction tools that may improve our understanding of the relationship between MCI and functional connectivity.

JMLR Journal 2013 Journal Article

A Risk Comparison of Ordinary Least Squares vs Ridge Regression

  • Paramveer S. Dhillon
  • Dean P. Foster
  • Sham M. Kakade
  • Lyle H. Ungar

We compare the risk of ridge regression to a simple variant of ordinary least squares, in which one simply projects the data onto a finite dimensional subspace (as specified by a principal component analysis) and then performs an ordinary (un- regularized) least squares regression in this subspace. This note shows that the risk of this ordinary least squares method (PCA-OLS) is within a constant factor (namely 4) of the risk of ridge regression (RR). [abs] [ pdf ][ bib ] &copy JMLR 2013. ( edit, beta )

JMLR Journal 2011 Journal Article

Minimum Description Length Penalization for Group and Multi-Task Sparse Learning

  • Paramveer S. Dhillon
  • Dean Foster
  • Lyle H. Ungar

We propose a framework MIC (Multiple Inclusion Criterion) for learning sparse models based on the information theoretic Minimum Description Length (MDL) principle. MIC provides an elegant way of incorporating arbitrary sparsity patterns in the feature space by using two-part MDL coding schemes. We present MIC based models for the problems of grouped feature selection (MIC-GROUP) and multi-task feature selection (MIC-MULTI). MIC-GROUP assumes that the features are divided into groups and induces two level sparsity, selecting a subset of the feature groups, and also selecting features within each selected group. MIC-MULTI applies when there are multiple related tasks that share the same set of potentially predictive features. It also induces two level sparsity, selecting a subset of the features, and then selecting which of the tasks each feature should be added to. Lastly, we propose a model, TRANSFEAT, that can be used to transfer knowledge from a set of previously learned tasks to a new task that is expected to share similar features. All three methods are designed for selecting a small set of predictive features from a large pool of candidate features. We demonstrate the effectiveness of our approach with experimental results on data from genomics and from word sense disambiguation problems. [abs] [ pdf ][ bib ] &copy JMLR 2011. ( edit, beta )

JMLR Journal 2006 Journal Article

Streamwise Feature Selection

  • Jing Zhou
  • Dean P. Foster
  • Robert A. Stine
  • Lyle H. Ungar

In streamwise feature selection, new features are sequentially considered for addition to a predictive model. When the space of potential features is large, streamwise feature selection offers many advantages over traditional feature selection methods, which assume that all features are known in advance. Features can be generated dynamically, focusing the search for new features on promising subspaces, and overfitting can be controlled by dynamically adjusting the threshold for adding features to the model. In contrast to traditional forward feature selection algorithms such as stepwise regression in which at each step all possible features are evaluated and the best one is selected, streamwise feature selection only evaluates each feature once when it is generated. We describe information-investing and α-investing, two adaptive complexity penalty methods for streamwise feature selection which dynamically adjust the threshold on the error reduction required for adding a new feature. These two methods give false discovery rate style guarantees against overfitting. They differ from standard penalty methods such as AIC, BIC and RIC, which always drastically over- or under-fit in the limit of infinite numbers of non-predictive features. Empirical results show that streamwise regression is competitive with (on small data sets) and superior to (on large data sets) much more compute-intensive feature selection methods such as stepwise regression, and allows feature selection on problems with millions of potential features. [abs] [ pdf ][ bib ] &copy JMLR 2006. ( edit, beta )

IROS Conference 2003 Conference Paper

Using policy gradient reinforcement learning on autonomous robot controllers

  • Gregory Z. Grudic
  • Vijay Kumar 0001
  • Lyle H. Ungar

Robot programmers can often quickly program a robot to approximately execute a task under specific environment conditions. However, achieving robust performance under more general conditions is significantly more difficult. We propose a framework that starts with an existing control system and uses reinforcement feedback from the environment to autonomously improve the controller's performance. We use the policy gradient reinforcement learning (PGRL) framework, which estimates a gradient (in controller space) of improved reward, allowing the controller parameters to be incrementally updated to autonomously achieve locally optimal performance. Our approach is experimentally verified on a Cye robot executing a room entry and observation task, showing significant reduction in task execution time and robustness with respect to un-modelled changes in the environment.

UAI Conference 2001 Conference Paper

Probabilistic Models for Unified Collaborative and Content-Based Recommendation in Sparse-Data Environments

  • Alexandrin Popescul
  • Lyle H. Ungar
  • David M. Pennock
  • Steve Lawrence

Recommender systems leverage product and community information to target products to consumers. Researchers have developed collaborative recommenders, content-based recommenders, and (largely ad-hoc) hybrid systems. We propose a unified probabilistic framework for merging collaborative and content-based recommendations. We extend Hofmann's [1999] aspect model to incorporate three-way co-occurrence data among users, items, and item content. The relative influence of collaboration data versus content data is not imposed as an exogenous parameter, but rather emerges naturally from the given data sources. Global probabilistic models coupled with standard Expectation Maximization (EM) learning algorithms tend to drastically overfit in sparse-data situations, as is typical in recommendation applications. We show that secondary content information can often be used to overcome sparsity. Experiments on data from the ResearchIndex library of Computer Science publications show that appropriate mixture models incorporating secondary data produce significantly better quality recommenders than k-nearest neighbors (k-NN). Global probabilistic models also allow more general inferences than local methods like k-NN.

v2026.09.13