Arrow Research search
Back to UAI

UAI 2015

Estimating the Partition Function by Discriminance Sampling

Conference Paper Accepted Paper Artificial Intelligence · Machine Learning · Uncertainty in Artificial Intelligence

Abstract

models. As such, efficient approximations via variational inference and Monte Carlo methods are of great interest. Importance sampling (IS) and its variant, annealed IS (AIS) have been widely used for estimating the partition function in graphical models, such as Markov random fields and deep generative models. However, IS tends to underestimate the partition function and is subject to high variance when the proposal distribution is more peaked than the target distribution. On the other hand, “reverse” versions of IS and AIS tend to overestimate the partition function, and degenerate when the target distribution is more peaked than the proposal distribution. In this work, we present a simple, general method that gives much more reliable and robust estimates than either IS (AIS) or reverse IS (AIS). Our method works by converting the estimation problem into a simple classification problem that discriminates between the samples drawn from the target and the proposal. We give extensive theoretical and empirical justification; in particular, we show that an annealed version of our method significantly outperforms both AIS and reverse AIS as proposed by Burda et al. (2015), which has been the stateof-the-art for likelihood evaluation in deep generative models. Importance sampling (IS) and its variants, such as annealed importance sampling (AIS) (Neal, 2001), are probably the most widely used Monte Carlo methods for estimating the partition function. IS works by drawing samples from a tractable proposal (or reference) distribution R p0 (x), and estimates the target partition function Z = f (x) by averaging the importance weights f (x)/p0 (x) across the samples. Unfortunately, the IS estimate often has very high variance if the choice of proposal distribution is very different from the target, especially when the proposal is more peaked or has thinner tails than the target. In addition, in practice IS often underestimates the partition function due to the heavy-tailed nature of the importance weights, leading to overly optimistic likelihood estimates when used for model evaluation (e. g. , Burda et al. , 2015).

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
Conference on Uncertainty in Artificial Intelligence
Archive span
1985-2025
Indexed papers
3717
Paper id
954081108659067787
v2026.09.13