Arrow Research search
Back to ICML

ICML 2025

Efficiently Vectorized MCMC on Modern Accelerators

Conference Paper Accept (spotlight poster) Artificial Intelligence · Machine Learning

Abstract

With the advent of automatic vectorization tools (e. g. , JAX’s vmap), writing multi-chain MCMC algorithms is often now as simple as invoking those tools on single-chain code. Whilst convenient, for various MCMC algorithms this results in a synchronization problem—loosely speaking, at each iteration all chains running in parallel must wait until the last chain has finished drawing its sample. In this work, we show how to design single-chain MCMC algorithms in a way that avoids synchronization overheads when vectorizing with tools like vmap, by using the framework of finite state machines (FSMs). Using a simplified model, we derive an exact theoretical form of the obtainable speed-ups using our approach, and use it to make principled recommendations for optimal algorithm design. We implement several popular MCMC algorithms as FSMs, including Elliptical Slice Sampling, HMC-NUTS, and Delayed Rejection, demonstrating speed-ups of up to an order of magnitude in experiments.

Authors

Keywords

  • Markov Chain Monte Carlo
  • MCMC
  • Parallel Computing

Context

Venue
International Conference on Machine Learning
Archive span
1993-2025
Indexed papers
16471
Paper id
608190682045691989
v2026.09.13