Arrow Research search
Back to RLC

RLC 2024

Posterior Sampling for Continuing Environments

Conference Paper RLC accepted paper Artificial Intelligence · Machine Learning · Reinforcement Learning

Abstract

Existing posterior sampling algorithms for continuing reinforcement learning (RL) rely on maintaining state-action visitation counts, making them unsuitable for complex environments with high-dimensional state spaces. We develop the first extension of posterior sampling for RL (PSRL) that is suited for a continuing agent-environment interface and integrates naturally into scalable agent designs. Our approach, continuing PSRL (CPSRL), determines when to resample a new model of the environment from the posterior distribution based on a simple randomization scheme. We establish an $\tilde{O}(\tau S \sqrt{A T})$ bound on the Bayesian regret in the tabular setting, where $S$ is the number of environment states, $A$ is the number of actions, and $\tau$ denotes the {\it reward averaging time}, which is a bound on the duration required to accurately estimate the average reward of any policy. Our work is the first to formalize and rigorously analyze this random resampling approach. Our simulations demonstrate CPSRL's effectiveness in high-dimensional state spaces where traditional algorithms fail.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
Reinforcement Learning Conference
Archive span
2024-2025
Indexed papers
228
Paper id
1364725544809134
v2026.09.13