Arrow Research search
Back to RLC

RLC 2024

Boosting Soft Q-Learning by Bounding

Conference Paper RLC accepted paper Artificial Intelligence · Machine Learning · Reinforcement Learning

Abstract

An agent’s ability to leverage past experience is critical for efficiently solving new tasks. Prior work has focused on using value function estimates to obtain zero-shot approximations for solutions to a new task. In soft $Q$-learning, we show how any value function estimate can also be used to derive double-sided bounds on the optimal value function. The derived bounds lead to new approaches for boosting training performance which we validate experimentally. Notably, we find that the proposed framework suggests an alternative method for updating the $Q$-function, leading to boosted performance.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
Reinforcement Learning Conference
Archive span
2024-2025
Indexed papers
228
Paper id
1150108906271725783
v2026.09.13