Arrow Research search
Back to IROS

IROS 2025

M3PO: Massively Multi-Task Model-Based Policy Optimization

Conference Paper Accepted Paper Artificial Intelligence · Robotics

Abstract

We introduce Massively Multi-Task Model-Based Policy Optimization (M3PO), a scalable model-based reinforcement learning (MBRL) framework designed to address the challenges of sample efficiency in single-task settings and generalization in multi-task domains. Existing model-based approaches like DreamerV3 rely on generative world models that prioritize pixel-level reconstruction, often at the cost of control-centric representations, while model-free methods such as PPO suffer from high sample complexity and limited exploration. M3PO integrates an implicit world model, trained to predict task outcomes without reconstructing observations, with a hybrid exploration strategy that combines model-based planning and model-free uncertainty-driven bonuses. This approach eliminates the bias-variance trade-off inherent in prior methods (e. g. , POME’s exploration bonuses) by using the discrepancy between model-based and model-free value estimates to guide exploration while maintaining stable policy updates via a trust-region optimizer. M3PO is introduced as an advanced alternative to existing model-based policy optimization methods.

Authors

Keywords

  • Training
  • Visualization
  • Computational modeling
  • Scalability
  • Reinforcement learning
  • Predictive models
  • Multitasking
  • Planning
  • Incentive schemes
  • Overfitting
  • Optimal Policy
  • Model-based Optimization
  • Model-Based Policy Optimization
  • Estimated Values
  • Sampling Efficiency
  • Implicit Model
  • Model-based Estimates
  • Policy Update
  • Model-free Methods
  • Policy Stability
  • Model-based Reinforcement Learning
  • Benchmark
  • Prediction Model
  • Learning Models
  • Value Function
  • State Space
  • Parallelization
  • Latent Model
  • Model Predictive Control
  • Markov Decision Process
  • Reinforcement Learning Algorithm
  • Diverse Tasks
  • Multi-task Learning
  • Average Return
  • State St
  • Continuous Action Space
  • Latent State
  • Policy Gradient
  • Trajectory Optimization
  • Model-based Algorithm

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
552981354632115524
v2026.09.13