Arrow Research search
Back to AAMAS

AAMAS 2026

Stochastically Dominant Preference Optimization: Policy Improvement For All

Conference Paper Extended Abstracts Autonomous Agents and Multiagent Systems

Abstract

Reinforcement learning from human feedback (RLHF) optimizes policies based on users’ rankings of output samples rather than using user-provided rewards. These methods typically assume users’ underlying utility functions are homogeneous and their rankings differ only due to noise, ultimately optimizing for the average user. Instead, we seek policies that guarantee improvement for all users with respect to their heterogeneous preferences. We introduce stochastic dominance as a stricter guiding criteria for policy optimization that guarantees improvement under any social welfare function. Our approach, stochastically dominant preference optimization (SDPO), avoids explicit reward function estimation while providing individual performance improvement guarantees for users with diverse preferences.

Authors

Keywords

  • RLHF
  • policy improvement
  • stochastic dominance

Context

Venue
International Conference on Autonomous Agents and Multiagent Systems
Archive span
2002-2026
Indexed papers
8043
Paper id
621496773854555491
v2026.09.13