AAMAS 2026
Stochastically Dominant Preference Optimization: Policy Improvement For All
Abstract
Reinforcement learning from human feedback (RLHF) optimizes policies based on users’ rankings of output samples rather than using user-provided rewards. These methods typically assume users’ underlying utility functions are homogeneous and their rankings differ only due to noise, ultimately optimizing for the average user. Instead, we seek policies that guarantee improvement for all users with respect to their heterogeneous preferences. We introduce stochastic dominance as a stricter guiding criteria for policy optimization that guarantees improvement under any social welfare function. Our approach, stochastically dominant preference optimization (SDPO), avoids explicit reward function estimation while providing individual performance improvement guarantees for users with diverse preferences.
Authors
Keywords
Context
- Venue
- International Conference on Autonomous Agents and Multiagent Systems
- Archive span
- 2002-2026
- Indexed papers
- 8043
- Paper id
- 621496773854555491