Arrow Research search

Author name cluster

Syed M. Abbas

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

1 paper
1 author row

Possible papers

1

AAMAS Conference 2026 Conference Paper

Stochastically Dominant Preference Optimization: Policy Improvement For All

  • Ali Farajzadeh
  • Syed M. Abbas
  • Aadirupa Saha
  • Brian D. Ziebart

Reinforcement learning from human feedback (RLHF) optimizes policies based on users’ rankings of output samples rather than using user-provided rewards. These methods typically assume users’ underlying utility functions are homogeneous and their rankings differ only due to noise, ultimately optimizing for the average user. Instead, we seek policies that guarantee improvement for all users with respect to their heterogeneous preferences. We introduce stochastic dominance as a stricter guiding criteria for policy optimization that guarantees improvement under any social welfare function. Our approach, stochastically dominant preference optimization (SDPO), avoids explicit reward function estimation while providing individual performance improvement guarantees for users with diverse preferences.

v2026.09.13