AAMAS 2026
Robust Direct Preference Optimization for Offline Learning
Abstract
Recent preference-based alignment methods, such as Direct Preference Optimization (DPO), perform well under comprehensive offline datasets. However, practical datasets often contain sparse, noisy, and unevenly distributed comparisons, which can degrade model performance. To address this, we first adopt the principle of pessimism and propose a Robust DPO framework that optimizes for the worst-case reward within some data-dependent uncertainty set, and show its effectiveness in offline problems. Moreover, we show that the resulting robust optimal policy can be obtained by directly fine-tuning a baseline DPO model, avoiding the need for retraining. We further construct an uncertainty set to tackle the data uncertainty, based on the graph Laplacian, and show the set contains the true underlying reward with high probability. We then further evaluate the effectiveness of our method in controlled tabular and LLM setting, which validate our theoretical finds.
Authors
Keywords
Context
- Venue
- International Conference on Autonomous Agents and Multiagent Systems
- Archive span
- 2002-2026
- Indexed papers
- 8043
- Paper id
- 120597508940576247