AAMAS Conference 2026 Conference Paper
Interactionless Inverse Reinforcement Learning: A Data-Centric Framework for Durable Alignment
- Elias Malomgré
- Pieter Simoens
AIalignmentisgrowinginimportance, yetmanycurrentapproaches learn safety behavior by directly modifying policy parameters, entangling normative constraints with the underlying policy. This oftenyieldsopaque, difficult-to-editalignmentartifactsandreduces their reuse across models or deployments, a failure mode we term Alignment Waste. We propose Interactionless Inverse Reinforcement Learning, a framework for learning inspectable, editable, and reusable reward artifacts separately from policy optimization. We further introduce the Alignment Flywheel, a human-in-the-loop lifecycle for iteratively auditing, patching, and hardening these artifacts through automated evaluation and refinement. Together, these ideas recast alignment from a disposable training expense into a durable, verifiable engineering asset.