Arrow Research search
Back to AAMAS

AAMAS 2026

Interactionless Inverse Reinforcement Learning: A Data-Centric Framework for Durable Alignment

Conference Paper Blue Sky Ideas Track Autonomous Agents and Multiagent Systems

Abstract

AIalignmentisgrowinginimportance, yetmanycurrentapproaches learn safety behavior by directly modifying policy parameters, entangling normative constraints with the underlying policy. This oftenyieldsopaque, difficult-to-editalignmentartifactsandreduces their reuse across models or deployments, a failure mode we term Alignment Waste. We propose Interactionless Inverse Reinforcement Learning, a framework for learning inspectable, editable, and reusable reward artifacts separately from policy optimization. We further introduce the Alignment Flywheel, a human-in-the-loop lifecycle for iteratively auditing, patching, and hardening these artifacts through automated evaluation and refinement. Together, these ideas recast alignment from a disposable training expense into a durable, verifiable engineering asset.

Authors

Keywords

  • AI alignment
  • AI safety
  • Inverse Reinforcement Learning
  • reward modeling
  • Alignment Waste
  • Alignment Flywheel

Context

Venue
International Conference on Autonomous Agents and Multiagent Systems
Archive span
2002-2026
Indexed papers
8043
Paper id
703494654422527794
v2026.09.13