Arrow Research search
Back to AAAI

AAAI 2023

Mitigating Adversarial Norm Training with Moral Axioms

Conference Paper AAAI Technical Track on Philosophy and Ethics of AI Artificial Intelligence

Abstract

This paper addresses the issue of adversarial attacks on ethical AI systems. We investigate using moral axioms and rules of deontic logic in a norm learning framework to mitigate adversarial norm training. This model of moral intuition and construction provides AI systems with moral guard rails yet still allows for learning conventions. We evaluate our approach by drawing inspiration from a study commonly used in moral development research. This questionnaire aims to test an agent's ability to reason to moral conclusions despite opposed testimony. Our findings suggest that our model can still correctly evaluate moral situations and learn conventions in an adversarial training environment. We conclude that adding axiomatic moral prohibitions and deontic inference rules to a norm learning model makes it less vulnerable to adversarial attacks.

Authors

Keywords

  • CMS: Social Cognition And Interaction
  • KRR: Belief Change
  • KRR: Reasoning with Beliefs
  • ML: Adversarial Learning & Robustness
  • PEAI: AI and Epistemology
  • PEAI: Morality and Value-Based AI
  • PEAI: Safety, Robustness & Trustworthiness
  • RU: Uncertainty Representations

Context

Venue
AAAI Conference on Artificial Intelligence
Archive span
1980-2026
Indexed papers
28718
Paper id
342406174102801449
v2026.09.13