Arrow Research search
Back to ICML

ICML 2025

Trustworthy Machine Learning through Data-Specific Indistinguishability

Conference Paper Accept (poster) Artificial Intelligence ยท Machine Learning

Abstract

This paper studies a range of AI/ML trust concepts, including memorization, data poisoning, and copyright, which can be modeled as constraints on the influence of data on a (trained) model, characterized by the outcome difference from a processing function (training algorithm). In this realm, we show that provable trust guarantees can be efficiently provided through a new framework termed Data-Specific Indistinguishability (DSI) to select trust-preserving randomization tightly aligning with targeted outcome differences, as a relaxation of the classic Input-Independent Indistinguishability (III). We establish both the theoretical and algorithmic foundations of DSI with the optimal multivariate Gaussian mechanism. We further show its applications to develop trustworthy deep learning with black-box optimizers. The experimental results on memorization mitigation, backdoor defense, and copyright protection show both the efficiency and effectiveness of the DSI noise mechanism.

Authors

Keywords

  • Differential trustworthiness
  • data-specific indistinguishability
  • memorization
  • backdoor attacks
  • copyright

Context

Venue
International Conference on Machine Learning
Archive span
1993-2025
Indexed papers
16471
Paper id
745583726813848921
v2026.09.13