Arrow Research search
Back to AAAI

AAAI 2023

Counterfactual Learning with General Data-Generating Policies

Conference Paper AAAI Technical Track on Machine Learning III Artificial Intelligence

Abstract

Off-policy evaluation (OPE) attempts to predict the performance of counterfactual policies using log data from a different policy. We extend its applicability by developing an OPE method for a class of both full support and deficient support logging policies in contextual-bandit settings. This class includes deterministic bandit (such as Upper Confidence Bound) as well as deterministic decision-making based on supervised and unsupervised learning. We prove that our method's prediction converges in probability to the true performance of a counterfactual policy as the sample size increases. We validate our method with experiments on partly and entirely deterministic logging policies. Finally, we apply it to evaluate coupon targeting policies by a major online platform and show how to improve the existing policy.

Authors

Keywords

  • APP: Business/Marketing/Advertising/E-Commerce
  • ML: Causal Learning
  • ML: Online Learning & Bandits
  • ML: Reinforcement Learning Algorithms
  • ML: Reinforcement Learning Theory

Context

Venue
AAAI Conference on Artificial Intelligence
Archive span
1980-2026
Indexed papers
28718
Paper id
404544186518364874
v2026.09.13