Arrow Research search

Author name cluster

Bruno da Silva

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
1 author row

Possible papers

6

NeurIPS Conference 2022 Conference Paper

Off-Policy Evaluation for Action-Dependent Non-stationary Environments

  • Yash Chandak
  • Shiv Shankar
  • Nathaniel Bastian
  • Bruno da Silva
  • Emma Brunskill
  • Philip S. Thomas

Methods for sequential decision-making are often built upon a foundational assumption that the underlying decision process is stationary. This limits the application of such methods because real-world problems are often subject to changes due to external factors (\textit{passive} non-stationarity), changes induced by interactions with the system itself (\textit{active} non-stationarity), or both (\textit{hybrid} non-stationarity). In this work, we take the first steps towards the fundamental challenge of on-policy and off-policy evaluation amidst structured changes due to active, passive, or hybrid non-stationarity. Towards this goal, we make a \textit{higher-order stationarity} assumption such that non-stationarity results in changes over time, but the way changes happen is fixed. We propose, OPEN, an algorithm that uses a double application of counterfactual reasoning and a novel importance-weighted instrument-variable regression to obtain both a lower bias and a lower variance estimate of the structure in the changes of a policy's past performances. Finally, we show promising results on how OPEN can be used to predict future performances for several domains inspired by real-world applications that exhibit non-stationarity.

NeurIPS Conference 2021 Conference Paper

Universal Off-Policy Evaluation

  • Yash Chandak
  • Scott Niekum
  • Bruno da Silva
  • Erik Learned-Miller
  • Emma Brunskill
  • Philip S. Thomas

When faced with sequential decision-making problems, it is often useful to be able to predict what would happen if decisions were made using a new policy. Those predictions must often be based on data collected under some previously used decision-making rule. Many previous methods enable such off-policy (or counterfactual) estimation of the expected value of a performance measure called the return. In this paper, we take the first steps towards a 'universal off-policy estimator' (UnO)---one that provides off-policy estimates and high-confidence bounds for any parameter of the return distribution. We use UnO for estimating and simultaneously bounding the mean, variance, quantiles/median, inter-quantile range, CVaR, and the entire cumulative distribution of returns. Finally, we also discuss UnO's applicability in various settings, including fully observable, partially observable (i. e. , with unobserved confounders), Markovian, non-Markovian, stationary, smoothly non-stationary, and discrete distribution shifts.

AAMAS Conference 2013 Conference Paper

Biasing the Behavior of Organizationally Adept Agents

  • Daniel Corkill
  • Chongjie Zhang
  • Bruno da Silva
  • Yoonheui Kim
  • Daniel Garant
  • Victor R. Lesser
  • Xiaoqin Zhang

An organizationally adept agent (OAA) adjusts its behavior when given annotated organizational guidelines. More importantly, it can also determine when such guidelines become ineffective and proactively adapt its behavior to better achieve organizational objectives. We present the high-level aspects of this architecture and analyze its effectiveness using call-center OAAs striving to extinguish fires in RoboCup Rescue scenarios.

AAAI Conference 2012 Conference Paper

TD-DeltaPi: A Model-Free Algorithm for Efficient Exploration

  • Bruno da Silva
  • Andrew Barto

We study the problem of finding efficient exploration policies for the case in which an agent is momentarily not concerned with exploiting, and instead tries to compute a policy for later use. We first formally define the Optimal Exploration Problem as one of sequential sampling and show that its solutions correspond to paths of minimum expected length in the space of policies. We derive a model-free, local linear approximation to such solutions and use it to construct efficient exploration policies. We compare our model-free approach to other exploration techniques, including one with the best known PAC bounds, and show that ours is both based on a well-defined optimization problem and empirically efficient.

AAMAS Conference 2010 Conference Paper

Using Spatial Hints to Improve Policy Reuse in a Reinforcement Learning Agent

  • Bruno da Silva
  • Alan Mackworth

We study the problem of knowledge reuse by a reinforcementlearning agent. We are interested in how an agent can exploitpolicies that were learned in the past to learn a new task moreefficiently in the present. Our approach is to elicit spatial hintsfrom an expert suggesting the world states in which each existingpolicy should be more relevant to the new task. By using thesehints with domain exploration, the agent is able to detect thoseportions of existing policies that are beneficial to the new task, therefore learning a new policy more efficiently. We call ourapproach Spatial Hints Policy Reuse (SHPR). Experimentsdemonstrate the effectiveness and robustness of our method. Ourresults encourage further study investigating how much moreefficacy can be gained from the elicitation of very simple advicefrom humans.

AAAI Conference 2007 Conference Paper

KA-CAPTCHA: An Opportunity for Knowledge Acquisition on the Web

  • Bruno da Silva

Any Web user is a potential knowledge contributor, but it remains a challenge to make them devote their time contributing to some purpose. In order to align individual with social interests, we selected the CAPTCHA Web resource protection application to embed knowledge elicitation within the users' main task of accessing a Web resource. Consequently, unlike previous knowledge acquisition approaches, no extra effort is expected from users since they are already willing to use a CAPTCHA to perform some particular task. We present an application where we extract pictorial knowledge from Web users, and experiments suggest that our approach enables knowledge acquisition while still satisfying CAPTCHA's security requirements.

v2026.09.13