Arrow Research search

Author name cluster

Lloyd Greenwald

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

2 papers
1 author row

Possible papers

2

AAAI Conference 1999 Conference Paper

Efficient Exploration for Optimizing Immediate Reward

  • Dale Schuurmans
  • University of Waterloo
  • Lloyd Greenwald
  • Drexel University

Weconsider the problemof learning an effective behavior strategy from reward. Althoughmuchstudied, the issue of howto use prior knowledgeto scale optimal behavior learning up to real-world problems remains an important open issue. Weinvestigate the inherent data-complexity of behavior-learning whenthe goal is simply to optimize immediate reward. Although easier than reinforcement learning, whereone must also cope with state dynamics, immediate rewardlearning is still a common problemand is fundamentallyharder than supervised learning. For optimizing immediatereward, prior knowledgecan be expressed either as a bias on the space of possible reward models, or a bias on the space of possible controllers. Weinvestigate the two paradigmatic learning approachesof indirect (reward-model)learning anddirect-control learning, and showthat neither uniformlydominatesthe other in general. Model-based learning has the advantage of generalizing reward experiences across states and actions, but direct-control learning has the advantage of focusing only on potentially optimal actions and avoidinglearning irrelevant worlddetails. Bothstrategies can be strongly advantageousin different circumstances. Weintroduce hybrid learning strategies that combinethe benefits of both approaches, and uniformlyimprovetheir learning efficiency.

AAAI Conference 1994 Short Paper

Time-Critical Scheduling in Stochastic Domains

  • Lloyd Greenwald

In this work we look at extending the work of (Dean et al. 1993) to handle more complicated scheduling problems in which the sources of complexity stem not only from large state spaces but from large action spaces as well. In these problems it is no longer tractable to compute optimal policies for restricted state spaces via policy iteration. We, instead, borrow from operations research in applying bottleneck-centered scheduling heuristics (Adams et al. 1988). Additionally, our techniques draw from the work of (Drummond and Bresina 1990).

v2026.09.13