Arrow Research search

Author name cluster

Baher Abdulhai

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

2 papers
2 author rows

Possible papers

2

ICLR Conference 2023 Conference Paper

Conservative Bayesian Model-Based Value Expansion for Offline Policy Optimization

  • Jihwan Jeong
  • Xiaoyu Wang 0018
  • Michael Gimelfarb
  • Hyunwoo Kim
  • Baher Abdulhai
  • Scott Sanner

Offline reinforcement learning (RL) addresses the problem of learning a performant policy from a fixed batch of data collected by following some behavior policy. Model-based approaches are particularly appealing in the offline setting since they can extract more learning signals from the logged dataset by learning a model of the environment. However, the performance of existing model-based approaches falls short of model-free counterparts, due to the compounding of estimation errors in the learned model. Driven by this observation, we argue that it is critical for a model-based method to understand when to trust the model and when to rely on model-free estimates, and how to act conservatively w.r.t. both. To this end, we derive an elegant and simple methodology called conservative Bayesian model-based value expansion for offline policy optimization (CBOP), that trades off model-free and model-based estimates during the policy evaluation step according to their epistemic uncertainties, and facilitates conservatism by taking a lower bound on the Bayesian posterior value estimate. On the standard D4RL continuous control tasks, we find that our method significantly outperforms previous model-based approaches: e.g., MOPO by $116.4$%, MOReL by $23.2$% and COMBO by $23.7$%. Further, CBOP achieves state-of-the-art performance on $11$ out of $18$ benchmark datasets while doing on par on the remaining datasets.

EAAI Journal 2012 Journal Article

Forecasting of short-term traffic-flow based on improved neurofuzzy models via emotional temporal difference learning algorithm

  • Javad Abdi
  • Behzad Moshiri
  • Baher Abdulhai
  • Ali Khaki Sedigh

Bounded rationally idea, rather that optimization idea, have result and better performance in decision making theory. Bounded rationality is the idea in decision making, rationality of individuals is limited by the information they have, the cognitive limitations of their minds, and the finite amount of time they have to make decisions. The emotional theory is an important topic presented in this field. The new methods in the direction of purposeful forecasting issues, which are based on cognitive limitations, are presented in this study. The presented algorithms in this study are emphasizes to rectify the learning the peak points, to increase the forecasting accuracy, to decrease the computational time and comply the multi-object forecasting in the algorithms. The structure of the proposed algorithms is based on approximation of its current estimate according to previously learned estimates. The short term traffic flow forecasting is a real benchmark that has been studied in this area. Traffic flow is a good measure of traffic activity. The time-series data used for fitting the proposed models are obtained from a two lane street I-494 in Minnesota City, USA. The research discuss the strong points of new method based on neurofuzzy and limbic system structure such as Locally Linear Neurofuzzy network (LLNF) and Brain Emotional Learning Based Intelligent Controller (BELBIC) models against classical and other intelligent methods such as Radial Basis Function (RBF), Takagi–Sugeno (T–S) neurofuzzy, and Multi-Layer Perceptron (MLP), and the effect of noise on the performance of the models is also considered. Finally, findings confirmed the significance of structural brain modeling beyond the classical artificial neural networks.

v2026.09.13