Arrow Research search

Author name cluster

Chi-Guhn Lee

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

17 papers
2 author rows

Possible papers

17

AAAI Conference 2026 Conference Paper

A Differential Perspective on Distributional Reinforcement Learning

  • Juan Sebastian Rojas
  • Chi-Guhn Lee

To date, distributional reinforcement learning (distributional RL) methods have exclusively focused on the discounted setting, where an agent aims to optimize a discounted sum of rewards over time. In this work, we extend distributional RL to the average-reward setting, where an agent aims to optimize the reward received per time step. In particular, we utilize a quantile-based approach to develop the first set of algorithms that can successfully learn and/or optimize the long-run per-step reward distribution, as well as the differential return distribution of an average-reward MDP. We derive proven-convergent tabular algorithms for both prediction and control, as well as a broader family of algorithms that have appealing scaling properties. Empirically, we find that these algorithms yield competitive and sometimes superior performance when compared to their non-distributional equivalents, while also capturing rich information about the long-run per-step reward and differential return distributions.

AAMAS Conference 2026 Conference Paper

Algorithmic Contract Design with Reinforcement Learning Agents

  • David Molina Concha
  • Kyeonghyeon Park
  • Hyun-Rok Lee
  • Taesik Lee
  • Chi-Guhn Lee

Designing incentive mechanisms for multi-agent systems in stochastic and dynamic environments is a critical challenge, as system outcomes emerge from the complex interplay of agent learning and environmentaluncertainty. Existingprincipal–multi-agentcontract design methods often assume static settings or ignore learning dynamics, limiting their applicability in multi-agent reinforcement learning (MARL). Furthermore, the contract design space is highly constrained by feasibility requirements, such as individual rationality and incentive compatibility, making it difficult to explore. Weintroducetheprincipal-MARLcontractdesignproblem, where a principal optimizes both recruitment and incentive contracts evaluated via MARL. To address this problem, we propose Constrained ParetoMaximumEntropySearch(cPMES), amulti-objectiveBayesian optimization framework that treats feasibility as an explicit objective and selects designs based on information gain over the Pareto front. Experiments in social dilemma environments demonstrate thatcPMESefficientlyidentifiesfeasible, high-performingcontracts, significantly improving coordination and system-level rewards.

ECAI Conference 2025 Conference Paper

A Contrastive Diffusion-Based Network (CDNet) for Time Series Classification

  • Yaoyu Zhang
  • Chi-Guhn Lee

Deep learning models are widely used for time series classification (TSC) due to their scalability and efficiency. However, their performance degrades under challenging data conditions such as class similarity, multimodal distributions, and noise. To address these limitations, we propose CDNet, a Contrastive Diffusion-based Network that enhances existing classifiers by generating informative positive and negative samples via a learned diffusion process. Unlike traditional diffusion models that denoise individual samples, CDNet learns transitions between samples—both within and across classes—through convolutional approximations of reverse diffusion steps. We introduce a theoretically grounded CNN-based mechanism to enable both denoising and mode coverage, and incorporate an uncertainty-weighted composite loss for robust training. Extensive experiments on the UCR Archive and simulated datasets demonstrate that CDNet significantly improves state-of-the-art (SOTA) deep learning classifiers, particularly under noisy, similar, and multimodal conditions.

RLC Conference 2025 Conference Paper

Burning RED: Unlocking Subtask-Driven Reinforcement Learning and Risk-Awareness in Average-Reward Markov Decision Processes

  • Juan Sebastian Rojas
  • Chi-Guhn Lee

Average-reward Markov decision processes (MDPs) provide a foundational framework for sequential decision-making under uncertainty. However, average-reward MDPs have remained largely unexplored in reinforcement learning (RL) settings, with the majority of RL-based efforts having been allocated to discounted MDPs. In this work, we study a unique structural property of average-reward MDPs and utilize it to introduce Reward-Extended Differential (or RED) reinforcement learning: a novel RL framework that can be used to effectively and efficiently solve various learning objectives, or subtasks, simultaneously in the average-reward setting. We introduce a family of RED learning algorithms for prediction and control, including proven-convergent algorithms for the tabular case. We then showcase the power of these algorithms by demonstrating how they can be used to learn a policy that optimizes, for the first time, the well-known conditional value-at-risk (CVaR) risk measure in a fully-online manner, without the use of an explicit bi-level optimization scheme or an augmented state-space.

RLJ Journal 2025 Journal Article

Burning RED: Unlocking Subtask-Driven Reinforcement Learning and Risk-Awareness in Average-Reward Markov Decision Processes

  • Juan Sebastian Rojas
  • Chi-Guhn Lee

Average-reward Markov decision processes (MDPs) provide a foundational framework for sequential decision-making under uncertainty. However, average-reward MDPs have remained largely unexplored in reinforcement learning (RL) settings, with the majority of RL-based efforts having been allocated to discounted MDPs. In this work, we study a unique structural property of average-reward MDPs and utilize it to introduce Reward-Extended Differential (or RED) reinforcement learning: a novel RL framework that can be used to effectively and efficiently solve various learning objectives, or subtasks, simultaneously in the average-reward setting. We introduce a family of RED learning algorithms for prediction and control, including proven-convergent algorithms for the tabular case. We then showcase the power of these algorithms by demonstrating how they can be used to learn a policy that optimizes, for the first time, the well-known conditional value-at-risk (CVaR) risk measure in a fully-online manner, without the use of an explicit bi-level optimization scheme or an augmented state-space.

EAAI Journal 2024 Journal Article

Alleviating confirmation bias in perpetually dynamic environments: Continuous unsupervised domain adaptation-based condition monitoring (CUDACoM)

  • Mohamed Abubakr Hassan
  • Chi-Guhn Lee

Motivation Deep learning (DL) has revolutionized condition monitoring (CoM) in mechanical systems by reducing manual signal processing. However, DL's industrial integration is limited due to low robustness against distribution shifts. Existing CoM approaches focus on adaptating to one-time distribution shift. We introduce Continuous Unsupervised Domain Adaptation-based CoM (CUDACoM), tackling continuous distribution shifts in systems under perpetually dynamic conditions. Methodology CUDACoM mitigates confirmation bias, detrimental in long-domain sequences, by introducing two novel strategies (1) Fresh Initialization and (2) In-Domain Pseudo Labeling. Fresh Initialization maintains high model plasticity, while In-Domain pseudo-labeling improves pseudo-label accuracy, enhancing model adaptability. These strategies reduce confirmation bias, crucial for robust self-training, making CUDACoM ideal for long-sequence domain adaptation in perpetually dynamic environments. Results CUDACoM outperforms state-of-the-art (SOTA) adversarial and self-training approaches. Validated through two practical case studies: a 200% (RPM) change and gradual sensor degradation across 40 noise levels. These challenging case studies show stronger data shifts than the commonly used standard benchmarking datasets. The second case is especially novel, formulating robustness to noise as a domain adaptation problem. CUDACoM achieved a test accuracy of 0. 937 in the RPM case (vs. SOTA's 0. 770) and 0. 849 in the sensor degradation case (vs. SOTA's 0. 751). Impact This study addresses the overlooked challenges of employing DL for CoM in perpetually dynamic environments, particularly, the confirmation bias. With CUDACoM's computational efficiency, we provide a practical solution that enhances reliability by facilitating the integration of robust DL into industrial CoM systems.

PRL Workshop 2023 Workshop Paper

A Learnable Similarity Metric for Transfer Learning with Dynamics Mismatch

  • Ram Ananth Sreenivasan
  • Hyun-Rok Lee
  • Yeonjeong Jeong
  • Jongseong Jang
  • Dongsub Shim
  • Chi-Guhn Lee

When transferring knowledge from previously mastered source tasks to a new target task, the similarity between the source and target tasks can play a key role in whether such transfer is beneficial or harmful. In this paper, we develop an upper-bound of difference in action value function of source and target tasks with dynamics mismatch, and use the bound as a metric for dissimilarity between two tasks. The proposed metric does not require additional samples and adds little extra computation to the reinforcement learning algorithm for the target task. Also, the metric is highly portable so that it can be integrated into a wide range of algorithms. We showcase the effectiveness of the metric by incorporating it as a gatekeeper in the knowledge transfer step of transfer reinforcement learning algorithms. Numerical results on a suite of transfer learning scenarios demonstrate the benefits of preventing negative transfer in case of severe mismatch while accelerating learning otherwise

ICLR Conference 2023 Conference Paper

Recursive Time Series Data Augmentation

  • Amine Mohamed Aboussalah
  • Min-Jae Kwon
  • Raj G. Patel
  • Cheng Chi
  • Chi-Guhn Lee

Time series observations can be seen as realizations of an underlying dynamical system governed by rules that we typically do not know. For time series learning tasks we create our model using available data. Training on available realizations, where data is limited, often induces severe over-fitting thereby preventing generalization. To address this issue, we introduce a general recursive framework for time series augmentation, which we call the Recursive Interpolation Method (RIM). New augmented time series are generated using a recursive interpolation function from the original time series for use in training. We perform theoretical analysis to characterize the proposed RIM and to guarantee its performance under certain conditions. We apply RIM to diverse synthetic and real-world time series cases to achieve strong performance over non-augmented data on a variety of learning tasks. Our method is also computationally more efficient and leads to better performance when compared to state of the art time series data augmentation.

IJCAI Conference 2022 Conference Paper

Multi-policy Grounding and Ensemble Policy Learning for Transfer Learning with Dynamics Mismatch

  • Hyun-Rok Lee
  • Ram Ananth Sreenivasan
  • Yeonjeong Jeong
  • Jongseong Jang
  • Dongsub Shim
  • Chi-Guhn Lee

We propose a new transfer learning algorithm between tasks with different dynamics. The proposed algorithm solves an Imitation from Observation problem (IfO) to ground the source environment to the target task before learning an optimal policy in the grounded environment. The learned policy is deployed in the target task without additional training. A particular feature of our algorithm is the employment of multiple rollout policies during training with a goal to ground the environment more globally; hence, it is named as Multi-Policy Grounding (MPG). The quality of final policy is further enhanced via ensemble policy learning. We demonstrate the superiority of the proposed algorithm analytically and numerically. Numerical studies show that the proposed multi-policy approach allows comparable grounding with single policy approach with a fraction of target samples, hence the algorithm is able to maintain the quality of obtained policy even as the number of interactions with the target environment becomes extremely small.

IROS Conference 2021 Conference Paper

A Marginal Log-Likelihood Approach for the Estimation of Discount Factors of Multiple Experts in Inverse Reinforcement Learning

  • Babatunde H. Giwa
  • Chi-Guhn Lee

We focus on multiple experts performing a task in a Markov decision process (MDP) environment. A probabilistic assignment of trajectories to clusters and a mathematical framework which leverages the utility function are employed to jointly estimate the discount factor and reward. We treat the number of clusters as a hyperparameter which can be "freely" selected by the problem designer. In this work, we specifically treat the cluster of trajectories as a latent variable in the adapted maximum entropy inverse reinforcement learning (IRL) formulation; the introduction of this latent variable adds to the complexity of the IRL problem. To manage such complexity, we optimize a marginal log-likelihood function via Expectation Maximization. To test our approach, we have utilized behavioral data generated from three MDP environments. Experimental works show that our approach is promising towards the estimation of discount factors in IRL for non-interacting multiple experts.

IJCAI Conference 2021 Conference Paper

Bayesian Experience Reuse for Learning from Multiple Demonstrators

  • Mike Gimelfarb
  • Scott Sanner
  • Chi-Guhn Lee

Learning from Demonstrations (LfD) is a powerful approach for incorporating advice from experts in the form of demonstrations. However, demonstrations often come from multiple sub-optimal experts with conflicting goals, rendering them difficult to incorporate effectively in online settings. To address this, we formulate a quadratic program whose solution yields an adaptive weighting over experts, that can be used to sample experts with relevant goals. In order to compare different source and target task goals safely, we model their uncertainty using normal-inverse-gamma priors, whose posteriors are learned from demonstrations using Bayesian neural networks with a shared encoder. Our resulting approach, which we call Bayesian Experience Reuse, can be applied for LfD in static and dynamic decision-making settings. We demonstrate its effectiveness for minimizing multi-modal functions, and optimizing a high-dimensional supply chain with cost uncertainty, where it is also shown to improve upon the performance of the demonstrators' policies.

UAI Conference 2021 Conference Paper

Contextual policy transfer in reinforcement learning domains via deep mixtures-of-experts

  • Michael Gimelfarb
  • Scott Sanner
  • Chi-Guhn Lee

In reinforcement learning, agents that consider the context or current state when transferring source policies have been shown to outperform context-free approaches. However, existing approaches suffer from limitations, including sensitivity to sparse or delayed rewards and estimation errors in values. One important insight is that explicit learned models of the source dynamics, when available, could benefit contextual transfer in such settings. In this paper, we assume a family of tasks with shared sub-goals but different dynamics, and availability of estimated dynamics and policies for source tasks. To deal with possible estimation errors in dynamics, we introduce a novel Bayesian mixture-of-experts for learning state-dependent beliefs over source task dynamics that match the target dynamics using state transitions collected from the target task. The mixture is easy to interpret, is robust to estimation errors in dynamics, and is compatible with most RL algorithms. We incorporate it into standard policy reuse frameworks and demonstrate its effectiveness on benchmarks from OpenAI gym.

PRL Workshop 2021 Workshop Paper

Discount Factor Estimation in a Model-Based Inverse Reinforcement Learning Framework

  • Babatunde Giwa
  • Chi-Guhn Lee

We consider the crucial task of estimating an expert’s discount factor in Inverse Reinforcement Learning (IRL) to facilitate a better synthesis towards the resulting optimal policy. Existing IRL algorithms have significantly overlooked the vital need to estimate the discount factor, experimental studies and theoretical intuitions show variability of the learnt reward function as the discount factor changes. In this work, we adapt the model-based maximum entropy IRL framework and optimize a utility-based softmax likelihood function via a featurebased gradient update to jointly learn the discount factor and reward. To test our approach, we utilize behavioral data from three Markov decision process (MDP) environments, namely, Grid-World, Mountain-Car Driving and Object-World. Experimental and numerical studies show that our approach is viable for the simultaneous estimation of the discount factor and reward function in IRL.

ICLR Conference 2021 Conference Paper

Incremental few-shot learning via vector quantization in deep embedded space

  • Kuilin Chen
  • Chi-Guhn Lee

The capability of incrementally learning new tasks without forgetting old ones is a challenging problem due to catastrophic forgetting. This challenge becomes greater when novel tasks contain very few labelled training samples. Currently, most methods are dedicated to class-incremental learning and rely on sufficient training data to learn additional weights for newly added classes. Those methods cannot be easily extended to incremental regression tasks and could suffer from severe overfitting when learning few-shot novel tasks. In this study, we propose a nonparametric method in deep embedded space to tackle incremental few-shot learning problems. The knowledge about the learned tasks are compressed into a small number of quantized reference vectors. The proposed method learns new tasks sequentially by adding more reference vectors to the model using few-shot samples in each novel task. For classification problems, we employ the nearest neighbor scheme to make classification on sparsely available data and incorporate intra-class variation, less forgetting regularization and calibration of reference vectors to mitigate catastrophic forgetting. In addition, the proposed learning vector quantization (LVQ) in deep embedded space can be customized as a kernel smoother to handle incremental few-shot regression tasks. Experimental results demonstrate that the proposed method outperforms other state-of-the-art methods in incremental learning.

NeurIPS Conference 2021 Conference Paper

Risk-Aware Transfer in Reinforcement Learning using Successor Features

  • Michael Gimelfarb
  • Andre Barreto
  • Scott Sanner
  • Chi-Guhn Lee

Sample efficiency and risk-awareness are central to the development of practical reinforcement learning (RL) for complex decision-making. The former can be addressed by transfer learning, while the latter by optimizing some utility function of the return. However, the problem of transferring skills in a risk-aware manner is not well-understood. In this paper, we address the problem of transferring policies between tasks in a common domain that differ only in their reward functions, in which risk is measured by the variance of reward streams. Our approach begins by extending the idea of generalized policy improvement to maximize entropic utilities, thus extending the dynamic programming's policy improvement operation to sets of policies \emph{and} levels of risk-aversion. Next, we extend the idea of successor features (SF), a value function representation that decouples the environment dynamics from the rewards, to capture the variance of returns. Our resulting risk-aware successor features (RaSF) integrate seamlessly within the RL framework, inherit the superior task generalization ability of SFs, while incorporating risk into the decision-making. Experiments on a discrete navigation domain and control of a simulated robotic arm demonstrate the ability of RaSFs to outperform alternative methods including SFs, when taking the risk of the learned policies into account.

UAI Conference 2019 Conference Paper

Epsilon-BMC: A Bayesian Ensemble Approach to Epsilon-Greedy Exploration in Model-Free Reinforcement Learning

  • Michael Gimelfarb
  • Scott Sanner
  • Chi-Guhn Lee

Resolving the exploration-exploitation trade-off remains a fundamental problem in the design and implementation of reinforcement learning (RL) algorithms. In this paper, we focus on model-free RL using the epsilon-greedy exploration policy, which despite its simplicity, remains one of the most frequently used forms of exploration. However, a key limitation of this policy is the specification of epsilon. In this paper, we provide a novel Bayesian perspective of epsilon as a measure of the uncertainty (and hence convergence) in the Q-value function. We introduce a closed-form Bayesian model update based on Bayesian model combination (BMC), based on this new perspective, which allows us to adapt epsilon using experiences from the environment in constant time with monotone convergence guarantees. We demonstrate that our proposed algorithm, epsilon-BMC, efficiently balances exploration and exploitation on different problems, performing comparably or outperforming the best tuned fixed annealing schedules and an alternative data-dependent epsilon adaptation scheme proposed in the literature.

NeurIPS Conference 2018 Conference Paper

Reinforcement Learning with Multiple Experts: A Bayesian Model Combination Approach

  • Michael Gimelfarb
  • Scott Sanner
  • Chi-Guhn Lee

Potential based reward shaping is a powerful technique for accelerating convergence of reinforcement learning algorithms. Typically, such information includes an estimate of the optimal value function and is often provided by a human expert or other sources of domain knowledge. However, this information is often biased or inaccurate and can mislead many reinforcement learning algorithms. In this paper, we apply Bayesian Model Combination with multiple experts in a way that learns to trust a good combination of experts as training progresses. This approach is both computationally efficient and general, and is shown numerically to improve convergence across discrete and continuous domains and different reinforcement learning algorithms.

v2026.09.13