Arrow Research search

Author name cluster

Sangwon Seo

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
1 author row

Possible papers

8

AAMAS Conference 2026 Conference Paper

Hierarchical Reward Design from Language: Enhancing Alignment of Agent Behavior with Human Specifications

  • Zhiqin Qian
  • Ryan Diaz
  • Sangwon Seo
  • Vaibhav Unhelkar

When training artificial intelligence (AI) to perform tasks, humans oftencarenotonlyaboutwhether ataskiscompletedbutalsohow it is performed. As AI agents tackle increasingly complex tasks, aligning their behavior with human-provided specifications becomes critical for responsible AI deployment. Reward design provides a direct channel for such alignment by translating human expectations into reward functions that guide reinforcement learning (RL). However, existing methods are often too limited to capture nuanced human preferences that arise in long-horizon tasks. Hence, we introduce Hierarchical Reward Design from Language (HRDL): a problem formulation that extends classical reward design to encode richer behavioral specifications for hierarchical RL agents. We further propose Language to Hierarchical Rewards (L2HR) as a solution to HRDL. Experiments show that AI agents trained with rewards designed via L2HR not only complete tasks effectively but also better adhere to human specifications. Together, HRDL and L2HR advance the research on human-aligned AI agents.

AAMAS Conference 2025 Conference Paper

$Socratic: $ Enhancing Human Teamwork via AI-enabled Coaching

  • Sangwon Seo
  • Bing Han
  • Rayan E. Harari
  • Roger D. Dias
  • Marco A. Zenati
  • Eduardo Salas
  • Vaibhav Unhelkar

Coaches are vital for effective collaboration, but cost and resource constraints often limit their availability during real-world tasks. This limitation poses serious challenges in life-critical domains that rely on effective teamwork, such as healthcare and disaster response. To address this gap, we propose and realize an innovative application of AI: task-time team coaching. Specifically, we introduce Socratic, a novel AI system that complements human coaches by providing real-time guidance during task execution. Socratic monitors team behavior, detects misalignments in team members’ shared understanding, and delivers automated interventions to improve team performance. We validated Socratic through two human subject experiments involving dyadic collaboration. The results demonstrate that the system significantly enhances team performance with minimal interventions. Participants also perceived Socratic as helpful and trustworthy, supporting its potential for adoption. Our findings also suggest promising directions both for AI research and its practical applications to enhance human teamwork.

AAMAS Conference 2025 Conference Paper

Hierarchical Imitation Learning of Team Behavior from Heterogeneous Demonstrations

  • Sangwon Seo
  • Vaibhav Unhelkar

Successful collaboration requires team members to stay aligned, especially in complex sequential tasks. Team members must dynamically coordinate which subtasks to perform and in what order. However, real-world constraints like partial observability and limited communication bandwidth often lead to suboptimal collaboration. Even among expert teams, the same task can be executed in multiple ways. To develop multi-agent systems and human-AI teams for such tasks, we are interested in data-driven learning of multimodal team behaviors. Multi-Agent Imitation Learning (MAIL) provides a promising framework for data-driven learning of team behavior from demonstrations, but existing methods struggle with heterogeneous demonstrations, as they assume that all demonstrations originate from a single team policy. Hence, in this work, we introduce DTIL: a hierarchical MAIL algorithm designed to learn multimodal team behaviors in complex sequential tasks. DTIL represents each team member with a hierarchical policy and learns these policies from heterogeneous team demonstrations in a factored manner. By employing a distribution-matching approach, DTIL mitigates compounding errors and scales effectively to long horizons and continuous state representations. Experimental results show that DTIL outperforms MAIL baselines and accurately models team behavior across a variety of collaborative scenarios.

AAAI Conference 2024 Short Paper

AI-Assisted Human Teamwork

  • Sangwon Seo

Effective teamwork translates to fewer preventable errors and higher task performance in collaborative tasks. However, in time-critical tasks, successful teamwork becomes highly challenging to attain. In such settings, often, team members have partial observability of their surroundings, incur high cost of communication, and have trouble estimating the state and intent of their teammates. To assist a team in improving teamwork at task time, my doctoral research proposes an automated task-time team intervention system. Grounded in the notion of shared mental models, the system first detects whether the team is on the same page or not. It then generates effective interventions to improve teamwork. Additionally, by leveraging past demonstrations to learn a model of team behavior, this system minimizes the need for domain experts to specify teamwork models and rules.

AAMAS Conference 2024 Conference Paper

IDIL: Imitation Learning of Intent-Driven Expert Behavior

  • Sangwon Seo
  • Vaibhav Unhelkar

When faced with accomplishing a task, human experts exhibit intentional behavior. Their unique intents shape their plans and decisions, resulting in experts demonstrating diverse behaviors to accomplish the same task. Due to the uncertainties encountered in the real world and their bounded rationality, experts sometimes adjust their intents, which in turn influences their behaviors during task execution. This paper introduces IDIL, a novel imitation learning algorithm to mimic these diverse intent-driven behaviors of experts. Iteratively, our approach estimates expert intent from heterogeneous demonstrations and then uses it to learn an intent-aware model of their behavior. Unlike contemporary approaches, IDIL is capable of addressing sequential tasks with high-dimensional state representations, while sidestepping the complexities and drawbacks associated with adversarial training (a mainstay of related techniques). Our empirical results suggest that the models generated by IDIL either match or surpass those produced by recent imitation learning benchmarks in metrics of task performance. Moreover, as it creates a generative model, IDIL demonstrates superior performance in intent inference metrics, crucial for human-agent interactions, and aptly captures a broad spectrum of expert behaviors.

AAMAS Conference 2023 Conference Paper

Automated Task-Time Interventions to Improve Teamwork using Imitation Learning

  • Sangwon Seo
  • Bing Han
  • Vaibhav Unhelkar

Effective human-human and human-autonomy teamwork is critical but often challenging to perfect. The challenge is particularly relevant in time-critical domains, such as healthcare and disaster response, where the time pressures can make coordination increasingly difficult to achieve and the consequences of imperfect coordination can be severe. To improve teamwork in these and other domains, we present TIC: an automated intervention approach for improving coordination between team members. Using BTIL, a multi-agent imitation learning algorithm, our approach first learns a generative model of team behavior from past task execution data. Next, it utilizes the learned generative model and team’s task objective (shared reward) to algorithmically generate execution-time interventions. We evaluate our approach in synthetic multi-agent teaming scenarios, where team members make decentralized decisions without full observability of the environment. The experiments demonstrate that the automated interventions can successfully improve team performance and shed light on the design of autonomous agents for improving teamwork.

IJCAI Conference 2022 Conference Paper

Semi-Supervised Imitation Learning of Team Policies from Suboptimal Demonstrations

  • Sangwon Seo
  • Vaibhav V. Unhelkar

We present Bayesian Team Imitation Learner (BTIL), an imitation learning algorithm to model the behavior of teams performing sequential tasks in Markovian domains. In contrast to existing multi-agent imitation learning techniques, BTIL explicitly models and infers the time-varying mental states of team members, thereby enabling learning of decentralized team policies from demonstrations of suboptimal teamwork. Further, to allow for sample- and label-efficient policy learning from small datasets, BTIL employs a Bayesian perspective and is capable of learning from semi-supervised demonstrations. We demonstrate and benchmark the performance of BTIL on synthetic multi-agent tasks as well as a novel dataset of human-agent teamwork. Our experiments show that BTIL can successfully learn team policies from demonstrations despite the influence of team members' (time-varying and potentially misaligned) mental states on their behavior.

JBHI Journal 2017 Journal Article

Sleep Period Time Estimation Based on Electrodermal Activity

  • Su Hwan Hwang
  • Kwang Suk Park
  • Sangwon Seo
  • Hee Nam Yoon
  • Da Woon Jung
  • Hyun Jae Baek
  • Jaegeol Cho
  • Jae Won Choi

We proposed and tested a method to estimate sleep period time (SPT) using electrodermal activity (EDA) signals. Eight healthy subjects and six obstructive sleep apnea patients participated in the experiments. Each subject's EDA signals were measured at the middle and ring fingers of the dominant hand during polysomnography (PSG). For nine of the 17 participants, wrist actigraphy was also measured for a quantitative comparison of EDA- and actigraphy-based methods. Based on the training data, we observed that sleep onset was accompanied by a gradual reduction of amplitude of the EDA signals, whereas sleep offset was accompanied by a rapid increase in amplitude of EDA signals. We developed a method based on these EDA fluctuations during sleep-wake transitions, and applied it to a test dataset. The performance of the method was assessed by comparing its results with those from a physician's sleep stage scores. The mean absolute errors in the obtained values for sleep onset, offset, and period time between the proposed method, and the results of the PSG were 4. 1, 3. 0, and 6. 1 min, respectively. Furthermore, there were no significant differences in the corresponding values between the methods. We compared these results with those obtained by applying actigraphic methods, and found that our algorithm outperformed these in terms of each estimated parameter of interest in SPT estimation. Long awakening periods were also detected based on sympathetic responses reflected in the EDA signals. The proposed method can be applied to a daily sleep monitoring system.

v2026.09.13