Arrow Research search

Author name cluster

Sangmin Lee

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

7 papers
1 author row

Possible papers

7

EAAI Journal 2026 Journal Article

HarmonyRouting with traffic impact prediction based on graph neural network for large-scale semiconductor fabrication

  • Younkook Kang
  • Kwangyoung Im
  • Sangmin Lee
  • Sungzoon Cho

Recently, automated material handling systems, which operate without human intervention, have been developed for smart semiconductor factories. An overhead hoist transport (OHT), a vehicle robot that transfers carriers between production equipment along a railway network, is an effective tool for automated material handling systems. To improve routing decisions, many studies have attempted to predict the traffic of OHTs, mainly focusing on routes from an individual OHT perspective because of the associated complexity. To overcome the limitations of existing studies, we propose HarmonyRouting, which is a prediction-based system-perspective routing method designed to optimize overall traffic flow by balancing individual vehicle routes. First, a partitioned attention-based graph neural network was proposed to efficiently and effectively predict traffic for large-scale fabrication plant layouts. Second, the concept of vehicle traffic impact caused by congestion on a path was defined and modelled. Third, traffic impact value was applied to the routing model for system-view decisions. Its architecture was designed to enhance practical application. The performance of the model was evaluated using simulations with the actual fabrication plant layout and the largest number of OHTs ever studied. Our model outperformed other models in terms of average delivery time and delay. Our model represents an innovative approach towards realizing a fully autonomous manufacturing factory.

AAAI Conference 2025 Conference Paper

LAMA-UT: Language Agnostic Multilingual ASR Through Orthography Unification and Language-Specific Transliteration

  • Sangmin Lee
  • Woojin Chung
  • Hong-Goo Kang

Building a universal multilingual automatic speech recognition (ASR) model that performs equitably across languages has long been a challenge due to its inherent difficulties. To address this task we introduce a Language-Agnostic Multilingual ASR pipeline through orthography Unification and language-specific Transliteration (LAMA-UT). LAMA-UT operates without any language-specific modules while matching the performance of state-of-the-art models trained on a minimal amount of data. Our pipeline consists of two key steps. First, we utilize a universal transcription generator to unify orthographic features into Romanized form and capture common phonetic characteristics across diverse languages. Second, we utilize a universal converter to transform these universal transcriptions into language-specific ones. In experiments, we demonstrate the effectiveness of our proposed method leveraging universal transcriptions for massively multilingual ASR. Our pipeline achieves a relative error reduction rate of 45% when compared to Whisper and performs comparably to MMS, despite being trained on only 0.1% of Whisper's training data. Furthermore, our pipeline does not rely on any language-specific modules. However, it performs on par with zero-shot ASR approaches which utilize additional language-specific lexicons and language models. We expect this framework to serve as a cornerstone for flexible multilingual ASR systems that are generalizable even to unseen languages.

NeurIPS Conference 2025 Conference Paper

Toward Human Deictic Gesture Target Estimation

  • Xu Cao
  • Pranav Virupaksha
  • Sangmin Lee
  • Bolin Lai
  • Wenqi Jia
  • Jintai Chen
  • James Rehg

Humans have a remarkable ability to use co-speech deictic gestures, such as pointing and showing, to enrich verbal communication and support social interaction. These gestures are so fundamental that infants begin to use them even before they acquire spoken language, which highlights their central role in human communication. Understanding the intended targets of another individual's deictic gestures enables inference of their intentions, comprehension of their current actions, and prediction of upcoming behaviors. Despite its significance, gesture target estimation remains an underexplored task within the computer vision community. In this paper, we introduce GestureTarget, a novel task designed specifically for comprehensive evaluation of social deictic gesture semantic target estimation. To address this task, we propose TransGesture, a set of Transformer-based gesture target prediction models. Given an input image and the spatial location of a person, our models predict the intended target of their gesture within the scene. Critically, our gaze-aware joint cross attention fusion model demonstrates how incorporating gaze-following cues significantly improves gesture target mask prediction IoU by 6% and gesture existence prediction accuracy by 10%. Our results underscore the complexity and importance of integrating gaze cues into deictic gesture intention understanding, advocating for increased research attention to this emerging area. All data, code will be made publicly available upon acceptance. Code of TransGesture is available at GitHub. com/IrohXu/TransGesture.

AAAI Conference 2025 Conference Paper

Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection

  • Sung Jin Um
  • Dongjin Kim
  • Sangmin Lee
  • Jung Uk Kim

The goal of video moment retrieval and highlight detection is to identify specific segments and highlights based on a given text query. With the rapid growth of video content and the overlap between these tasks, recent works have addressed both simultaneously. However, they still struggle to fully capture the overall video context, making it challenging to determine which words are most relevant. In this paper, we present a novel Video Context-aware Keyword Attention module that overcomes this limitation by capturing keyword variation within the context of the entire video. To achieve this, we introduce a video context clustering module that provides concise representations of the overall video context, thereby enhancing the understanding of keyword dynamics. Furthermore, we propose a keyword weight detection module with keyword-aware contrastive learning that incorporates keyword information to enhance fine-grained alignment between visual and textual features. Extensive experiments on the QVHighlights, TVSum, and Charades-STA benchmarks demonstrate that our proposed method significantly improves performance in moment retrieval and highlight detection tasks compared to existing approaches.

AAAI Conference 2021 Conference Paper

Explaining Convolutional Neural Networks through Attribution-Based Input Sampling and Block-Wise Feature Aggregation

  • Sam Sattarzadeh
  • Mahesh Sudhakar
  • Anthony Lem
  • Shervin Mehryar
  • Konstantinos N Plataniotis
  • Jongseong Jang
  • Hyunwoo Kim
  • Yeonjeong Jeong

As an emerging field in Machine Learning, Explainable AI (XAI) has been offering remarkable performance in interpreting the decisions made by Convolutional Neural Networks (CNNs). To achieve visual explanations for CNNs, methods based on class activation mapping and randomized input sampling have gained great popularity. However, the attribution methods based on these techniques provide low-resolution and blurry explanation maps that limit their explanation ability. To circumvent this issue, visualization based on various layers is sought. In this work, we collect visualization maps from multiple layers of the model based on an attributionbased input sampling technique and aggregate them to reach a fine-grained and complete explanation. We also propose a layer selection strategy that applies to the whole family of CNN-based models, based on which our extraction framework is applied to visualize the last layers of each convolutional block of the model. Moreover, we perform an empirical analysis of the efficacy of derived lower-level information to enhance the represented attributions. Comprehensive experiments conducted on shallow and deep models trained on natural and industrial datasets, using both ground-truth and model-truth based evaluation metrics validate our proposed algorithm by meeting or outperforming the state-of-the-art methods in terms of explanation ability and visual quality, demonstrating that our method shows stability regardless of the size of objects or instances to be explained.

AAAI Conference 2021 Conference Paper

Towards a Better Understanding of VR Sickness: Physical Symptom Prediction for VR Contents

  • Hak Gu Kim
  • Sangmin Lee
  • Seongyeop Kim
  • Heoun-taek Lim
  • Yong Man Ro

We address the black-box issue of VR sickness assessment (VRSA) by evaluating the level of physical symptoms of VR sickness. For the VR contents inducing the similar VR sickness level, the physical symptoms can vary depending on the characteristics of the contents. Most of existing VRSA methods focused on assessing the overall VR sickness score. To make better understanding of VR sickness, it is required to predict and provide the level of major symptoms of VR sickness rather than overall degree of VR sickness. In this paper, we predict the degrees of main physical symptoms affecting the overall degree of VR sickness, which are disorientation, nausea, and oculomotor. In addition, we introduce a new large-scale dataset for VRSA including 360 videos with various frame rates, physiological signals, and subjective scores. On VRSA benchmark and our newly collected dataset, our approach shows a potential to not only achieve the highest correlation with subjective scores, but also to better understand which symptoms are the main causes of VR sickness.

AAAI Conference 2021 Conference Paper

Visual Comfort Aware-Reinforcement Learning for Depth Adjustment of Stereoscopic 3D Images

  • Hak Gu Kim
  • Minho Park
  • Sangmin Lee
  • Seongyeop Kim
  • Yong Man Ro

Depth adjustment aims to enhance the visual experience of stereoscopic 3D (S3D) images, which accompanied with improving visual comfort and depth perception. For a human expert, the depth adjustment procedure is a sequence of iterative decision making. The human expert iteratively adjusts the depth until he is satisfied with the both levels of visual comfort and the perceived depth. In this work, we present a novel deep reinforcement learning (DRL)-based approach for depth adjustment named VCA-RL (Visual Comfort Aware Reinforcement Learning) to explicitly model human sequential decision making in depth editing operations. We formulate the depth adjustment process as a Markov decision process where actions are defined as camera movement operations to control the distance between the left and right cameras. Our agent is trained based on the guidance of an objective visual comfort assessment metric to learn the optimal sequence of camera movement actions in terms of perceptual aspects in stereoscopic viewing. With extensive experiments and user studies, we show the effectiveness of our VCA-RL model on three different S3D databases.

v2026.09.13