Arrow Research search

Author name cluster

Sunghyun Park

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
2 author rows

Possible papers

9

NeurIPS Conference 2025 Conference Paper

MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans

  • Shubhankar Borse
  • Seokeon Choi
  • Sunghyun Park
  • Jeongho Kim
  • Shreya Kadambi
  • Risheek Garrepalli
  • Sungrack Yun
  • Durga Malladi

Generation of images containing multiple humans, performing complex actions, while preserving their facial identities, is a significant challenge. A major factor contributing to this is the lack of a a dedicated benchmark. To address this, we introduce MultiHuman-Testbench, a novel benchmark for rigorously evaluating generative models for multi-human generation. The benchmark comprises 1800 samples, including carefully curated text prompts, describing a range of simple to complex human actions. These prompts are matched with a total of 5, 550 unique human face images, sampled uniformly to ensure diversity across age, ethnic background, and gender. Alongside captions, we provide human-selected pose conditioning images which accurately match the prompt. We propose a multi-faceted evaluation suite employing four key metrics to quantify face count, ID similarity, prompt alignment, and action detection. We conduct a thorough evaluation of a diverse set of models, including zero-shot approaches and training-based methods, with and without regional priors. We also propose novel techniques to incorporate image and region isolation using human segmentation and Hungarian matching, significantly improving ID similarity. Our proposed benchmark and key findings provide valuable insights and a standardized tool for advancing research in multi-human image generation.

AAAI Conference 2025 Conference Paper

What to Preserve and What to Transfer: Faithful, Identity-Preserving Diffusion-based Hairstyle Transfer

  • Chaeyeon Chung
  • Sunghyun Park
  • Jeongho Kim
  • Jaegul Choo

Hairstyle transfer is a challenging task in the image editing field that modifies the hairstyle of a given face image while preserving its other appearance and background features. The existing hairstyle transfer approaches heavily rely on StyleGAN, which is pre-trained on cropped and aligned face images. Hence, they struggle to generalize under challenging conditions such as extreme variations of head poses or focal lengths. To address this issue, we propose a one-stage hairstyle transfer diffusion model, HairFusion, that applies to real-world scenarios. Specifically, we carefully design a hair-agnostic representation as the input of the model, where the original hair information is thoroughly eliminated. Next, we introduce a hair align cross-attention (Align-CA) to accurately align the reference hairstyle with the face image while considering the difference in their head poses. To enhance the preservation of the face image’s original features, we leverage adaptive hair blending during the inference, where the output’s hair regions are estimated by the cross-attention map in Align-CA and blended with non-hair areas of the face image. Our experimental results show that our method achieves state-of-the-art performance compared to the existing methods in preserving the integrity of both the transferred hairstyle and the surrounding features.

YNICL Journal 2024 Journal Article

Resting-state functional connectivity of amygdala subregions across different symptom subtypes of obsessive–compulsive disorder patients

  • Harah Kwon
  • Minji Ha
  • Sunah Choi
  • Sunghyun Park
  • Moonyoung Jang
  • Minah Kim
  • Jun Soo Kwon

AIM: Obsessive-compulsive disorder (OCD) is a heterogeneous condition characterized by distinct symptom subtypes, each with varying pathophysiologies and treatment responses. Recent research has highlighted the role of the amygdala, a brain region that is central to emotion processing, in these variations. However, the role of amygdala subregions with distinct functions has not yet been fully elucidated. In this study, we aimed to clarify the biological mechanisms underlying OCD subtype heterogeneity by investigating the functional connectivity (FC) of amygdala subregions across distinct OCD symptom subtypes. METHODS: Resting-state functional magnetic resonance images were obtained from 107 medication-free OCD patients and 110 healthy controls (HCs). Using centromedial, basolateral, and superficial subregions of the bilateral amygdala as seed regions, whole-brain FC was compared between OCD patients and HCs and among patients with different OCD symptom subtypes, which included contamination fear and washing, obsessive (i.e., harm due to injury, aggression, sexual, and religious), and compulsive (i.e., symmetry, ordering, counting, and checking) subtypes. RESULTS: Compared to HCs, compulsive-type OCD patients exhibited hypoconnectivity between the left centromedial amygdala (CMA) and bilateral superior frontal gyri. Compared with patients with contamination fear and washing OCD subtypes, patients with compulsive-type OCD showed hypoconnectivity between the left CMA and left frontal cortex. CONCLUSIONS: CMA-frontal cortex hypoconnectivity may contribute to the compulsive presentation of OCD through impaired control of behavioral responses to negative emotions. Our findings underscored the potential significance of the distinct neural underpinnings of different OCD manifestations, which could pave the way for more targeted treatment strategies in the future.

AAAI Conference 2024 Conference Paper

When Model Meets New Normals: Test-Time Adaptation for Unsupervised Time-Series Anomaly Detection

  • Dongmin Kim
  • Sunghyun Park
  • Jaegul Choo

Time-series anomaly detection deals with the problem of detecting anomalous timesteps by learning normality from the sequence of observations. However, the concept of normality evolves over time, leading to a "new normal problem", where the distribution of normality can be changed due to the distribution shifts between training and test data. This paper highlights the prevalence of the new normal problem in unsupervised time-series anomaly detection studies. To tackle this issue, we propose a simple yet effective test-time adaptation strategy based on trend estimation and a self-supervised approach to learning new normalities during inference. Extensive experiments on real-world benchmarks demonstrate that incorporating the proposed strategy into the anomaly detector consistently improves the model's performances compared to the existing baselines, leading to robustness to the distribution shifts.

AAAI Conference 2024 Conference Paper

YTCommentQA: Video Question Answerability in Instructional Videos

  • Saelyne Yang
  • Sunghyun Park
  • Yunseok Jang
  • Moontae Lee

Instructional videos provide detailed how-to guides for various tasks, with viewers often posing questions regarding the content. Addressing these questions is vital for comprehending the content, yet receiving immediate answers is difficult. While numerous computational models have been developed for Video Question Answering (Video QA) tasks, they are primarily trained on questions generated based on video content, aiming to produce answers from within the content. However, in real-world situations, users may pose questions that go beyond the video's informational boundaries, highlighting the necessity to determine if a video can provide the answer. Discerning whether a question can be answered by video content is challenging due to the multi-modal nature of videos, where visual and verbal information are intertwined. To bridge this gap, we present the YTCommentQA dataset, which contains naturally-generated questions from YouTube, categorized by their answerability and required modality to answer -- visual, script, or both. Experiments with answerability classification tasks demonstrate the complexity of YTCommentQA and emphasize the need to comprehend the combined role of visual and script information in video reasoning. The dataset is available at https://github.com/lgresearch/YTCommentQA.

ICRA Conference 2022 Conference Paper

Semi-Autonomous Teleoperation via Learning Non-Prehensile Manipulation Skills

  • Sangbeom Park
  • Yoonbyung Chai
  • Sunghyun Park
  • Jeongeun Park 0002
  • Kyungjae Lee 0001
  • Sungjoon Choi

In this paper, we present a semi-autonomous teleoperation framework for a pick-and-place task using an RGB-D sensor. In particular, we assume that the target object is located in a cluttered environment where both prehensile grasping and non-prehensile manipulation are combined for efficient teleoperation. A trajectory-based reinforcement learning is utilized for learning the non-prehensile manipulation to rearrange the objects for enabling direct grasping. From the depth image of the cluttered environment and the location of the goal object, the learned policy can provide multiple options of non-prehensile manipulation to the human operator. We carefully design a reward function for the rearranging task where the policy is trained in a simulational environment. Then, the trained policy is transferred to a real-world and evaluated in a number of real-world experiments with the varying number of objects where we show that the proposed method outperforms manual keyboard control in terms of the time duration for the grasping.

AAAI Conference 2021 Conference Paper

Vid-ODE: Continuous-Time Video Generation with Neural Ordinary Differential Equation

  • Sunghyun Park
  • Kangyeol Kim
  • Junsoo Lee
  • Jaegul Choo
  • Joonseok Lee
  • Sookyung Kim
  • Edward Choi

Video generation models often operate under the assumption of fixed frame rates, which leads to suboptimal performance when it comes to handling flexible frame rates (e. g. , increasing the frame rate of the more dynamic portion of the video as well as handling missing video frames). To resolve the restricted nature of existing video generation models’ ability to handle arbitrary timesteps, we propose continuous-time video generation by combining neural ODE (Vid-ODE) with pixellevel video processing techniques. Using ODE-ConvGRU as an encoder, a convolutional version of the recently proposed neural ODE, which enables us to learn continuous-time dynamics, Vid-ODE can learn the spatio-temporal dynamics of input videos of flexible frame rates. The decoder integrates the learned dynamics function to synthesize video frames at any given timesteps, where the pixel-level composition technique is used to maintain the sharpness of individual frames. With extensive experiments on four real-world video datasets, we verify that the proposed Vid-ODE outperforms state-ofthe-art approaches under various video generation settings, both within the trained time range (interpolation) and beyond the range (extrapolation). To the best of our knowledge, Vid- ODE is the first work successfully performing continuous-time video generation using real-world videos.

ICML Conference 2019 Conference Paper

Learning Context-dependent Label Permutations for Multi-label Classification

  • Jinseok Nam
  • Young-Bum Kim
  • Eneldo Loza Mencía
  • Sunghyun Park
  • Ruhi Sarikaya
  • Johannes Fürnkranz

A key problem in multi-label classification is to utilize dependencies among the labels. Chaining classifiers are a simple technique for addressing this problem but current algorithms all assume a fixed, static label ordering. In this work, we propose a multi-label classification approach which allows to choose a dynamic, context-dependent label ordering. Our proposed approach consists of two sub-components: a simple EM-like algorithm which bootstraps the learned model, and a more elaborate approach based on reinforcement learning. Our experiments on three public multi-label classification benchmarks show that our proposed dynamic label ordering approach based on reinforcement learning outperforms recurrent neural networks with fixed label ordering across both bipartition and ranking measures on all the three datasets. As a result, we obtain a powerful sequence prediction-based algorithm for multi-label classification, which is able to efficiently and explicitly exploit label dependencies.

AAAI Conference 2019 Conference Paper

Paraphrase Diversification Using Counterfactual Debiasing

  • Sunghyun Park
  • Seung-won Hwang
  • Fuxiang Chen
  • Jaegul Choo
  • Jung-Woo Ha
  • Sunghun Kim
  • Jinyeong Yim

The problem of generating a set of diverse paraphrase sentences while (1) not compromising the original meaning of the original sentence, and (2) imposing diversity in various semantic aspects, such as a lexical or syntactic structure, is examined. Existing work on paraphrase generation has focused more on the former, and the latter was trained as a fixed style transfer, such as transferring from positive to negative sentiments, even at the cost of losing semantics. In this work, we consider style transfer as a means of imposing diversity, with a paraphrasing correctness constraint that the target sentence must remain a paraphrase of the original sentence. However, our goal is to maximize the diversity for a set of k generated paraphrases, denoted as the diversified paraphrase (DP) problem. Our key contribution is deciding the style guidance at generation towards the direction of increasing the diversity of output with respect to those generated previously. As pre-materializing training data for all style decisions is impractical, we train with biased data, but with debiasing guidance. Compared to state-of-the-art methods, our proposed model can generate more diverse and yet semantically consistent paraphrase sentences. That is, our model, trained with the MSCOCO dataset, achieves the highest embedding scores, .94/. 95/. 86, similar to state-of-the-art results, but with a lower mBLEU score (more diverse) by 8. 73%.

v2026.09.13