Arrow Research search

Author name cluster

Minkyu Kim

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

18 papers
2 author rows

Possible papers

18

AAAI Conference 2026 Conference Paper

Transparent Networks for Multivariate Time Series

  • Minkyu Kim
  • Suan Lee
  • Jinho Kim

Transparent models, which provide inherently interpretable predictions, are receiving significant attention in high-stakes domains. However, despite much real-world data being collected as time series, there is a lack of studies on transparent time series models. To address this gap, we propose a novel transparent neural network model for time series called Generalized Additive Time Series Model (GATSM). GATSM consists of two parts: 1) independent feature networks to learn feature representations, and 2) a transparent temporal module to learn temporal patterns across different time steps using the feature representations. This structure allows GATSM to effectively capture temporal patterns and handle varying-length time series while preserving transparency. Empirical experiments show that GATSM significantly outperforms existing generalized additive models and achieves comparable performance to black-box time series models, such as recurrent neural networks and Transformer. In addition, we demonstrate that GATSM finds interesting patterns in time series.

NeurIPS Conference 2025 Conference Paper

Energy-based generator matching: A neural sampler for general state space

  • Dongyeop Woo
  • Minsu Kim
  • Minkyu Kim
  • Kiyoung Seong
  • Sungsoo Ahn

We propose Energy-based generator matching (EGM), a modality-agnostic approach to train generative models from energy functions in the absence of data. Extending the recently proposed generator matching, EGM enables training of arbitrary continuous-time Markov processes, e. g. , diffusion, flow, and jump, and can generate data from continuous, discrete, and a mixture of two modalities. To this end, we propose estimating the generator matching loss using self-normalized importance sampling with an additional bootstrapping trick to reduce variance in the importance weight. We validate EGM on both discrete and multimodal tasks up to 100 and 20 dimensions, respectively.

NeurIPS Conference 2025 Conference Paper

On scalable and efficient training of diffusion samplers

  • Minkyu Kim
  • Kiyoung Seong
  • Dongyeop Woo
  • Sungsoo Ahn
  • Minsu Kim

We address the challenge of training diffusion models to sample from unnormalized energy distributions in the absence of data, the so-called diffusion samplers. Although these approaches have shown promise, they struggle to scale in more demanding scenarios where energy evaluations are expensive and the sampling space is high-dimensional. To address this limitation, we propose a scalable and sample-efficient framework that properly harmonizes the powerful classical sampling method and the diffusion sampler. Specifically, we utilize Monte Carlo Markov chain (MCMC) samplers with a novelty-based auxiliary energy as a Searcher to collect off-policy samples, using an auxiliary energy function to compensate for exploring modes the diffusion sampler rarely visits. These off-policy samples are then combined with on-policy data to train the diffusion sampler, thereby expanding its coverage of the energy landscape. Furthermore, we identify primacy bias, i. e. , the preference of samplers for early experience during training, as the main cause of mode collapse during training, and introduce a periodic re-initialization trick to resolve this issue. Our method significantly improves sample efficiency on standard benchmarks for diffusion samplers and also excels at higher-dimensional problems and real-world molecular conformer generation.

ICLR Conference 2025 Conference Paper

Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance

  • Dongmin Park
  • Sebin Kim
  • Taehong Moon
  • Minkyu Kim
  • Kangwook Lee 0001
  • Jaewoong Cho

State-of-the-art text-to-image (T2I) diffusion models often struggle to generate rare compositions of concepts, e.g., objects with unusual attributes. In this paper, we show that the compositional generation power of diffusion models on such rare concepts can be significantly enhanced by the Large Language Model (LLM) guidance. We start with empirical and theoretical analysis, demonstrating that exposing frequent concepts relevant to the target rare concepts during the diffusion sampling process yields more accurate concept composition. Based on this, we propose a training-free approach, R2F, that plans and executes the overall rare-to-frequent concept guidance throughout the diffusion inference by leveraging the abundant semantic knowledge in LLMs. Our framework is flexible across any pre-trained diffusion models and LLMs, and can be seamlessly integrated with the region-guided diffusion approaches. Extensive experiments on three datasets, including our newly proposed benchmark, RareBench, containing various prompts with rare compositions of concepts, R2F significantly surpasses existing models including SD3.0 and FLUX by up to 28.1%p in T2I alignment. Code is available at https://github.com/krafton-ai/Rare-to-Frequent.

TMLR Journal 2025 Journal Article

SuFP: Piecewise Bit Allocation Floating-Point for Robust Neural Network Quantization

  • Geonwoo Ko
  • Sungyeob Yoo
  • Seri Ham
  • Seeyeon Kim
  • Minkyu Kim
  • Joo-Young Kim

The rapid growth in model size and computational demand of Deep Neural Networks (DNNs) has led to significant challenges in memory and compute efficiency, necessitating the adoption of lower bit-width data types to enhance hardware performance. Floating-point 8 (FP8) has emerged as a promising solution, supported by the latest AI processors, due to its potential for reducing memory usage and computational load. However, each application often requires its own optimal FP8 configuration to achieve high performance, resulting in inconsistent performance and increased hardware complexity. To address these limitations, we introduce Super Floating-Point (SuFP), an innovative data type that integrates various floating-point configurations into a single representation through a piecewise bit allocation. This approach enables SuFP to effectively capture both dense regions near zero and sparse regions with outliers, thereby minimizing quantization errors and ensuring full-precision floating-point performance across different models. Furthermore, SuFP’s processing element design is optimized to reduce the hardware overhead. Our experimental results demonstrate the robustness and accuracy of SuFP over various neural networks in the vision and natural language processing domains. Remarkably, SuFP shows its superiority in large models such as large language model (Llama 2) and text-to-image generative model (Stable Diffusion v2). We also verify training feasibility on ResNet models and highlight the structural design of SuFP for general applicability.

ICLR Conference 2025 Conference Paper

Test-time Alignment of Diffusion Models without Reward Over-optimization

  • Sunwoo Kim 0007
  • Minkyu Kim
  • Dongmin Park

Diffusion models excel in generative tasks, but aligning them with specific objectives while maintaining their versatility remains challenging. Existing fine-tuning methods often suffer from reward over-optimization, while approximate guidance approaches fail to optimize target rewards effectively. Addressing these limitations, we propose a training-free, test-time method based on Sequential Monte Carlo (SMC) to sample from the reward-aligned target distribution. Our approach, tailored for diffusion sampling and incorporating tempering techniques, achieves comparable or superior target rewards to fine-tuning methods while preserving diversity and cross-reward generalization. We demonstrate its effectiveness in single-reward optimization, multi-objective scenarios, and online black-box optimization. This work offers a robust solution for aligning diffusion models with diverse downstream objectives without compromising their general capabilities. Code is available at https://github.com/krafton-ai/DAS.

EAAI Journal 2025 Journal Article

Thermoeconomic optimization of climate-adaptive solar and wind multi-generation systems using artificial intelligence and thermal energy recovery

  • Ehsanolah Assareh
  • Nima Izadyar
  • Emad Tandis
  • Mehdi Khiadani
  • Amir shahavand
  • Neha Agarwal
  • Arian Gerami
  • Ahmed Rezk

This study presents a hybrid multi-generation energy system designed to overcome solar intermittency while meeting the global demand for integrated delivery of electricity, water, cooling, and sustainable fuels in the transition to decarbonization. The engineering application integrates solar thermal and wind energy with a modified Brayton cycle, a Steam Rankine Cycle (SRC), and a Thermoelectric Generator (TEG) to simultaneously produce electricity, fresh water via Reverse Osmosis (RO), hydrogen and oxygen via Proton Exchange Membrane Electrolyzer (PEME), and cooling (via absorption chiller) within a unified optimization framework. The system was modeled using Engineering Equation Solver (EES) and optimized via Response Surface Methodology (RSM) based on 11 decision variables. To address the complexity of optimization, a second phase applied Artificial Intelligence (AI) techniques: Adaptive Boosting (AdaBoost) for predictive modelling and Particle Swarm Optimization (PSO) for global optimization. Under optimal conditions, the Response Surface Methodology yielded an exergy efficiency of 45. 8 % with a cost rate of 576. 76 United States Dollars per hour (USD/h), while AI reduced costs to 211. 2 USD/h with a moderate efficiency trade-off. Simulation of the optimized configuration across eight diverse climates identified Quebec as most viable, generating 22, 629. 6 Megawatt-hours per year (MWh/year) of electricity and avoiding 4616. 4 tons of Carbon Dioxide (CO2) emissions annually. Integration of wind energy stabilizes solar variability, enhancing performance. AI contributes to optimizing complex interactions, nonlinear constraints, and multiple conflicting objectives. The methodology offers a scalable, generalizable framework for designing intelligent, climate-resilient infrastructures. Future research includes AI-enabled real-time control, experimental validation, and broader deployment strategies.

ICLR Conference 2024 Conference Paper

Image Clustering Conditioned on Text Criteria

  • Sehyun Kwon
  • Jaeseung Park
  • Minkyu Kim
  • Jaewoong Cho
  • Ernest K. Ryu
  • Kangwook Lee 0001

Classical clustering methods do not provide users with direct control of the clustering results, and the clustering results may not be consistent with the relevant criterion that a user has in mind. In this work, we present a new methodology for performing image clustering based on user-specified criteria in the form of text by leveraging modern Vision-Language Models and Large Language Models. We call our method Image Clustering Conditioned on Text Criteria (IC$|$TC), and it represents a different paradigm of image clustering. IC$|$TC requires a minimal and practical degree of human intervention and grants the user significant control over the clustering results in return. Our experiments show that IC$|$TC can effectively cluster images with various criteria, such as human action, physical location, or the person's mood, significantly outperforming baselines.

AAAI Conference 2023 Conference Paper

Balanced Column-Wise Block Pruning for Maximizing GPU Parallelism

  • Cheonjun Park
  • Mincheol Park
  • Hyun Jae Oh
  • Minkyu Kim
  • Myung Kuk Yoon
  • Suhyun Kim
  • Won Woo Ro

Pruning has been an effective solution to reduce the number of computations and the memory requirement in deep learning. The pruning unit plays an important role in exploiting the GPU resources efficiently. The filter is proposed as a simple pruning unit of structured pruning. However, since the filter is quite large as pruning unit, the accuracy drop is considerable with a high pruning ratio. GPU rearranges the weight and input tensors into tiles (blocks) for efficient computation. To fully utilize GPU resources, this tile structure should be considered, which is the goal of block pruning. However, previous block pruning prunes both row vectors and column vectors. Pruning of row vectors in a tile corresponds to filter pruning, and it also interferes with column-wise block pruning of the following layer. In contrast, column vectors are much smaller than row vectors and can achieve lower accuracy drop. Additionally, if the pruning ratio for each tile is different, GPU utilization can be limited by imbalanced workloads by irregular-sized blocks. The same pruning ratio for the weight tiles processed in parallel enables the actual inference process to fully utilize the resources without idle time. This paper proposes balanced column-wise block pruning, named BCBP, to satisfy two conditions: the column-wise minimal size of the pruning unit and balanced workloads. We demonstrate that BCBP is superior to previous pruning methods through comprehensive experiments.

YNIMG Journal 2023 Journal Article

Cortical maps of somatosensory perception in human

  • Seokyun Ryun
  • Minkyu Kim
  • June Sic Kim
  • Chun Kee Chung

Tactile and movement-related somatosensory perceptions are crucial for our daily lives and survival. Although the primary somatosensory cortex is thought to be the key structure of somatosensory perception, various cortical downstream areas are also involved in somatosensory perceptual processing. However, little is known about whether cortical networks of these downstream areas can be dissociated depending on each perception, especially in human. We address this issue by combining data from direct cortical stimulation (DCS) for eliciting somatosensation and data from high-gamma band (HG) elicited during tactile stimulation and movement tasks. We found that artificial somatosensory perception is elicited not only from conventional somatosensory-related areas such as the primary and secondary somatosensory cortices but also from a widespread network including superior/inferior parietal lobules and premotor cortex. Interestingly, DCS on the dorsal part of the fronto-parietal area including superior parietal lobule and dorsal premotor cortex often induces movement-related somatosensations, whereas that on the ventral one including inferior parietal lobule and ventral premotor cortex generally elicits tactile sensations. Furthermore, the HG mapping results of the movement and passive tactile stimulation tasks revealed considerable similarity in the spatial distribution between the HG and DCS functional maps. Our findings showed that macroscopic neural processing for tactile and movement-related perceptions could be segregated.

NeurIPS Conference 2023 Conference Paper

S-CLIP: Semi-supervised Vision-Language Learning using Few Specialist Captions

  • Sangwoo Mo
  • Minkyu Kim
  • Kyungmin Lee
  • Jinwoo Shin

Vision-language models, such as contrastive language-image pre-training (CLIP), have demonstrated impressive results in natural image domains. However, these models often struggle when applied to specialized domains like remote sensing, and adapting to such domains is challenging due to the limited number of image-text pairs available for training. To address this, we propose S-CLIP, a semi-supervised learning method for training CLIP that utilizes additional unpaired images. S-CLIP employs two pseudo-labeling strategies specifically designed for contrastive learning and the language modality. The caption-level pseudo-label is given by a combination of captions of paired images, obtained by solving an optimal transport problem between unpaired and paired images. The keyword-level pseudo-label is given by a keyword in the caption of the nearest paired image, trained through partial label learning that assumes a candidate set of labels for supervision instead of the exact one. By combining these objectives, S-CLIP significantly enhances the training of CLIP using only a few image-text pairs, as demonstrated in various specialist domains, including remote sensing, fashion, scientific figures, and comics. For instance, S-CLIP improves CLIP by 10% for zero-shot classification and 4% for image-text retrieval on the remote sensing benchmark, matching the performance of supervised CLIP while using three times fewer image-text pairs.

NeurIPS Conference 2021 Conference Paper

SmoothMix: Training Confidence-calibrated Smoothed Classifiers for Certified Robustness

  • Jongheon Jeong
  • Sejun Park
  • Minkyu Kim
  • Heung-Chang Lee
  • Do-Guk Kim
  • Jinwoo Shin

Randomized smoothing is currently a state-of-the-art method to construct a certifiably robust classifier from neural networks against $\ell_2$-adversarial perturbations. Under the paradigm, the robustness of a classifier is aligned with the prediction confidence, i. e. , the higher confidence from a smoothed classifier implies the better robustness. This motivates us to rethink the fundamental trade-off between accuracy and robustness in terms of calibrating confidences of a smoothed classifier. In this paper, we propose a simple training scheme, coined SmoothMix, to control the robustness of smoothed classifiers via self-mixup: it trains on convex combinations of samples along the direction of adversarial perturbation for each input. The proposed procedure effectively identifies over-confident, near off-class samples as a cause of limited robustness in case of smoothed classifiers, and offers an intuitive way to adaptively set a new decision boundary between these samples for better robustness. Our experimental results demonstrate that the proposed method can significantly improve the certified $\ell_2$-robustness of smoothed classifiers compared to existing state-of-the-art robust training methods.

IROS Conference 2019 Conference Paper

Toward Achieving Formal Guarantees for Human-Aware Controllers in Human-Robot Interactions

  • Rachel Schlossman
  • Minkyu Kim
  • Ufuk Topcu
  • Luis Sentis

With the primary objective of human-robot interaction being to support humans’ goals, there exists a need to formally synthesize robot controllers that can provide the desired service. Synthesis techniques have the benefit of providing formal guarantees for specification satisfaction. There is potential to apply these techniques for devising robot controllers whose specifications are coupled with human needs. This paper explores the use of formal methods to construct human-aware robot controllers to support the productivity requirements of humans. We tackle these types of scenarios via human workload-informed models and reactive synthesis. This strategy allows us to synthesize controllers that fulfill formal specifications that are expressed as linear temporal logic formulas. We present a case study in which we reason about a work delivery and pickup task such that the robot increases worker productivity, but not stress induced by high work backlog. We demonstrate our controller using the Toyota HSR, a mobile manipulator robot. The results demonstrate the realization of a robust robot controller that is guaranteed to properly reason and react in collaborative tasks with human partners.

JBHI Journal 2016 Journal Article

Reconstruction of Precordial Lead Electrocardiogram From Limb Leads Using the State-Space Model

  • Jaehyeok Lee
  • Minkyu Kim
  • Jungkuk Kim

A new electrocardiogram (ECG) reconstruction method based on a state-space model is presented. This method was applied to reconstruct precordial leads from limb leads (lead I, II, III) for its validity verification. The system matrices of the state-space model were estimated at the model estimation stage by considering the limb lead signals as the input of the system and precordial lead signals as the output. To evaluate the performance of the proposed method, all of the 549 records of the Physikalisch Technische Bundesanstalt diagnostic ECG database were used, and the correlation coefficients (CC) and root-mean-square errors between reconstructed ECG and measured ECG were calculated. For a more objective evaluation, the results were compared with those of linear regression model that has been typically used for ECG reconstruction. The mean and median values of CCs were higher than 0. 988 and 0. 995, respectively, for healthy subject data, and also higher than 0. 981 and 0. 993, respectively, for cardiac patient data and comparable to those by linear regression model. In addition, it was found that the reconstruction performance depended on the type of disease rather than lead type. Among cardiac patient data, hypertrophy, myocarditis, valvular heart disease, and stable heart angina showed higher CC (>0. 990), while unstable angina and heart failure showed lower CC of 0. 932 and 0. 914, respectively. Moreover, when ECG contaminated with the noise was used for reconstruction, the proposed method demonstrated better performance than linear regression model in general.

ICRA Conference 2016 Conference Paper

Tele-operation system with reliable grasping force estimation to compensate for the time-varying sEMG feature

  • Minkyu Kim
  • Jaemin Lee
  • Keehoon Kim

This paper presents a real-time framework for tele-manipulation by using sEMG signals to estimate both human motion and force intention. Our previous study showed that the ability to detect discrete force levels was not applicable to complex tasks such as grasping, holding, and manipulating various objects with variable force. Consequently, we identified the need to simultaneously track the arm and hand configurations and estimate the grasping force. However, it is difficult to continuously estimate the grasping force because of the time-varying nature of surface Electromyogram (sEMG) signals, even if a force remains constant. To solve such a problem, this study proposes a new regression strategy to enable continuous and proportional measurements and transmission of the grasping force by using sEMG signals in transient and steady-states. A 7-DOF robot arm with a robotic hand was able to remotely imitate a subject via an easily-wearable sEMG and inertia measurement units sensor interface. The experimental results verified that the motion and force capturing system successfully enabled interaction tasks, such as grasping, holding, and releasing motions with objects, with reliable and continuous force estimation.

IROS Conference 2015 Conference Paper

A robust control method of multi-DOF power-assistant robots for unknown external perturbation using sEMG signals

  • Jaemin Lee
  • Minkyu Kim
  • Keehoon Kim

This paper presents a control method of multi-DOF power assistant robots for anatomical multi-axis joints such as the wrist and the ankle. It is difficult to calculate the accurate direction of human motion intention during manipulating an object due to discrepancy between the calculated force from F/T sensor and the real human intention. Only using an sEMG is not an adequate method of power assistance for unknown external perturbation in the anatomical multi-axis joint, because the sEMG signal cannot figure out where the intention vector exists during interactions. This paper proposes a robust control method of power-assistant robots for unknown external perturbation during manipulating an object by using both the F/T sensor and sEMG. The specific purpose of this study to control the exoskeleton robot for the wrist motion during manipulating an object, although the accurate intention vector of the wrist joint is unknown. It was verified that the proposed method generates the assisted power to follow the human motion intention even in the case of unknown external forces through experiments.

ICRA Conference 2014 Conference Paper

Implementation of real-time motion and force capturing system for tele-manipulation based on sEMG signals and IMU motion data

  • Minkyu Kim
  • Kwanghyun Ryu
  • Yonghwan Oh
  • Sang-Rok Oh
  • Keehoon Kim

In this paper, we present a real-time motion and force capturing system for tele-operated robotic manipulation that combines surface-electromyogram (sEMG) pattern recognition with an inertia measurement unit(IMU) for motion calculation. The purpose of this system is to deliver the human motion and intended force to a remote robotic manipulator and to realize multi-fingered activities-of-daily-living (ADL) tasks that require motion and force commands simultaneously and instantaneously. The proposed system combines two different sensors: (i) the IMU captures arm motion, (ii) and the sEMG detects the hand motion and force. We propose an algorithm to calculate the human arm motion using IMU sensors and a pattern recognition algorithm for a multi-grasp myoelectric control method that uses sEMG signals to determine the hand postures and grasping force information. In order to validate the proposed motion and force capturing system, we used the in-house developed robotic arm, K-Arm, which has seven degrees-of-freedom (three for shoulder, one for elbow, and three for wrist), and a sixteen degrees-of-freedom robotic hand. Transmission Control Protocol Internet Protocol (TCP/IP)-based network communication was implemented for total system integration. The experimental results verified the effectiveness of the proposed method, although some open problems encountered.

IROS Conference 2014 Conference Paper

Integrated control method for power-assisted rehabilitation: Ellipsoid regression and impedance control

  • Jaemin Lee
  • Minkyu Kim
  • Sang-Rok Oh
  • Keehoon Kim

This paper proposes an integrated control method including learning with an ellipsoid function and impedance controller for rehabilitation using power-assisted robotic devices. The proposed controller consists of two parts of a primary algorithm, which are ellipsoid regression method for re-designing trajectory and impedance controller with pseudo mass/inertia. The ellipsoid regression method generates reference impedance profiles though acquiring motion and force trajectories during rehabilitation tasks assisted by therapists. The assisted force is controlled by impedance controller during execution of rehabilitation task using a concept of pseudo mass/inertia. The proposed method offers the power-assisted rehabilitation as guided by therapist, without consistent help from the therapist or other assisters. The proposed control method is validated by experiments throughout a 2-DOF rehabilitation robot, KULEX-2DOF(KIST Upper Limb Exoskeleton - 2DOF).

v2026.09.13