Arrow Research search

Author name cluster

Sungwon Kim

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
2 author rows

Possible papers

8

NeurIPS Conference 2025 Conference Paper

Disentangling Hyperedges through the Lens of Category Theory

  • Yoonho Lee
  • Junseok Lee
  • Sangwoo Seo
  • Sungwon Kim
  • Yeongmin Kim
  • Chanyoung Park

Despite the promising results of disentangled representation learning in discovering latent patterns in graph-structured data, few studies have explored disentanglement for hypergraph-structured data. Integrating hyperedge disentanglement into hypergraph neural networks enables models to leverage hidden hyperedge semantics, such as unannotated relations between nodes, that are associated with labels. This paper presents an analysis of hyperedge disentanglement from a category-theoretical perspective and proposes a novel criterion for disentanglement derived from the naturality condition. Our proof-of-concept model experimentally showed the potential of the proposed criterion by successfully capturing functional relations of genes (nodes) in genetic pathways (hyperedges).

ICML Conference 2025 Conference Paper

Thickness-aware E(3)-Equivariant 3D Mesh Neural Networks

  • Sungwon Kim
  • Namkyeong Lee
  • Yunyoung Doh
  • Seungmin Shin
  • Guimok Cho
  • Seung-Won Jeon
  • Sangkook Kim
  • Chanyoung Park

Mesh-based 3D static analysis methods have recently emerged as efficient alternatives to traditional computational numerical solvers, significantly reducing computational costs and runtime for various physics-based analyses. However, these methods primarily focus on surface topology and geometry, often overlooking the inherent thickness of real-world 3D objects, which exhibits high correlations and similar behavior between opposing surfaces. This limitation arises from the disconnected nature of these surfaces and the absence of internal edge connections within the mesh. In this work, we propose a novel framework, the Thickness-aware E(3)-Equivariant 3D Mesh Neural Network (T-EMNN), that effectively integrates the thickness of 3D objects while maintaining the computational efficiency of surface meshes. Additionally, we introduce data-driven coordinates that encode spatial information while preserving E(3)-equivariance or invariance properties, ensuring consistent and robust analysis. Evaluations on a real-world industrial dataset demonstrate the superior performance of T-EMNN in accurately predicting node-level 3D deformations, effectively capturing thickness effects while maintaining computational efficiency.

NeurIPS Conference 2023 Conference Paper

Density of States Prediction of Crystalline Materials via Prompt-guided Multi-Modal Transformer

  • Namkyeong Lee
  • Heewoong Noh
  • Sungwon Kim
  • Dongmin Hyun
  • Gyoung S. Na
  • Chanyoung Park

The density of states (DOS) is a spectral property of crystalline materials, which provides fundamental insights into various characteristics of the materials. While previous works mainly focus on obtaining high-quality representations of crystalline materials for DOS prediction, we focus on predicting the DOS from the obtained representations by reflecting the nature of DOS: DOS determines the general distribution of states as a function of energy. That is, DOS is not solely determined by the crystalline material but also by the energy levels, which has been neglected in previous works. In this paper, we propose to integrate heterogeneous information obtained from the crystalline materials and the energies via a multi-modal transformer, thereby modeling the complex relationships between the atoms in the crystalline materials and various energy levels for DOS prediction. Moreover, we propose to utilize prompts to guide the model to learn the crystal structural system-specific interactions between crystalline materials and energies. Extensive experiments on two types of DOS, i. e. , Phonon DOS and Electron DOS, with various real-world scenarios demonstrate the superiority of DOSTransformer. The source code for DOSTransformer is available at https: //github. com/HeewoongNoh/DOSTransformer.

EAAI Journal 2023 Journal Article

Improving the accuracy of daily solar radiation prediction by climatic data using an efficient hybrid deep learning model: Long short-term memory (LSTM) network coupled with wavelet transform

  • Meysam Alizamir
  • Jalal Shiri
  • Ahmad Fakheri Fard
  • Sungwon Kim
  • AliReza Docheshmeh Gorgij
  • Salim Heddam
  • Vijay P. Singh

Accurate daily solar radiation prediction is a crucial task for the management and generation of solar energy as one of the alternatives to fossil fuels. In this study, the prediction accuracy of new machine learning methods, wavelet long short-term memory (WLSTM), wavelet multi-layer perceptron artificial neural network (WMLPANN), long short-term memory (LSTM), multi-layer perceptron artificial neural network (MLPANN), and multivariate adaptive regression splines (MARS), was assessed for modeling daily solar radiation using various input combinations of climatic data of maximum and minimum relative humidity, potential evapotranspiration, maximum and minimum temperature, precipitation and wind speed from two stations, Brownstown and Carbondale located in Illinois, USA. For accurate assessment of prediction accuracy of the proposed models, four reliable statistical metrics, root mean squared error (RMSE), Nash–Sutcliffe efficiency coefficient (NSE), coefficient of determination (R2), and One-Tailed Wilcoxon Signed-Rank Test were employed. Comparison of results, based on the RMSE values, indicated that the WLSTM method performed better than the WMLPANN, LSTM, MLPANN and MARS methods in the estimation of solar radiation values at both stations. The average RMSE values of WMLPANN, LSTM, MLPANN, and MARS approaches was decreased by 6%, 4%, 7. 3%, and 13. 5% using WLSTM method at Brownstown Station, by 6%, 5. 2%, 13. 2%, and 14. 3% at Carbondale Station, respectively. The overall results during the testing phase of both stations revealed the successful application of hybridization of the LSTM model with the wavelet transform technique for improving the prediction accuracy of solar radiation based on climatic parameters.

NeurIPS Conference 2023 Conference Paper

Interpretable Prototype-based Graph Information Bottleneck

  • Sangwoo Seo
  • Sungwon Kim
  • Chanyoung Park

The success of Graph Neural Networks (GNNs) has led to a need for understanding their decision-making process and providing explanations for their predictions, which has given rise to explainable AI (XAI) that offers transparent explanations for black-box models. Recently, the use of prototypes has successfully improved the explainability of models by learning prototypes to imply training graphs that affect the prediction. However, these approaches tend to provide prototypes with excessive information from the entire graph, leading to the exclusion of key substructures or the inclusion of irrelevant substructures, which can limit both the interpretability and the performance of the model in downstream tasks. In this work, we propose a novel framework of explainable GNNs, called interpretable Prototype-based Graph Information Bottleneck (PGIB) that incorporates prototype learning within the information bottleneck framework to provide prototypes with the key subgraph from the input graph that is important for the model prediction. This is the first work that incorporates prototype learning into the process of identifying the key subgraphs that have a critical impact on the prediction performance. Extensive experiments, including qualitative analysis, demonstrate that PGIB outperforms state-of-the-art methods in terms of both prediction performance and explainability.

NeurIPS Conference 2023 Conference Paper

P-Flow: A Fast and Data-Efficient Zero-Shot TTS through Speech Prompting

  • Sungwon Kim
  • Kevin Shih
  • rohan badlani
  • Joao Felipe Santos
  • Evelina Bakhturina
  • Mikyas Desta
  • Rafael Valle
  • Sungroh Yoon

While recent large-scale neural codec language models have shown significant improvement in zero-shot TTS by training on thousands of hours of data, they suffer from drawbacks such as a lack of robustness, slow sampling speed similar to previous autoregressive TTS methods, and reliance on pre-trained neural codec representations. Our work proposes P-Flow, a fast and data-efficient zero-shot TTS model that uses speech prompts for speaker adaptation. P-Flow comprises a speech-prompted text encoder for speaker adaptation and a flow matching generative decoder for high-quality and fast speech synthesis. Our speech-prompted text encoder uses speech prompts and text input to generate speaker-conditional text representation. The flow matching generative decoder uses the speaker-conditional output to synthesize high-quality personalized speech significantly faster than in real-time. Unlike the neural codec language models, we specifically train P-Flow on LibriTTS dataset using a continuous mel-representation. Through our training method using continuous speech prompts, P-Flow matches the speaker similarity performance of the large-scale zero-shot TTS models with two orders of magnitude less training data and has more than 20$\times$ faster sampling speed. Our results show that P-Flow has better pronunciation and is preferred in human likeness and speaker similarity to its recent state-of-the-art counterparts, thus defining P-Flow as an attractive and desirable alternative. We provide audio samples on our demo page: [https: //research. nvidia. com/labs/adlr/projects/pflow](https: //research. nvidia. com/labs/adlr/projects/pflow)

NeurIPS Conference 2020 Conference Paper

Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search

  • Jaehyeon Kim
  • Sungwon Kim
  • Jungil Kong
  • Sungroh Yoon

Recently, text-to-speech (TTS) models such as FastSpeech and ParaNet have been proposed to generate mel-spectrograms from text in parallel. Despite the advantage, the parallel TTS models cannot be trained without guidance from autoregressive TTS models as their external aligners. In this work, we propose Glow-TTS, a flow-based generative model for parallel TTS that does not require any external aligner. By combining the properties of flows and dynamic programming, the proposed model searches for the most probable monotonic alignment between text and the latent representation of speech on its own. We demonstrate that enforcing hard monotonic alignments enables robust TTS, which generalizes to long utterances, and employing generative flows enables fast, diverse, and controllable speech synthesis. Glow-TTS obtains an order-of-magnitude speed-up over the autoregressive model, Tacotron 2, at synthesis with comparable speech quality. We further show that our model can be easily extended to a multi-speaker setting.

NeurIPS Conference 2020 Conference Paper

NanoFlow: Scalable Normalizing Flows with Sublinear Parameter Complexity

  • Sang-gil Lee
  • Sungwon Kim
  • Sungroh Yoon

Normalizing flows (NFs) have become a prominent method for deep generative models that allow for an analytic probability density estimation and efficient synthesis. However, a flow-based network is considered to be inefficient in parameter complexity because of reduced expressiveness of bijective mapping, which renders the models unfeasibly expensive in terms of parameters. We present an alternative parameterization scheme called NanoFlow, which uses a single neural density estimator to model multiple transformation stages. Hence, we propose an efficient parameter decomposition method and the concept of flow indication embedding, which are key missing components that enable density estimation from a single neural network. Experiments performed on audio and image models confirm that our method provides a new parameter-efficient solution for scalable NFs with significant sublinear parameter complexity.

v2026.09.13