Arrow Research search

Author name cluster

Guang Lin

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

13 papers
1 author row

Possible papers

13

TMLR Journal 2026 Journal Article

Adversarial Vulnerability from On-Manifold Inseparability and Poor Off-Manifold Convergence

  • Rajdeep Haldar
  • Yue Xing
  • Qifan Song
  • Guang Lin

We introduce a new perspective on adversarial vulnerability in image classification: fragility can arise from poor convergence in off-manifold directions. We model data as lying on low-dimensional manifolds, where on-manifold directions correspond to high-variance, data-aligned features and off-manifold directions capture low-variance, nuanced features. Standard first-order optimizers, such as gradient descent, are inherently ill-conditioned, leading to slow or incomplete convergence in off-manifold directions. When data is inseparable along the on-manifold direction, robustness depends on learning these subtle off-manifold features, and failure to converge leaves models exposed to adversarial perturbations. On the theoretical side, we formalize this mechanism through convergence analyses of logistic regression and two-layer linear networks under first-order methods. These results highlight how ill-conditioning slows or prevents convergence in off-manifold directions, thereby motivating the use of second-order methods which mitigate ill-conditioning and achieve convergence across all directions. Empirically, we demonstrate that even without adversarial training, robustness improves significantly with extended training or second-order optimization, underscoring convergence as a central factor. As an auxiliary empirical finding, we observe that batch normalization suppresses these robustness gains, consistent with its implicit bias toward uniform-margin rather than max-margin solutions. By introducing the notions of on- and off-manifold convergence, this work provides a novel theoretical explanation for adversarial vulnerability.

AAAI Conference 2026 Conference Paper

Exploring Non-Convex Discrete Energy Landscapes: An Efficient Langevin-Like Sampler with Replica Exchange

  • Haoyang Zheng
  • Hengrong Du
  • Ruqi Zhang
  • Guang Lin

Gradient-based Discrete Samplers (GDSs) are effective for sampling discrete energy landscapes. However, they often stagnate in complex, non-convex settings. To improve exploration, we introduce the Discrete Replica EXchangE Langevin (DREXEL) sampler and its variant with Adjusted Metropolis (DREAM). These samplers use two GDSs at different temperatures and step sizes: one focuses on local exploitation, while the other explores broader energy landscapes. When energy differences are significant, sample swaps occur, governed by a mechanism tailored for discrete sampling to ensure detailed balance. Theoretically, we prove that the proposed samplers satisfy detailed balance and converge to the target distribution under mild conditions. Experiments across 2d synthetic simulations, sampling from Ising models and restricted Boltzmann machines, and training deep energy-based models further confirm their efficiency in exploring non-convex discrete energy landscapes.

EAAI Journal 2025 Journal Article

A self-adaptive energy-based learning rate for stochastic gradient descent via Vector Auxiliary Variable method

  • Jiahao Zhang
  • Christian Moya
  • Guang Lin

Optimizing the learning rate remains a critical challenge in machine learning, essential for achieving model stability and efficient convergence. This paper introduces the Vector Auxiliary Variable (VAV) algorithm, an energy-based self-adaptive learning rate optimization method for stochastic gradient descent(SGD). Unlike traditional methods with fixed learning rates, the VAV method dynamically adjusts learning rates using an auxiliary variable linked directly to the training loss, ensuring significantly larger stable learning rates, rapid convergence, and substantial reductions in training time. Empirical results from regression (Physics-Informed Neural Networks for Burgers’ equation), image classification (on CIFAR-10, CIFAR-100 datasets), and natural language processing (SST-2 sentiment classification) tasks demonstrate improved stability, faster convergence rates (e. g. , achieving 88% accuracy on CIFAR-10 in 30 epochs vs. 80 epochs for SGD), and competitive accuracy performance. Theoretical analysis confirms the unconditional energy dissipation property and convergence guarantees under reasonable assumptions. The VAV method offers practical advantages, particularly beneficial for large-scale real-time machine learning applications in scientific computing and deep learning.

EAAI Journal 2025 Journal Article

High-quality three-dimensional cartoon avatar reconstruction with Gaussian splatting

  • MinHyuk Jang
  • Jong Wook Kim
  • Youngdong Jang
  • Donghyun Kim
  • Wonseok Roh
  • InYong Hwang
  • Guang Lin
  • Sangpil Kim

The growth of the augmented reality industry has increased demand for three-dimensional (3D) cartoon avatars, requiring expertise from computer graphics designers. Recent 3D Gaussian splatting methods have successfully reconstructed 3D avatars from videos, establishing them as a promising solution for this task. However, these methods primarily focus on real-world videos, limiting their effectiveness in the cartoon domain. In this paper, we present an artificial intelligence (AI)-based method for 3D avatar reconstruction from animated cartoon videos, addressing the physically unrealistic and unstructured geometries of cartoons, as well as the varying texture styles across frames. Our surface fitting module models the unstructured geometry of cartoon characters by integrating the surfaces observed from multiple views into a 3D avatar. We design a style normalizer that adjusts color distributions to reduce texture color inconsistencies in each frame of animated cartoons. Additionally, to better capture the simplified color distributions of cartoons, we design a frequency transform loss that focuses on low-frequency components. Our method significantly outperforms state-of-the-art methods, achieving approximately a 25% improvement in Learned Perceptual Image Patch Similarity (LPIPS) with a score of 0. 052 over baselines across the Cartoon Neuman and ToonVid datasets, which comprise 10 videos with diverse styles and poses. Consequently, this paper presents a promising solution to meet the growing demand for high-quality 3D cartoon avatar modeling.

NeurIPS Conference 2025 Conference Paper

LLM Safety Alignment is Divergence Estimation in Disguise

  • Rajdeep Haldar
  • Ziyi Wang
  • Guang Lin
  • Yue Xing
  • Qifan Song

We present a theoretical framework showing that popular LLM alignment methods—including RLHF and its variants—can be understood as divergence estimators between aligned (safe or preferred) and unaligned (harmful or less-preferred) distributions. This perspective explains the emergence of separation in the latent space between safe and harmful prompts after alignment. As an application of our general divergence framework, we propose KLDO, a novel KL divergence-based alignment method, and empirically validate its effectiveness. We further show that using compliance–refusal datasets, rather than standard preference-based datasets, leads to stronger separation and improved safety alignment. Finally, to quantify the separation effect, we propose a distance-based metric in the prompt representation space, which also acts as a statistically significant indicator for model safety.

AAAI Conference 2024 Conference Paper

Federated X-armed Bandit

  • Wenjie Li
  • Qifan Song
  • Jean Honorio
  • Guang Lin

This work establishes the first framework of federated X-armed bandit, where different clients face heterogeneous local objective functions defined on the same domain and are required to collaboratively figure out the global optimum. We propose the first federated algorithm for such problems, named Fed-PNE. By utilizing the topological structure of the global objective inside the hierarchical partitioning and the weak smoothness property, our algorithm achieves sublinear cumulative regret with respect to both the number of clients and the evaluation budget. Meanwhile, it only requires logarithmic communications between the central server and clients, protecting the client privacy. Experimental results on synthetic functions and real datasets validate the advantages of Fed-PNE over various centralized and federated baseline algorithms.

AIIM Journal 2024 Journal Article

Triplet-branch network with contrastive prior-knowledge embedding for disease grading

  • Yuexiang Li
  • Yanping Wang
  • Guang Lin
  • Yawen Huang
  • Jingxin Liu
  • Yi Lin
  • Dong Wei
  • Qirui Zhang

Since different disease grades require different treatments from physicians, i. e. , the low-grade patients may recover with follow-up observations whereas the high-grade may need immediate surgery, the accuracy of disease grading is pivotal in clinical practice. In this paper, we propose a Triplet-Branch Network with ContRastive priOr-knoWledge embeddiNg (TBN-CROWN) for the accurate disease grading, which enables physicians to accordingly take appropriate treatments. Specifically, our TBN-CROWN has three branches, which are implemented for representation learning, classifier learning and grade-related prior-knowledge learning, respectively. The former two branches deal with the issue of class-imbalanced training samples, while the latter one embeds the grade-related prior-knowledge via a novel auxiliary module, termed contrastive embedding module. The proposed auxiliary module takes the features embedded by different branches as input, and accordingly constructs positive and negative embeddings for the model to deploy grade-related prior-knowledge via contrastive learning. Extensive experiments on our private and two publicly available disease grading datasets show that our TBN-CROWN can effectively tackle the class-imbalance problem and yield a satisfactory grading accuracy for various diseases, such as fatigue fracture, ulcerative colitis, and diabetic retinopathy.

TMLR Journal 2023 Journal Article

Federated High-Dimensional Online Decision Making

  • Chi-Hua Wang
  • Wenjie Li
  • Guang Lin

We resolve the main challenge of federated bandit policy design via exploration-exploitation trade-off delineation under data decentralization with a local privacy protection argument. Such a challenge is practical in domain-specific applications and admits another layer of complexity in applications of medical decision-making and web marketing, where high- dimensional decision contexts are sensitive but important to inform decision-making. Exist- ing (low dimensional) federated bandits suffer super-linear theoretical regret upper bound in high-dimensional scenarios and are at risk of client information leakage due to their in- ability to separate exploration from exploitation. This paper proposes a class of bandit policy design, termed Fedego Lasso, to complete the task of federated high-dimensional online decision-making with sub-linear theoretical regret and local client privacy argument. Fedego Lasso relies on a novel multi-client teamwork-selfish bandit policy design to per- form decentralized collaborative exploration and federated egocentric exploration with log- arithmic communication costs. Experiments demonstrate the effectiveness of the proposed algorithms on both synthetic and real-world datasets.

EAAI Journal 2023 Journal Article

Learning the dynamical response of nonlinear non-autonomous dynamical systems with deep operator neural networks

  • Guang Lin
  • Christian Moya
  • Zecheng Zhang

We propose using operator learning to approximate the dynamical response of non-autonomous systems, such as nonlinear control systems. Unlike classical function learning, operator learning maps between two function spaces, does not require discretization of the output function, and provides flexibility in data preparation and solution prediction. Particularly, we apply and redesign the Deep Operator Neural Network (DeepONet) to recursively learn the solution trajectories of the dynamical systems. Our approach involves constructing and training a DeepONet that approximates the system’s local solution operator. We then develop a numerical scheme that recursively simulates the system’s long/medium-term dynamic response for given inputs and initial conditions, using the trained DeepONet. We accompany the proposed scheme with an estimate for the error bound of the associated cumulative error. Moreover, we propose a data-driven Runge–Kutta (RK) explicit scheme that leverages the DeepONet’s forward pass and automatic differentiation to better approximate the system’s response when the numerical scheme’s step size is small. Numerical experiments on the predator–prey, pendulum, and cart pole systems demonstrate that our proposed DeepONet framework effectively learns to approximate the dynamical response of non-autonomous systems with time-dependent inputs.

AAAI Conference 2023 Conference Paper

Non-reversible Parallel Tempering for Deep Posterior Approximation

  • Wei Deng
  • Qian Zhang
  • Qi Feng
  • Faming Liang
  • Guang Lin

Parallel tempering (PT), also known as replica exchange, is the go-to workhorse for simulations of multi-modal distributions. The key to the success of PT is to adopt efficient swap schemes. The popular deterministic even-odd (DEO) scheme exploits the non-reversibility property and has successfully reduced the communication cost from quadratic to linear given the sufficiently many chains. However, such an innovation largely disappears in big data due to the limited chains and few bias-corrected swaps. To handle this issue, we generalize the DEO scheme to promote non-reversibility and propose a few solutions to tackle the underlying bias caused by the geometric stopping time. Notably, in big data scenarios, we obtain a nearly linear communication cost based on the optimal window size. In addition, we also adopt stochastic gradient descent (SGD) with large and constant learning rates as exploration kernels. Such a user-friendly nature enables us to conduct approximation tasks for complex posteriors without much tuning costs.

TMLR Journal 2022 Journal Article

Deformation Robust Roto-Scale-Translation Equivariant CNNs

  • Liyao Gao
  • Guang Lin
  • Wei Zhu

Incorporating group symmetry directly into the learning process has proved to be an effective guideline for model design. By producing features that are guaranteed to transform covariantly to the group actions on the inputs, group-equivariant convolutional neural networks (G-CNNs) achieve significantly improved generalization performance in learning tasks with intrinsic symmetry. General theory and practical implementation of G-CNNs have been studied for planar images under either rotation or scaling transformation, but only individually. We present, in this paper, a roto-scale-translation equivariant CNN ($\mathcal{RST}$-CNN), that is guaranteed to achieve equivariance jointly over these three groups via coupled group convolutions. Moreover, as symmetry transformations in reality are rarely perfect and typically subject to input deformation, we provide a stability analysis of the equivariance of representation to input distortion, which motivates the truncated expansion of the convolutional filters under (pre-fixed) low-frequency spatial modes. The resulting model provably achieves deformation-robust $\mathcal{RST}$ equivariance, i.e., the $\mathcal{RST}$ symmetry is still "approximately” preserved when the transformation is "contaminated” by a nuisance data deformation, a property that is especially important for out-of-distribution generalization. Numerical experiments on MNIST, Fashion-MNIST, and STL-10 demonstrate that the proposed model yields remarkable gains over prior arts, especially in the small data regime where both rotation and scaling variations are present within the data.

NeurIPS Conference 2020 Conference Paper

A Contour Stochastic Gradient Langevin Dynamics Algorithm for Simulations of Multi-modal Distributions

  • Wei Deng
  • Guang Lin
  • Faming Liang

We propose an adaptively weighted stochastic gradient Langevin dynamics algorithm (SGLD), so-called contour stochastic gradient Langevin dynamics (CSGLD), for Bayesian learning in big data statistics. The proposed algorithm is essentially a scalable dynamic importance sampler, which automatically flattens the target distribution such that the simulation for a multi-modal distribution can be greatly facilitated. Theoretically, we prove a stability condition and establish the asymptotic convergence of the self-adapting parameter to a unique fixed-point, regardless of the non-convexity of the original energy function; we also present an error analysis for the weighted averaging estimators. Empirically, the CSGLD algorithm is tested on multiple benchmark datasets including CIFAR10 and CIFAR100. The numerical results indicate its superiority over the existing state-of-the-art algorithms in training deep neural networks.

NeurIPS Conference 2019 Conference Paper

An Adaptive Empirical Bayesian Method for Sparse Deep Learning

  • Wei Deng
  • Xiao Zhang
  • Faming Liang
  • Guang Lin

We propose a novel adaptive empirical Bayesian (AEB) method for sparse deep learning, where the sparsity is ensured via a class of self-adaptive spike-and-slab priors. The proposed method works by alternatively sampling from an adaptive hierarchical posterior distribution using stochastic gradient Markov Chain Monte Carlo (MCMC) and smoothly optimizing the hyperparameters using stochastic approximation (SA). The convergence of the proposed method to the asymptotically correct distribution is established under mild conditions. Empirical applications of the proposed method lead to the state-of-the-art performance on MNIST and Fashion MNIST with shallow convolutional neural networks (CNN) and the state-of-the-art compression performance on CIFAR10 with Residual Networks. The proposed method also improves resistance to adversarial attacks.

v2026.09.13