Arrow Research search

Author name cluster

Yue Xing

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
1 author row

Possible papers

11

TMLR Journal 2026 Journal Article

Adversarial Vulnerability from On-Manifold Inseparability and Poor Off-Manifold Convergence

  • Rajdeep Haldar
  • Yue Xing
  • Qifan Song
  • Guang Lin

We introduce a new perspective on adversarial vulnerability in image classification: fragility can arise from poor convergence in off-manifold directions. We model data as lying on low-dimensional manifolds, where on-manifold directions correspond to high-variance, data-aligned features and off-manifold directions capture low-variance, nuanced features. Standard first-order optimizers, such as gradient descent, are inherently ill-conditioned, leading to slow or incomplete convergence in off-manifold directions. When data is inseparable along the on-manifold direction, robustness depends on learning these subtle off-manifold features, and failure to converge leaves models exposed to adversarial perturbations. On the theoretical side, we formalize this mechanism through convergence analyses of logistic regression and two-layer linear networks under first-order methods. These results highlight how ill-conditioning slows or prevents convergence in off-manifold directions, thereby motivating the use of second-order methods which mitigate ill-conditioning and achieve convergence across all directions. Empirically, we demonstrate that even without adversarial training, robustness improves significantly with extended training or second-order optimization, underscoring convergence as a central factor. As an auxiliary empirical finding, we observe that batch normalization suppresses these robustness gains, consistent with its implicit bias toward uniform-margin rather than max-margin solutions. By introducing the notions of on- and off-manifold convergence, this work provides a novel theoretical explanation for adversarial vulnerability.

NeurIPS Conference 2025 Conference Paper

Keeping an Eye on LLM Unlearning: The Hidden Risk and Remedy

  • Jie Ren
  • Zhenwei Dai
  • Xianfeng Tang
  • Yue Xing
  • Shenglai Zeng
  • Jingying Zeng
  • Qiankun Peng
  • Samarth Varshney

Although Large Language Models (LLMs) have demonstrated impressive capabilities across a wide range of tasks, growing concerns have emerged over the misuse of sensitive, copyrighted, or harmful data during training. To address these concerns, unlearning techniques have been developed to remove the influence of specific data without retraining from scratch. However, this paper reveals a critical vulnerability in fine-tuning-based unlearning: a malicious user can craft a manipulated forgetting request that stealthily degrades the model’s utility for benign users. We demonstrate this risk through a red-teaming Stealthy Attack (SA), which is inspired by two key limitations of existing unlearning—the inability to constrain the scope of unlearning effect and the failure to distinguish benign tokens from unlearning signals. Prior work has shown that unlearned models tend to memorize forgetting data as unlearning signals, and respond with hallucinations or feigned ignorance when unlearning signals appear in the input. By subtly increasing the presence of common benign tokens in the forgetting data, SA enhances the connection between benign tokens and unlearning signals. As a result, when normal users include such tokens in their prompts, the model exhibits unlearning behaviors, leading to unintended utility degradation. To address this vulnerability, we propose Scope-aware Unlearning (SU), a lightweight enhancement that introduces a scope term into the unlearning objective, encouraging the model to localize the forgetting effect. Our method requires no additional data processing, integrates seamlessly with existing fine-tuning frameworks, and significantly improves robustness against SA. Extensive experiments validate the effectiveness of both SA and SU.

NeurIPS Conference 2025 Conference Paper

LLM Safety Alignment is Divergence Estimation in Disguise

  • Rajdeep Haldar
  • Ziyi Wang
  • Guang Lin
  • Yue Xing
  • Qifan Song

We present a theoretical framework showing that popular LLM alignment methods—including RLHF and its variants—can be understood as divergence estimators between aligned (safe or preferred) and unaligned (harmful or less-preferred) distributions. This perspective explains the emergence of separation in the latent space between safe and harmful prompts after alignment. As an application of our general divergence framework, we propose KLDO, a novel KL divergence-based alignment method, and empirically validate its effectiveness. We further show that using compliance–refusal datasets, rather than standard preference-based datasets, leads to stronger separation and improved safety alignment. Finally, to quantify the separation effect, we propose a distance-based metric in the prompt representation space, which also acts as a statistically significant indicator for model safety.

TMLR Journal 2024 Journal Article

Stealthy Backdoor Attack via Confidence-driven Sampling

  • Pengfei He
  • Yue Xing
  • Han Xu
  • Jie Ren
  • Yingqian Cui
  • Shenglai Zeng
  • Jiliang Tang
  • Makoto Yamada

Backdoor attacks facilitate unauthorized control in the testing stage by carefully injecting harmful triggers during the training phase of deep neural networks. Previous works have focused on improving the stealthiness of the trigger while randomly selecting samples to attack. However, we find that random selection harms the stealthiness of the model. In this paper, we identify significant pitfalls of random sampling, which make the attacks more detectable and easier to defend against. To improve the stealthiness of existing attacks, we introduce a method of strategically poisoning samples near the model's decision boundary, aiming to minimally alter the model's behavior (decision boundary) before and after backdooring. Our main insight for detecting boundary samples is exploiting the confidence scores as a metric for being near the decision boundary and selecting those to poison (inject) the attack. The proposed approach makes it significantly harder for defenders to identify the attacks. Our method is versatile and independent of any specific trigger design. We provide theoretical insights and conduct extensive experiments to demonstrate the effectiveness of the proposed method.

NeurIPS Conference 2022 Conference Paper

Phase Transition from Clean Training to Adversarial Training

  • Yue Xing
  • Qifan Song
  • Guang Cheng

Adversarial training is one important algorithm to achieve robust machine learning models. However, numerous empirical results show a great performance degradation from clean training to adversarial training (e. g. , 90+\% vs 67\% testing accuracy on CIFAR-10 dataset), which does not match the theoretical guarantee delivered by the existing studies. Such a gap inspires us to explore the existence of an (asymptotic) phase transition phenomenon with respect to the attack strength: adversarial training is as well behaved as clean training in the small-attack regime, but there is a sharp transition from clean training to adversarial training in the large-attack regime. We validate this conjecture in linear regression models, and conduct comprehensive experiments in deep neural networks.

NeurIPS Conference 2022 Conference Paper

Why Do Artificially Generated Data Help Adversarial Robustness

  • Yue Xing
  • Qifan Song
  • Guang Cheng

In the adversarial training framework of \cite{carmon2019unlabeled, gowal2021improving}, people use generated/real unlabeled data with pseudolabels to improve adversarial robustness. We provide statistical insights to explain why the artificially generated data improve adversarial training. In particular, we study how the attack strength and the quality of the unlabeled data affect adversarial robustness in this framework. Our results show that with a high-quality unlabeled data generator, adversarial training can benefit greatly from this framework under large attack strength, while a poor generator can still help to some extent. To make adaptions concerning the quality of generated data, we propose an algorithm that performs online adjustment to the weight between the labeled real data and the generated data, aiming to optimize the adversarial risk. Numerical studies are conducted to verify our theories and show the effectiveness of the proposed algorithm.

YNIMG Journal 2021 Journal Article

Characterizing the seizure onset zone and epileptic network using EEG-fMRI in a rat seizure model

  • Junling Wang
  • Bin Jing
  • Ru Liu
  • Donghong Li
  • Wei Wang
  • Jiaoyang Wang
  • Jianfeng Lei
  • Yue Xing

Accurate epileptogenic zone (EZ) or seizure onset zone (SOZ) localization is crucial for epilepsy surgery optimization. Previous animal and human studies on epilepsy have reported that changes in blood oxygen level-dependent (BOLD) signals induced by epileptic events could be used as diagnostic markers for EZ or SOZ localization. Simultaneous electroencephalography and functional magnetic resonance imaging (EEG-fMRI) recording is gaining interest as a non-invasive tool for preoperative epilepsy evaluation. However, EEG-fMRI studies have reported inconsistent and ambiguous findings. Therefore, it remains unclear whether BOLD responses can be used for accurate EZ or SOZ localization. In this study, we used simultaneous EEG-fMRI recording in a rat model of 4-aminopyridine-induced acute focal seizures to assess the spatial concordance between individual BOLD responses and the SOZ. This was to determine the optimal use of simultaneous EEG-fMRI recording in the SOZ localization. We observed a high spatial consistency between BOLD responses and the SOZ. Further, dynamic BOLD responses were consistent with the regions where the seizures were propagated. These results suggested that simultaneous EEG-fMRI recording could be used as a noninvasive clinical diagnostic technique for localizing the EZ or SOZ and could be an effective tool for mapping epileptic networks.

NeurIPS Conference 2021 Conference Paper

On the Algorithmic Stability of Adversarial Training

  • Yue Xing
  • Qifan Song
  • Guang Cheng

The adversarial training is a popular tool to remedy the vulnerability of deep learning models against adversarial attacks, and there is rich theoretical literature on the training loss of adversarial training algorithms. In contrast, this paper studies the algorithmic stability of a generic adversarial training algorithm, which can further help to establish an upper bound for generalization error. By figuring out the stability upper bound and lower bound, we argue that the non-differentiability issue of adversarial training causes worse algorithmic stability than their natural counterparts. To tackle this problem, we consider a noise injection method. While the non-differentiability problem seriously affects the stability of adversarial training, injecting noise enables the training trajectory to avoid the occurrence of non-differentiability with dominating probability, hence enhancing the stability performance of adversarial training. Our analysis also studies the relation between the algorithm stability and numerical approximation error of adversarial attacks.

NeurIPS Conference 2020 Conference Paper

Directional Pruning of Deep Neural Networks

  • Shih-Kang Chao
  • Zhanyu Wang
  • Yue Xing
  • Guang Cheng

In the light of the fact that the stochastic gradient descent (SGD) often finds a flat minimum valley in the training loss, we propose a novel directional pruning method which searches for a sparse minimizer in or close to that flat region. The proposed pruning method does not require retraining or the expert knowledge on the sparsity level. To overcome the computational formidability of estimating the flat directions, we propose to use a carefully tuned $\ell_1$ proximal gradient algorithm which can provably achieve the directional pruning with a small learning rate after sufficient training. The empirical results demonstrate the promising results of our solution in highly sparse regime (92% sparsity) among many existing pruning methods on the ResNet50 with the ImageNet, while using only a slightly higher wall time and memory footprint than the SGD. Using the VGG16 and the wide ResNet 28x10 on the CIFAR-10 and CIFAR-100, we demonstrate that our solution reaches the same minima valley as the SGD, and the minima found by our solution and the SGD do not deviate in directions that impact the training loss. The code that reproduces the results of this paper is available at https: //github. com/donlan2710/gRDA-Optimizer/tree/master/directional_pruning.

YNICL Journal 2018 Journal Article

Parkinson's disease related signal change in the nigrosomes 1–5 and the substantia nigra using T2* weighted 7T MRI

  • Stefan Theodor Schwarz
  • Olivier Mougin
  • Yue Xing
  • Anna Blazejewska
  • Nin Bajaj
  • Dorothee P. Auer
  • Penny Gowland

Improved markers for the progression of Parkinson's disease (PD) are required. Previous work has proven that iron dependent MRI scans can detect the largest Nigrosome (N1) within the substantia nigra (SN) pars compacta and changes in PD. Histopathological studies have shown that N1 is particularly affected in early PD whereas the other nigrosomes (N2–N5) and the surrounding iron-rich SN are affected later. In this study we aimed to determine whether MRI can detect the smaller nigrosomes (N2–N5) and whether graded signal alterations can be detected on T2*-weighted MRI at different disease stages consistent with histopathological changes. An observational prospective study was performed within the research imaging centre at the University of Nottingham, UK. Altogether 26 individuals with confirmed PD (median Hoehn&Yahr stage = 1, Unified PD Rating Scale [UPDRS] = 12. 5) and 15 healthy controls participated. High resolution T2*weighted 7T MRI of the brain was performed and visibility of N1-N5 within the SN was qualitatively rated. Normalised T2*weighted signal intensities in manually segmented N1–N5 regions and iron-rich SN were calculated. We performed group comparisons and correlations with severity based on UPDRS. Qualitative measures were a nigrosome visibility score and a confidence score for identification. Quantitative measures were T2*weighted contrast of N1–5 and iron-rich SN relative to white matter. We found that visual assessment of the SN for N1–N5 revealed normal range visibility scores in 14 of 15 controls. N1 was identified with the highest confidence and visibility was in abnormal range in all 26 PD patients. The other nigrosomes were less well visible and less confidently identified. There was a larger PD induced signal reduction in all nigrosomes than in the iron-rich SN (median signal difference N1–5 PD compared to controls: 19. 4% [IQR = 24%], iron-rich SN 11% [IQR = 24%, p = 0. 017]). The largest PD induced signal reduction was in N1: 37. 2% [IQR = 19%] which inversely correlated with UPDRS in PD (R2 = 0. 19). All nigrosomes can be detected using 7T MRI, and PD induced T2*weighted signal reduction was greatest in the nigrosomes (especially N1). The graded T2*weighted signal alterations in the nigrosomes match previously described differential histopathological effects of PD. N1 was identified with the highest confidence and T2*weighted signal in N1 correlated with UPDRS confirming N1 as the most promising SN marker of PD pathology.

YNICL Journal 2018 Journal Article

Patterns of grey matter loss associated with motor subscores in early Parkinson's disease

  • Xingfeng Li
  • Yue Xing
  • Antonio Martin-Bastida
  • Paola Piccini
  • Dorothee P. Auer

Classical motor symptoms of Parkinson's disease (PD) such as tremor, rigidity, bradykinesia, and axial symptoms are graded in the Movement Disorders Society Unified Parkinson's Disease Rating Scale (MDS-UPDRS) III. It is yet to be ascertained whether parkinsonian motor symptoms are associated with different anatomical patterns of neurodegeneration as reflected by brain grey matter (GM) alteration. This study aimed to investigate associations between motor subscores and brain GM at voxel level. High resolution structural MRI T1 scans from the Parkinson's Progression Markers Initiative (PPMI) repository were employed to estimate brain GM intensity of PD subjects. Correlations between GM intensity and total MDS-UPDRS III and its four subscores were computed. The total MDS-UPDRS III score was significantly negatively correlated bilaterally with putamen and caudate GM density. Lower anterior striatal GM intensity was significantly associated with higher rigidity subscores, whereas left-sided anterior striatal and precentral cortical GM reduction were correlated with severity of axial symptoms. No significant morphometric associations were demonstrated for tremor subscores. In conclusion, we provide evidence for neuroanatomical patterns underpinning motor symptoms in early PD.

v2026.09.13