Arrow Research search

Author name cluster

Yuan He

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

23 papers
2 author rows

Possible papers

23

AAAI Conference 2026 Conference Paper

Dual-Seed Evolutionary Algorithm for Noise Optimization in Diffusion Models

  • Yuzheng Tan
  • Yuan He
  • Yao Zhu
  • Tianlin Huo
  • Huanqian Yan
  • Hang Su
  • Shuxin Zhang
  • Guangneng Hu

Diffusion models have emerged as state-of-the-art generative methods, particularly excelling in conditional tasks such as prompt-driven image synthesis. While recent research emphasizes the pivotal role of noise seeds in enhancing text-image alignment and generating human-preferred outputs,these works predominantly rely on random Gaussian noise or heuristic local adjustments,, overlooking the potential of global optimization trategies to systematically improve generation quality. To bridge this gap, we propose Seed Optimization based on Evolution (SOE), a hybrid framework that integrates global evolutionary search with local semantic refinement. The global evolutionary stage conducts seed selection by jointly optimizing text-image alignment (via CLIP-Score) and human preference estimation (via ImageReward), while the local stage employs diffusion inversion to inject conditional semantics into the noise seed. Together, these components constitute a model-agnostic, training-free optimization framework for conditional diffusion models. Extensive experiments across various diffusion models demonstrate that SOE consistently improves semantic fidelity and visual quality, highlighting its generalizability and potential as a plug-and-play enhancement for generative diffusion pipelines.

JBHI Journal 2026 Journal Article

PGST: A prototype-guided parameter-efficient network for spatial transcriptomics prediction

  • Yuan He
  • Kaimiao Hu
  • Changming Sun
  • Leyi Wei
  • Ran Su

Spatial transcriptomics (ST) aims to decode spatially resolved gene expression patterns while preserving tissue morphology. Current methods tend to use lower-cost deep learning approaches for gene expression prediction, yet face severe challenges. First, existing methods fail to give sufficient consideration to the spatial specificity of positional encoding inherent in ST; second, they neglect to leverage spatially coherent co-expression patterns across different domains; third, their reliance on linearly weighted aggregation induces vulnerability to noise and distribution shifts; and finally, these architectures exhibit limited parameter efficiency. To address these issues, we introduce prototype-guided network for spatial transcriptomics (PGST), which includes four parts: (1) oriented signal propagation through polar embedding strategy for spatial transcriptomics (PEST); (2) prototype-guided aggregation for global co-feature preservation; (3) global consistency enforcement via shared decoder with reconstruction loss; and (4) lightweight architectural design. Our framework integrates contrastive learning with graph neural networks to balance local-global spatial dependencies and cross-modal consistency. Experimental results on multiple datasets from ST demonstrate the superior performance of our PGST model than existing methods. Our source code is available at: https://github.com/RanSuLab/PGST https://github.com/RanSuLab/PGST.

JBHI Journal 2025 Journal Article

BFGTP: A BERT-Guided Two-Stage Molecular Representation Learning Framework for Toxicity Prediction

  • Kaimiao Hu
  • Yuan He
  • Jianguo Wei
  • Changming Sun
  • Jie Geng
  • Leyi Wei
  • Ran Su

Accurate prediction of molecular toxicity is vital for drug development. Most mainstream methods rely on fingerprints or graph-based feature extraction, the emergence of large language models (LLMs) offers new prospects for molecular representation learning in toxicity prediction. Although several studies attempt to leverage LLMs to integrate molecular sequence data for pretraining molecular representations, certain limitations remain. Current LLM-based approaches usually utilize solely on class embedding features, overlooking the rich information in sequence embedding. Moreover, integrating pre-trained molecular representations with multi-modal molecular data may further enhance performance in toxicity prediction. To address these challenges, we propose BFGTP, a BERT-guided two-stage molecular representation learning framework for toxicity prediction. Firstly, we design independent encoders for molecular descriptions of three modalities, where the fingerprint encoder with dual level attention mechanisms effectively integrates multi-category fingerprints. Then, the two-stage guide strategy is introduced to fully utilize the prior knowledge of LLMs, employing contrastive learning to align and fuse the tri-modal representations and knowledge distillation to align predicted value distributions. BFGTP ultimately combines fingerprint and graph representations to predict molecular toxicity. Experiments on seven toxicity datasets show that BFGTP outperforms baselines, achieving the highest AUC on five datasets and the best average performance across five evaluation metrics. Ablation studies, t-SNE visualization and case study confirm the effectiveness of BFGTP's components and its ability to capture meaningful molecular representations.

TMLR Journal 2025 Journal Article

DyGMamba: Efficiently Modeling Long-Term Temporal Dependency on Continuous-Time Dynamic Graphs with State Space Models

  • Zifeng Ding
  • Yifeng Li
  • Yuan He
  • Antonio Norelli
  • Jingcheng Wu
  • Volker Tresp
  • Michael M. Bronstein
  • Yunpu Ma

Learning useful representations for continuous-time dynamic graphs (CTDGs) is challenging, due to the concurrent need to span long node interaction histories and grasp nuanced temporal details. In particular, two problems emerge: (1) Encoding longer histories requires more computational resources, making it crucial for CTDG models to maintain low computational complexity to ensure efficiency; (2) Meanwhile, more powerful models are needed to identify and select the most critical temporal information within the extended context provided by longer histories. To address these problems, we propose a CTDG representation learning model named DyGMamba, originating from the popular Mamba state space model (SSM). DyGMamba first leverages a node-level SSM to encode the sequence of historical node interactions. Another time-level SSM is then employed to exploit the temporal patterns hidden in the historical graph, where its output is used to dynamically select the critical information from the interaction history. We validate DyGMamba experimentally on the dynamic link prediction task. The results show that our model achieves state-of-the-art in most cases. DyGMamba also maintains high efficiency in terms of computational resources, making it possible to capture long temporal dependencies with a limited computation budget.

NeurIPS Conference 2025 Conference Paper

Impact of Dataset Properties on Membership Inference Vulnerability of Deep Transfer Learning

  • Marlon Tobaben
  • Hibiki Ito
  • Joonas Jälkö
  • Yuan He
  • Antti Honkela

Membership inference attacks (MIAs) are used to test practical privacy of machine learning models. MIAs complement formal guarantees from differential privacy (DP) under a more realistic adversary model. We analyse MIA vulnerability of fine-tuned neural networks both empirically and theoretically, the latter using a simplified model of fine-tuning. We show that the vulnerability of non-DP models when measured as the attacker advantage at a fixed false positive rate reduces according to a simple power law as the number of examples per class increases. A similar power-law applies even for the most vulnerable points, but the dataset size needed for adequate protection of the most vulnerable points is very large.

YNICL Journal 2025 Journal Article

Precision targeting of right dorsolateral prefrontal cortex with neuronavigated rTMS alleviates chronic insomnia via functional connectivity reorganization: a randomized neuroimaging trial

  • Liang Gong
  • Xi Yang
  • Yuan He
  • Haoyu Li
  • Wen Zhou
  • Duan Liu
  • Bei Zhang
  • Chunhua Xi

BACKGROUND: Repetitive transcranial magnetic stimulation (rTMS) offers a promising approach for the treatment of insomnia; however, the precise targets and underlying neural mechanisms remain unclear. This randomized, wait-controlled trial aimed to evaluate the clinical efficacy of neuronavigated rTMS targeting the right dorsolateral prefrontal cortex (DLPFC) in chronic insomnia disorder (CID) and to identify potential neural mechanisms associated with therapeutic outcomes. METHODS: Fifty patients with CID were randomized to receive 20 sessions of 1 Hz rTMS targeting the right DLPFC or to a waitlist control group. Stimulation coordinates were selected (MNI: 40,39,11) based on our previous neuroimaging meta-analysis, and were precisely localized using MRI-guided neuronavigation. Clinical assessments and resting-state fMRI were conducted before and after the intervention, respectively. Target-based functional connectivity (FC) analysis was used to map rTMS-associated network changes, while causal mediation analysis was used to examine the relationships between neural changes and clinical improvements. RESULTS: Compared to waitlist controls, the rTMS group showed greater improvements in insomnia and mood symptoms (all p < 0.001), with higher response rates (54.55 % vs. 9.09 %) and remission rates (68.18 % vs. 13.64 %). FC analysis showed significant group × time effects on the bilateral DLPFC, middle cingulate cortex, and right anterior cerebellar vermis. Mediation analysis indicated that FC changes in the right DLPFC mediated 24 % of the improvement in insomnia severity (Insomnia Severity Index, p = 0.048). CONCLUSION: These preliminary findings suggest that precision neuronavigated rTMS targeting the right DLPFC may alleviate insomnia symptoms, with the observed clinical improvements potentially related to the reorganization of the DLPFC network. While these results are encouraging, further research based on placebo-controlled study designs is required to confirm these effects and better understand the underlying mechanisms. This study provides preliminary evidence supporting the integration of precision targeting with neuroimaging to explore the mechanisms underlying the effects of rTMS in insomnia treatment.

IROS Conference 2025 Conference Paper

Stability Enhancement in Variable Morphing Multi-body AUVs for Underwater Structure Maintenance

  • Shuai Kang
  • Yuan He
  • Jin Zhang
  • Yuxi Gao
  • Yunfei Bai
  • Longchuan Li

This paper presents a Variable Morphing Multi-Body AUVs (VMMAUVs) concept, designed for underwater structure maintenance. This robot is capable of dynamically adjusting their structure to adapt to varying operational scenarios. The study explores two key stability mechanisms: buoyancy adjustment and aperture angle control, both aimed at optimizing the metacentric height. Through simulations and experiments with different buoyancy configurations and aperture angles, the results show that the proposed methods significantly enhance the system ’ s stability, enabling faster convergence and better posture retention. The feasibility of the control strategies is validated through various numerical simulations, demonstrating the effectiveness of angle tracking control and buoyancy adjustment in maintaining stability under dynamic oceanic conditions.

ICRA Conference 2025 Conference Paper

The Devil is in the Quality: Exploring Informative Samples for Semi-Supervised Monocular 3D Object Detection

  • Zhipeng Zhang
  • Zhenyu Li 0007
  • Hanshi Wang
  • Yuan He
  • Ke Wang
  • Heng Fan 0001

This paper tackles the challenging problem of semi-supervised monocular 3D object detection with a general framework. In specific, having observed that the bottleneck of this task lies in lacking reliable and informative samples from unlabeled data for detector learning, we introduce a novel simple yet effective ‘Augment and Criticize’ pipeline that mines abundant informative samples for robust detection. To be more specific, in the ‘Augment’ stage, we present the Augmentation-based Prediction aGgregation (APG), which applies automatically learned transformations to unlabeled images and aggregates detections from various augmented views as pseudo labels. Since not all the pseudo labels from APG are beneficially informative, the subsequent ‘Criticize’ phase is introduced. Particularly, we present the Critical Retraining Strategy (CRS) that, unlike simply filtering pseudo labels using a fixed threshold, employs a learnable network to evaluate the contribution of unlabeled images at different training timestamps. This way, the noisy samples prohibitive to model evolution can be effectively suppressed. In order to validate ‘Augment-Criticize’, we apply it to MonoDLE [1] and MonoFlex [2], and the two new detectors, dubbed 3DSeMo DLE and 3DSeMo FLEX, achieve state-of-the-art results with consistent improvements, evidencing its effectiveness and generality.

JBHI Journal 2025 Journal Article

TPNET: A Time-Sensitive Small Sample Multimodal Network for Cardiotoxicity Risk Prediction

  • Yuan He
  • Fengyun Zhang
  • Kaimiao Hu
  • Changming Sun
  • Jie Geng
  • Ning Ren
  • Ran Su

Cancer therapy-related cardiac dysfunction (CTRCD) is a potential complication associated with cancer treatment, particularly in patients with breast cancer, requiring monitoring of cardiac health during the treatment process. Tissue Doppler imaging (TDI) is a remarkable technique that can provide a comprehensive reflection of the left ventricle's physiological status. We hypothesized that the combination of TDI features with deep learning techniques could be utilized to predict CTRCD. To evaluate the hypothesis, we developed a temporal-multimodal pattern network for efficient training (TPNET) model to predict the incidence of CTRCD over a 24-month period based on TDI, function, and clinical data from 270 patients. Our model achieved an area under curve (AUC) of 0. 83 and sensitivity of 0. 88, demonstrating greater robustness compared to other existing visual models. To further translate our model's findings into practical applications, we utilized the integrated gradients (IG) attribution to perform a detailed evaluation of all the features. This analysis has identified key pathogenic signs that may have remained unnoticed, providing a viable option for implementing our model in preoperative breast cancer patients. Additionally, our findings demonstrate the potential of TPNET in discovering new causative agents for CTRCD.

NeurIPS Conference 2024 Conference Paper

Language Models as Hierarchy Encoders

  • Yuan He
  • Zhangdie Yuan
  • Jiaoyan Chen
  • Ian Horrocks

Interpreting hierarchical structures latent in language is a key limitation of current language models (LMs). While previous research has implicitly leveraged these hierarchies to enhance LMs, approaches for their explicit encoding are yet to be explored. To address this, we introduce a novel approach to re-train transformer encoder-based LMs as Hierarchy Transformer encoders (HiTs), harnessing the expansive nature of hyperbolic space. Our method situates the output embedding space of pre-trained LMs within a Poincaré ball with a curvature that adapts to the embedding dimension, followed by re-training on hyperbolic clustering and centripetal losses. These losses are designed to effectively cluster related entities (input as texts) and organise them hierarchically. We evaluate HiTs against pre-trained LMs, standard fine-tuned LMs, and several hyperbolic embedding baselines, focusing on their capabilities in simulating transitive inference, predicting subsumptions, and transferring knowledge across hierarchies. The results demonstrate that HiTs consistently outperform all baselines in these tasks, underscoring the effectiveness and transferability of our re-trained hierarchy encoders.

NeurIPS Conference 2023 Conference Paper

A Unified Generalization Analysis of Re-Weighting and Logit-Adjustment for Imbalanced Learning

  • Zitai Wang
  • Qianqian Xu
  • Zhiyong Yang
  • Yuan He
  • Xiaochun Cao
  • Qingming Huang

Real-world datasets are typically imbalanced in the sense that only a few classes have numerous samples, while many classes are associated with only a few samples. As a result, a naive ERM learning process will be biased towards the majority classes, making it difficult to generalize to the minority classes. To address this issue, one simple but effective approach is to modify the loss function to emphasize the learning on minority classes, such as re-weighting the losses or adjusting the logits via class-dependent terms. However, existing generalization analysis of such losses is still coarse-grained and fragmented, failing to explain some empirical results. To bridge this gap between theory and practice, we propose a novel technique named data-dependent contraction to capture how these modified losses handle different classes. On top of this technique, a fine-grained generalization bound is established for imbalanced learning, which helps reveal the mystery of re-weighting and logit-adjustment in a unified manner. Furthermore, a principled learning algorithm is developed based on the theoretical insights. Finally, the empirical results on benchmark datasets not only validate the theoretical results but also demonstrate the effectiveness of the proposed method.

AAAI Conference 2023 Conference Paper

Towards Decision-Friendly AUC: Learning Multi-Classifier with AUCµ

  • Peifeng Gao
  • Qianqian Xu
  • Peisong Wen
  • Huiyang Shao
  • Yuan He
  • Qingming Huang

Area Under the ROC Curve (AUC) is a widely used ranking metric in imbalanced learning due to its insensitivity to label distributions. As a well-known multiclass extension of AUC, Multiclass AUC (MAUC, a.k.a. M-metric) measures the average AUC of multiple binary classifiers. In this paper, we argue that simply optimizing MAUC is far from enough for imbalanced multi-classification. More precisely, MAUC only focuses on learning scoring functions via ranking optimization, while leaving the decision process unconsidered. Therefore, scoring functions being able to make good decisions might suffer from low performance in terms of MAUC. To overcome this issue, we turn to explore AUCµ, another multiclass variant of AUC, which further takes the decision process into consideration. Motivated by this fact, we propose a surrogate risk optimization framework to improve model performance from the perspective of AUCµ. Practically, we propose a two-stage training framework for multi-classification, where at the first stage a scoring function is learned maximizing AUCµ, and at the second stage we seek for a decision function to improve the F1-metric via our proposed soft F1. Theoretically, we first provide sufficient conditions that optimizing the surrogate losses could lead to the Bayes optimal scoring function. Afterward, we show that the proposed surrogate risk enjoys a generalization bound in order of O(1/√N). Experimental results on four benchmark datasets demonstrate the effectiveness of our proposed method in both AUCµ and F1-metric.

AAAI Conference 2022 Conference Paper

BERTMap: A BERT-Based Ontology Alignment System

  • Yuan He
  • Jiaoyan Chen
  • Denvar Antonyrajah
  • Ian Horrocks

Ontology alignment (a. k. a ontology matching (OM)) plays a critical role in knowledge integration. Owing to the success of machine learning in many domains, it has been applied in OM. However, the existing methods, which often adopt adhoc feature engineering or non-contextual word embeddings, have not yet outperformed rule-based systems especially in an unsupervised setting. In this paper, we propose a novel OM system named BERTMap which can support both unsupervised and semi-supervised settings. It first predicts mappings using a classifier based on fine-tuning the contextual embedding model BERT on text semantics corpora extracted from ontologies, and then refines the mappings through extension and repair by utilizing the ontology structure and logic. Our evaluation with three alignment tasks on biomedical ontologies demonstrates that BERTMap can often perform better than the leading OM systems LogMap and AML.

NeurIPS Conference 2022 Conference Paper

Exploring the Algorithm-Dependent Generalization of AUPRC Optimization with List Stability

  • Peisong Wen
  • Qianqian Xu
  • Zhiyong Yang
  • Yuan He
  • Qingming Huang

Stochastic optimization of the Area Under the Precision-Recall Curve (AUPRC) is a crucial problem for machine learning. Although various algorithms have been extensively studied for AUPRC optimization, the generalization is only guaranteed in the multi-query case. In this work, we present the first trial in the single-query generalization of stochastic AUPRC optimization. For sharper generalization bounds, we focus on algorithm-dependent generalization. There are both algorithmic and theoretical obstacles to our destination. From an algorithmic perspective, we notice that the majority of existing stochastic estimators are biased when the sampling strategy is biased, and is leave-one-out unstable due to the non-decomposability. To address these issues, we propose a sampling-rate-invariant unbiased stochastic estimator with superior stability. On top of this, the AUPRC optimization is formulated as a composition optimization problem, and a stochastic algorithm is proposed to solve this problem. From a theoretical perspective, standard techniques of the algorithm-dependent generalization analysis cannot be directly applied to such a listwise compositional optimization problem. To fill this gap, we extend the model stability from instancewise losses to listwise losses and bridge the corresponding generalization and stability. Additionally, we construct state transition matrices to describe the recurrence of the stability, and simplify calculations by matrix spectrum. Practically, experimental results on three image retrieval datasets on speak to the effectiveness and soundness of our framework.

NeurIPS Conference 2022 Conference Paper

OpenAUC: Towards AUC-Oriented Open-Set Recognition

  • Zitai Wang
  • Qianqian Xu
  • Zhiyong Yang
  • Yuan He
  • Xiaochun Cao
  • Qingming Huang

Traditional machine learning follows a close-set assumption that the training and test set share the same label space. While in many practical scenarios, it is inevitable that some test samples belong to unknown classes (open-set). To fix this issue, Open-Set Recognition (OSR), whose goal is to make correct predictions on both close-set samples and open-set samples, has attracted rising attention. In this direction, the vast majority of literature focuses on the pattern of open-set samples. However, how to evaluate model performance in this challenging task is still unsolved. In this paper, a systematic analysis reveals that most existing metrics are essentially inconsistent with the aforementioned goal of OSR: (1) For metrics extended from close-set classification, such as Open-set F-score, Youden's index, and Normalized Accuracy, a poor open-set prediction can escape from a low performance score with a superior close-set prediction. (2) Novelty detection AUC, which measures the ranking performance between close-set and open-set samples, ignores the close-set performance. To fix these issues, we propose a novel metric named OpenAUC. Compared with existing metrics, OpenAUC enjoys a concise pairwise formulation that evaluates open-set performance and close-set performance in a coupling manner. Further analysis shows that OpenAUC is free from the aforementioned inconsistency properties. Finally, an end-to-end learning method is proposed to minimize the OpenAUC risk, and the experimental results on popular benchmark datasets speak to its effectiveness.

NeurIPS Conference 2022 Conference Paper

OTKGE: Multi-modal Knowledge Graph Embeddings via Optimal Transport

  • Zongsheng Cao
  • Qianqian Xu
  • Zhiyong Yang
  • Yuan He
  • Xiaochun Cao
  • Qingming Huang

Multi-modal knowledge graph embeddings (KGE) have caught more and more attention in learning representations of entities and relations for link prediction tasks. Different from previous uni-modal KGE approaches, multi-modal KGE can leverage expressive knowledge from a wealth of modalities (image, text, etc. ), leading to more comprehensive representations of real-world entities. However, the critical challenge along this course lies in that the multi-modal embedding spaces are usually heterogeneous. In this sense, direct fusion will destroy the inherent spatial structure of different modal embeddings. To overcome this challenge, we revisit multi-modal KGE from a distributional alignment perspective and propose optimal transport knowledge graph embeddings (OTKGE). Specifically, we model the multi-modal fusion procedure as a transport plan moving different modal embeddings to a unified space by minimizing the Wasserstein distance between multi-modal distributions. Theoretically, we show that by minimizing the Wasserstein distance between the individual modalities and the unified embedding space, the final results are guaranteed to maintain consistency and comprehensiveness. Moreover, experimental results on well-established multi-modal knowledge graph completion benchmarks show that our OTKGE achieves state-of-the-art performance.

IJCAI Conference 2022 Conference Paper

RMGN: A Regional Mask Guided Network for Parser-free Virtual Try-on

  • Chao Lin
  • Zhao Li
  • Sheng Zhou
  • Shichang Hu
  • Jialun Zhang
  • Linhao Luo
  • Jiarun Zhang
  • Longtao Huang

Virtual try-on (VTON) aims at fitting target clothes to reference person images, which is widely adopted in e-commerce. Existing VTON approaches can be narrowly categorized into Parser-Based (PB) and Parser-Free (PF) by whether relying on the parser information to mask the persons’clothes and synthesize try-on images. Although abandoning parser information has improved the applicability of PF methods, the ability of detail synthesizing has also been sacrificed. As a result, the distraction from original cloth may persist in synthesized images, especially in complicated postures and high resolution applications. To address the aforementioned issue, we propose a novel PF method named Regional Mask Guided Network (RMGN). More specifically, a regional mask is proposed to explicitly fuse the features of target clothes and reference persons so that the persisted distraction can be eliminated. A posture awareness loss and a multi-level feature extractor are further proposed to handle the complicated postures and synthesize high resolution images. Extensive experiments demonstrate that our proposed RMGN outperforms both state-of-the-art PB and PF methods. Ablation studies further verify the effectiveness of modules in RMGN. Code is available at https: //github. com/jokerlc/RMGN-VITON.

NeurIPS Conference 2022 Conference Paper

The Minority Matters: A Diversity-Promoting Collaborative Metric Learning Algorithm

  • Shilong Bao
  • Qianqian Xu
  • Zhiyong Yang
  • Yuan He
  • Xiaochun Cao
  • Qingming Huang

Collaborative Metric Learning (CML) has recently emerged as a popular method in recommendation systems (RS), closing the gap between metric learning and Collaborative Filtering. Following the convention of RS, existing methods exploit unique user representation in their model design. This paper focuses on a challenging scenario where a user has multiple categories of interests. Under this setting, we argue that the unique user representation might induce preference bias, especially when the item category distribution is imbalanced. To address this issue, we propose a novel method called Diversity-Promoting Collaborative Metric Learning (DPCML), with the hope of considering the commonly ignored minority interest of the user. The key idea behind DPCML is to include a multiple set of representations for each user in the system. Based on this embedding paradigm, user preference toward an item is aggregated from different embeddings by taking the minimum item-user distance among the user embedding set. Furthermore, we observe that the diversity of the embeddings for the same user also plays an essential role in the model. To this end, we propose a diversity control regularization term to accommodate the multi-vector representation strategy better. Theoretically, we show that DPCML could generalize well to unseen test data by tackling the challenge of the annoying operation that comes from the minimum value. Experiments over a range of benchmark datasets speak to the efficacy of DPCML.

AAAI Conference 2021 Conference Paper

Composite Adversarial Attacks

  • Xiaofeng Mao
  • Yuefeng Chen
  • Shuhui Wang
  • Hang Su
  • Yuan He
  • Hui Xue

Adversarial attack is a technique for deceiving Machine Learning (ML) models, which provides a way to evaluate the adversarial robustness. In practice, attack algorithms are artificially selected and tuned by human experts to break a ML system. However, manual selection of attackers tends to be sub-optimal, leading to a mistakenly assessment of model security. In this paper, a new procedure called Composite Adversarial Attack (CAA) is proposed for automatically searching the best combination of attack algorithms and their hyperparameters from a candidate pool of 32 base attackers. We design a search space where attack policy is represented as an attacking sequence, i. e. , the output of the previous attacker is used as the initialization input for successors. Multiobjective NSGA-II genetic algorithm is adopted for finding the strongest attack policy with minimum complexity. The experimental result shows CAA beats 10 top attackers on 11 diverse defenses with less elapsed time (6 × faster than AutoAttack), and achieves the new state-of-the-art on l∞, l2 and unrestricted adversarial attacks.

NeurIPS Conference 2021 Conference Paper

When False Positive is Intolerant: End-to-End Optimization with Low FPR for Multipartite Ranking

  • Peisong Wen
  • Qianqian Xu
  • Zhiyong Yang
  • Yuan He
  • Qingming Huang

Multipartite ranking is a basic task in machine learning, where the Area Under the receiver operating characteristics Curve (AUC) is generally applied as the evaluation metric. Despite that AUC reflects the overall performance of the model, it is inconsistent with the expected performance in some application scenarios, where only a low False Positive Rate (FPR) is meaningful. To leverage high performance under low FPRs, we consider an alternative metric for multipartite ranking evaluating the True Positive Rate (TPR) at a given FPR, denoted as TPR@FPR. Unfortunately, the key challenge of direct TPR@FPR optimization is two-fold: \textbf{a)} the original objective function is not differentiable, making gradient backpropagation impossible; \textbf{b)} the loss function could not be written as a sum of independent instance-wise terms, making mini-batch based optimization infeasible. To address these issues, we propose a novel framework on top of the deep learning framework named \textit{Cross-Batch Approximation for Multipartite Ranking (CBA-MR)}. In face of \textbf{a)}, we propose a differentiable surrogate optimization problem where the instances having a short-time effect on FPR are rendered with different weights based on the random walk hypothesis. To tackle \textbf{b)}, we propose a fast ranking estimation method, where the full-batch loss evaluation is replaced by a delayed update scheme with the help of an embedding cache. Finally, experimental results on four real-world benchmarks are provided to demonstrate the effectiveness of the proposed method.

AAAI Conference 2020 Conference Paper

Fine-Grained Fashion Similarity Learning by Attribute-Specific Embedding Network

  • Zhe Ma
  • Jianfeng Dong
  • Zhongzi Long
  • Yao Zhang
  • Yuan He
  • Hui Xue
  • Shouling Ji

This paper strives to learn fine-grained fashion similarity. In this similarity paradigm, one should pay more attention to the similarity in terms of a specific design/attribute among fashion items, which has potential values in many fashion related applications such as fashion copyright protection. To this end, we propose an Attribute-Specific Embedding Network (ASEN) to jointly learn multiple attributespecific embeddings in an end-to-end manner, thus measure the fine-grained similarity in the corresponding space. With two attention modules, i. e. , Attribute-aware Spatial Attention and Attribute-aware Channel Attention, ASEN is able to locate the related regions and capture the essential patterns under the guidance of the specified attribute, thus make the learned attribute-specific embeddings better reflect the fine-grained similarity. Extensive experiments on four fashion-related datasets show the effectiveness of ASEN for fine-grained fashion similarity learning and its potential for fashion reranking. Code and data are available at https: //github. com/Maryeon/asen.

NeurIPS Conference 2020 Conference Paper

Heuristic Domain Adaptation

  • Shuhao Cui
  • Xuan Jin
  • Shuhui Wang
  • Yuan He
  • Qingming Huang

In visual domain adaptation (DA), separating the domain-specific characteristics from the domain-invariant representations is an ill-posed problem. Existing methods apply different kinds of priors or directly minimize the domain discrepancy to address this problem, which lack flexibility in handling real-world situations. Another research pipeline expresses the domain-specific information as a gradual transferring process, which tends to be suboptimal in accurately removing the domain-specific properties. In this paper, we address the modeling of domain-invariant and domain-specific information from the heuristic search perspective. We identify the characteristics in the existing representations that lead to larger domain discrepancy as the heuristic representations. With the guidance of heuristic representations, we formulate a principled framework of Heuristic Domain Adaptation (HDA) with well-founded theoretical guarantees. To perform HDA, the cosine similarity scores and independence measurements between domain-invariant and domain-specific representations are cast into the constraints at the initial and final states during the learning procedure. Similar to the final condition of heuristic search, we further derive a constraint enforcing the final range of heuristic network output to be small. Accordingly, we propose Heuristic Domain Adaptation Network (HDAN), which explicitly learns the domain-invariant and domain-specific representations with the above mentioned constraints. Extensive experiments show that HDAN has exceeded state-of-the-art on unsupervised DA, multi-source DA and semi-supervised DA. The code is available at https: //github. com/cuishuhao/HDA.

TCS Journal 2018 Journal Article

Group mutual exclusion in linear time and space

  • Yuan He
  • K. Gopalakrishnan
  • Eli Gafni

We present two algorithms for the Group Mutual Exclusion (GME) Problem that satisfy the properties of Mutual Exclusion, Starvation Freedom, Bounded Exit, Concurrent Entry and First Come First Served. Both our algorithms use only simple read and write instructions, have O ( N ) Shared Space complexity and O ( N ) Remote Memory Reference (RMR) complexity in the Cache Coherency (CC) model. Our first algorithm is developed by generalizing the well-known Lamport's Bakery Algorithm for the classical mutual exclusion problem, while preserving its simplicity and elegance. However, it uses unbounded shared registers. Our second algorithm uses only bounded registers and is developed by generalizing Taubenfeld's Black and White Bakery Algorithm to solve the classical mutual exclusion problem using only bounded shared registers. We show that contrary to common perception our algorithms are the first to achieve these properties with this combination of complexities.

v2026.09.13