Arrow Research search

Author name cluster

Shuang Yang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

16 papers
2 author rows

Possible papers

16

TMLR Journal 2026 Journal Article

UniRec: Unified Multimodal Encoding for LLM-Based Recommendations

  • Zijie Lei
  • Tao Feng
  • Zhigang Hua
  • Yan Xie
  • Guanyu Lin
  • Shuang Yang
  • Ge Liu
  • Jiaxuan You

Large language models (LLMs) have recently shown promise for multimodal recommen- dation, particularly with text and image inputs. Yet real-world recommendation signals extend far beyond these modalities. To reflect this, we formalize recommendation features into four modalities: text, images, categorical features, and numerical attributes, and em- phasize unique challenges this heterogeneity poses for LLMs in understanding multimodal information. In particular, these challenges arise not only across modalities but also within them, as attributes (e.g., price, rating, time) may all be numeric yet carry distinct meanings. Beyond this intra-modality ambiguity, another major challenge is the nested structure of recommendation signals, where user histories are sequences of items, each carrying multiple attributes. To address these challenges, we propose UniRec, a unified multimodal encoder for LLM-based recommendation. UniRec first employs modality-specific encoders to produce consistent embeddings across heterogeneous signals. It then applies a triplet representa- tion—comprising attribute name, type, and value—to separate schema from raw inputs and preserve semantic distinctions. Finally, a hierarchical Q-Former models the nested structure of user interactions while maintaining their layered organization. On multiple real-world benchmarks, UniRec outperforms state-of-the-art multimodal and LLM-based recommenders by up to 15%, while extensive ablation studies further validate the contribu- tions of each component.

EAAI Journal 2025 Journal Article

An effective convolutional and transformer cooperation network for underwater acoustic target recognition

  • Anqi Jin
  • Shuang Yang
  • Menghui Lei
  • Xiangyang Zeng
  • Haitao Wang

Underwater acoustic target recognition (UATR) is a key technology in the field of underwater acoustic information processing. In recent years, models based on convolutional neural networks (CNN) have shown excellent performance in the domain of UATR. However, CNN have limitations in capturing the global information of underwater acoustic features. Due to its advantages in modeling global dependencies, the Transformer model is gradually gaining attention from researchers. In order to capture the time-frequency dependencies in acoustic spectrograms more effectively, this paper proposes a recognition model based on the Mel spectrogram that combines CNN with the Transformer, named the underwater acoustics CNN-Transformer cooperation network (UACTC). Compared to the Transformer alone, this model is more efficient in extracting local features. The CNN module uses a residual network based on the efficient channel attention (ECA) module for efficient deep feature extraction. Additionally, the ECA module is introduced into the Transformer block to enhance the channel feature extraction of the Transformer. Experiments prove that the ECA module effectively improves the performance of the recognition system. The effectiveness of the proposed model has been validated on two public datasets, achieving 98. 05 % and 96. 96 % on the ShipsEar and DeepShip datasets, respectively.

NeurIPS Conference 2025 Conference Paper

AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration

  • Andy Zhou
  • Kevin Wu
  • Francesco Pinto
  • Zhaorun Chen
  • Yi Zeng
  • Yu Yang
  • Shuang Yang
  • Sanmi Koyejo

As large language models (LLMs) become increasingly capable, security and safety evaluation are crucial. While current red teaming approaches have made strides in assessing LLM vulnerabilities, they often rely heavily on human input and lack comprehensive coverage of emerging attack vectors. This paper introduces AutoRedTeamer, a novel framework for fully automated, end-to-end red teaming against LLMs. AutoRedTeamer combines a multi-agent architecture with a memory-guided attack selection mechanism to enable continuous discovery and integration of new attack vectors. The dual-agent framework consists of a red teaming agent that can operate from high-level risk categories alone to generate and execute test cases, and a strategy proposer agent that autonomously discovers and implements new attacks by analyzing recent research. This modular design allows AutoRedTeamer to adapt to emerging threats while maintaining strong performance on existing attack vectors. We demonstrate AutoRedTeamer’s effectiveness across diverse evaluation settings, achieving 20% higher attack success rates on HarmBench against Llama-3. 1-70B while reducing computational costs by 46% compared to existing approaches. AutoRedTeamer also matches the diversity of human-curated benchmarks in generating test cases, providing a comprehensive, scalable, and continuously evolving framework for evaluating the security of AI systems.

EAAI Journal 2025 Journal Article

Partial convolution-simple attention mechanism-SegFormer: An accurate and robust model for landslide identification

  • Shuang Yang
  • Yuzhu Wang
  • Kai Zhao
  • Xiaocai Liu
  • Jingqin Mu
  • Xupeng Zhao

To achieve accurate and robust landslide identification, this study presents an advanced model based on the SegFormer architecture, named Partial Convolution-Simple Attention Mechanism-SegFormer (PConv-simAM-SegFormer). The model integrates partial convolution (PConv) to enhance spatial feature extraction capabilities and employs a simple attention mechanism (simAM), inspired by neuroscience, to optimize attention allocation. Additionally, a novel reverse transfer learning strategy is introduced to leverage features from older landslides, thereby improving the detection of new landslides. Experimental results on three large-scale publicly available datasets validate the modelś effectiveness: a mean Intersection over Union (mIoU) of 76. 05% on the Landslide Research for Sichuan–Tibet Transportation Corridor dataset (LRSTTC), representing a 3. 04% improvement over the baseline; an mIoU of 90. 35% on the Bijie Landslide dataset (Bijie), with a 2. 88% enhancement; and an mIoU of 97. 34% on the Tibetan Plateau Lakes dataset (TPL), with a 0. 81% enhancement. These results demonstrate the modelś high accuracy and robustness across diverse datasets and application scenarios, showcasing its significant potential for practical applications in the field of landslide detection.

ICML Conference 2025 Conference Paper

UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning

  • Jiawei Zhang 0002
  • Shuang Yang
  • Bo Li 0026

Large Language Model (LLM) agents equipped with external tools have become increasingly powerful for complex tasks such as web shopping, automated email replies, and financial trading. However, these advancements amplify the risks of adversarial attacks, especially when agents can access sensitive external functionalities. Nevertheless, manipulating LLM agents into performing targeted malicious actions or invoking specific tools remains challenging, as these agents extensively reason or plan before executing final actions. In this work, we present UDora, a unified red teaming framework designed for LLM agents that dynamically hijacks the agent’s reasoning processes to compel malicious behavior. Specifically, UDora first generates the model’s reasoning trace for the given task, then automatically identifies optimal points within this trace to insert targeted perturbations. The resulting perturbed reasoning is then used as a surrogate response for optimization. By iteratively applying this process, the LLM agent will then be induced to undertake designated malicious actions or to invoke specific malicious tools. Our approach demonstrates superior effectiveness compared to existing methods across three LLM agent datasets. The code is available at https: //github. com/AI-secure/UDora.

EAAI Journal 2024 Journal Article

Spatio-temporal features for fast early warning of unplanned self-extubation in ICU

  • Yang Chen
  • Ling Wang
  • Guorong Wang
  • Shuang Yang
  • Yingying Wang
  • MingFang Xiang
  • Xuan Zhang
  • Hui Chen

Patients’ behaviors in the Intensive Care Units (ICU) have garnered research attention, particularly regarding the impact of Unplanned Extubation (UEX). However, there is currently no existing report on methods for early warning of UEX action in RGB video. Applying traditional human action recognition algorithms to UEX in the complex ICU environment proves challenging. To address the above issue, we propose a novel feature for early warning of UEX action in patients using RGB videos. Firstly, we employ the YOLOv3 detection method to extract the region of interest (ROI), which corresponds to the region where the patient is located. Subsequently, we develop a spatio-temporal (ST) feature for human action tracking by using the L-K optical flow algorithm. This ST feature encompasses optical flow corner number, trajectory distance, and wavelet transform features. Finally, we utilize support vector machine (SVM) for patient action classification and early warning. Experimental results on the ICU monitoring dataset demonstrate the superior performance of the proposed feature in UEX prediction.

EAAI Journal 2024 Journal Article

Underwater acoustic target recognition based on sub-band concatenated Mel spectrogram and multidomain attention mechanism

  • Shuang Yang
  • Anqi Jin
  • Xiangyang Zeng
  • Haitao Wang
  • Xi Hong
  • Menghui Lei

Underwater acoustic target recognition is extremely challenging because of the pronounced background noise and intricate sound propagation patterns inherent to maritime environments. Herein, we propose a sub-band concatenated Mel spectrogram to amplify low-frequency ship-radiated noise. This method enhances features through multispectrogram concatenation. Furthermore, we introduce a multidomain attention mechanism to enhance the performance of a simple residual network to develop a lightweight CFTANet model. The recognition accuracies of the recognition system are 90. 60% and 96. 40% on two open datasets. On the DeepShip dataset, the recognition accuracy is 7. 06% higher than those of previous state-of-the-art methods.

EAAI Journal 2023 Journal Article

A K-Net-based hybrid semantic segmentation method for extracting lake water bodies

  • Cong Chen
  • Yuzhu Wang
  • Shuang Yang
  • Xiaohui Ji
  • Gongwen Wang

Lakes have a crucial impact on natural disaster prevention, resource recycling, maintenance of agricultural production and daily life. The traditional way of acquiring lake water body information lacks efficiency, is dangerous, and is not suitable for lake water body information acquisition and real-time monitoring. For this reason, the automated lake water body extraction method based on deep learning semantic segmentation model is gradually becoming a mainstream method. However, most of the semantic segmentation models used for lake extraction today express features through static semantics, while ignoring the extraction relationships of different convolutional kernels for these features. In order to better extract lake water bodies from remote sensing images, this paper proposes a hybrid semantic segmentation method based on K-Net, which achieves high accuracy extraction of lake water bodies by introducing dynamic semantic kernels to iteratively refine the feature information. The superiority of the K-Net-based hybrid model on a Google remote sensing image dataset of lakes is validated. The experimental results show that (1) the hybrid model is able to achieve accurate extraction of lake water bodies, with the UperNet + K-Net model using Swin-l performing the best among all six evaluation metrics, with mean intersection over union (mIoU) reaching 97. 77%; and that (2) after incorporating the K-Net module, all tested models obtain a larger mIoU than before.

YNIMG Journal 2023 Journal Article

Deep learning-assisted identification and quantification of aneurysmal subarachnoid hemorrhage in non-contrast CT scans: Development and external validation of Hybrid 2D/3D UNet

  • Ping Hu
  • Haizhu Zhou
  • Tengfeng Yan
  • Hongping Miu
  • Feng Xiao
  • Xinyi Zhu
  • Lei Shu
  • Shuang Yang

Accurate stroke assessment and consequent favorable clinical outcomes rely on the early identification and quantification of aneurysmal subarachnoid hemorrhage (aSAH) in non-contrast computed tomography (NCCT) images. However, hemorrhagic lesions can be complex and difficult to distinguish manually. To solve these problems, here we propose a novel Hybrid 2D/3D UNet deep-learning framework for automatic aSAH identification and quantification in NCCT images. We evaluated 1824 consecutive patients admitted with aSAH to four hospitals in China between June 2018 and May 2022. Accuracy and precision, Dice scores and intersection over union (IoU), and interclass correlation coefficients (ICC) were calculated to assess model performance, segmentation performance, and correlations between automatic and manual segmentation, respectively. A total of 1355 patients with aSAH were enrolled: 931, 101, 179, and 144 in four datasets, of whom 326 were scanned with Siemens, 640 with Philips, and 389 with GE Medical Systems scanners. Our proposed deep-learning method accurately identified (accuracies 0.993-0.999) and segmented (Dice scores 0.550-0.897) hemorrhage in both the internal and external datasets, even combinations of hemorrhage subtypes. We further developed a convenient AI-assisted platform based on our algorithm to assist clinical workflows, whose performance was comparable to manual measurements by experienced neurosurgeons (ICCs 0.815-0.957) but with greater efficiency and reduced cost. While this tool has not yet been prospectively tested in clinical practice, our innovative hybrid network algorithm and platform can accurately identify and quantify aSAH, paving the way for fast and cheap NCCT interpretation and a reliable AI-based approach to expedite clinical decision-making for aSAH patients.

AAMAS Conference 2022 Conference Paper

Characterizing Attacks on Deep Reinforcement Learning

  • Xinlei Pan
  • Chaowei Xiao
  • Warren He
  • Shuang Yang
  • Jian Peng
  • Mingjie Sun
  • Mingyan Liu
  • Bo Li

Recent studies show that Deep Reinforcement Learning (DRL) models are vulnerable to adversarial attacks, which attack DRL models by adding small perturbations to the observations. However, some attacks assume full availability of the victim model, and some require a huge amount of computation, making them less feasible for real world applications. In this work, we make further explorations of the vulnerabilities of DRL by studying other aspects of attacks on DRL using realistic and e�cient attacks. First, we adapt and propose e�cient black-box attacks when we do not have access to DRL model parameters. Second, to address the high computational demands of existing attacks, we introduce e�cient online sequential attacks that exploit temporal consistency across consecutive steps. Third, we explore the possibility of an attacker perturbing other aspects in the DRL setting, such as the environment dynamics. Finally, to account for imperfections in how an attacker would inject perturbations in the physical world, we devise a method for generating a robust physical perturbations to be printed. The attack is evaluated on a real-world robot under various conditions. We conduct extensive experiments both in simulation such as Atari games, robotics and autonomous driving, and on real-world robotics, to compare the e�ectiveness of the proposed attacks with baseline approaches. To the best of our knowledge, we are the�rst to apply adversarial attacks on DRL systems to physical robots.

NeurIPS Conference 2021 Conference Paper

A Bi-Level Framework for Learning to Solve Combinatorial Optimization on Graphs

  • Runzhong Wang
  • Zhigang Hua
  • Gan Liu
  • Jiayi Zhang
  • Junchi Yan
  • Feng Qi
  • Shuang Yang
  • Jun Zhou

Combinatorial Optimization (CO) has been a long-standing challenging research topic featured by its NP-hard nature. Traditionally such problems are approximately solved with heuristic algorithms which are usually fast but may sacrifice the solution quality. Currently, machine learning for combinatorial optimization (MLCO) has become a trending research topic, but most existing MLCO methods treat CO as a single-level optimization by directly learning the end-to-end solutions, which are hard to scale up and mostly limited by the capacity of ML models given the high complexity of CO. In this paper, we propose a hybrid approach to combine the best of the two worlds, in which a bi-level framework is developed with an upper-level learning method to optimize the graph (e. g. add, delete or modify edges in a graph), fused with a lower-level heuristic algorithm solving on the optimized graph. Such a bi-level approach simplifies the learning on the original hard CO and can effectively mitigate the demand for model capacity. The experiments and results on several popular CO problems like Directed Acyclic Graph scheduling, Graph Edit Distance and Hamiltonian Cycle Problem show its effectiveness over manually designed heuristics and single-level learning methods.

ICML Conference 2021 Conference Paper

Progressive-Scale Boundary Blackbox Attack via Projective Gradient Estimation

  • Jiawei Zhang 0013
  • Linyi Li 0001
  • Huichen Li
  • Xiaolu Zhang
  • Shuang Yang
  • Bo Li 0026

Boundary based blackbox attack has been recognized as practical and effective, given that an attacker only needs to access the final model prediction. However, the query efficiency of it is in general high especially for high dimensional image data. In this paper, we show that such efficiency highly depends on the scale at which the attack is applied, and attacking at the optimal scale significantly improves the efficiency. In particular, we propose a theoretical framework to analyze and show three key characteristics to improve the query efficiency. We prove that there exists an optimal scale for projective gradient estimation. Our framework also explains the satisfactory performance achieved by existing boundary black-box attacks. Based on our theoretical framework, we propose Progressive-Scale enabled projective Boundary Attack (PSBA) to improve the query efficiency via progressive scaling techniques. In particular, we employ Progressive-GAN to optimize the scale of projections, which we call PSBA-PGAN. We evaluate our approach on both spatial and frequency scales. Extensive experiments on MNIST, CIFAR-10, CelebA, and ImageNet against different models including a real-world face recognition API show that PSBA-PGAN significantly outperforms existing baseline attacks in terms of query efficiency and attack success rate. We also observe relatively stable optimal scales for different models and datasets. The code is publicly available at https: //github. com/AI-secure/PSBA.

ICLR Conference 2020 Conference Paper

A Learning-based Iterative Method for Solving Vehicle Routing Problems

  • Hao Lu
  • Xingwen Zhang
  • Shuang Yang

This paper is concerned with solving combinatorial optimization problems, in particular, the capacitated vehicle routing problems (CVRP). Classical Operations Research (OR) algorithms such as LKH3 \citep{helsgaun2017extension} are inefficient and difficult to scale to larger-size problems. Machine learning based approaches have recently shown to be promising, partly because of their efficiency (once trained, they can perform solving within minutes or even seconds). However, there is still a considerable gap between the quality of a machine learned solution and what OR methods can offer (e.g., on CVRP-100, the best result of learned solutions is between 16.10-16.80, significantly worse than LKH3's 15.65). In this paper, we present ``Learn to Improve'' (L2I), the first learning based approach for CVRP that is efficient in solving speed and at the same time outperforms OR methods. Starting with a random initial solution, L2I learns to iteratively refine the solution with an improvement operator, selected by a reinforcement learning based controller. The improvement operator is selected from a pool of powerful operators that are customized for routing problems. By combining the strengths of the two worlds, our approach achieves the new state-of-the-art results on CVRP, e.g., an average cost of 15.57 on CVRP-100.

NeurIPS Conference 2020 Conference Paper

Bandit Samplers for Training Graph Neural Networks

  • Ziqi Liu
  • Zhengwei Wu
  • Zhiqiang Zhang
  • Jun Zhou
  • Shuang Yang
  • Le Song
  • Yuan Qi

Several sampling algorithms with variance reduction have been proposed for accelerating the training of Graph Convolution Networks (GCNs). However, due to the intractable computation of optimal sampling distribution, these sampling algorithms are suboptimal for GCNs and are not applicable to more general graph neural networks (GNNs) where the message aggregator contains learned weights rather than fixed weights, such as Graph Attention Networks (GAT). The fundamental reason is that the embeddings of the neighbors or learned weights involved in the optimal sampling distribution are \emph{changing} during the training and \emph{not known a priori}, but only \emph{partially observed} when sampled, thus making the derivation of an optimal variance reduced samplers non-trivial. In this paper, we formulate the optimization of the sampling variance as an adversary bandit problem, where the rewards are related to the node embeddings and learned weights, and can vary constantly. Thus a good sampler needs to acquire variance information about more neighbors (exploration) while at the same time optimizing the immediate sampling variance (exploit). We theoretically show that our algorithm asymptotically approaches the optimal variance within a factor of 3. We show the efficiency and effectiveness of our approach on multiple datasets.

ECAI Conference 2020 Conference Paper

Generating Natural Language Adversarial Examples on a Large Scale with Generative Models

  • Yankun Ren
  • Jianbin Lin
  • Siliang Tang
  • Jun Zhou 0011
  • Shuang Yang
  • Yuan (Alan) Qi
  • Xiang Ren 0001

Today text classification models have been widely used. However, these classifiers are found to be easily fooled by adversarial examples. Fortunately, standard attacking methods generate adversarial texts in a pair-wise way, that is, an adversarial text can only be created from a real-world text by replacing a few words. In many applications, these texts are limited in numbers, therefore their corresponding adversarial examples are often not diverse enough and sometimes hard to read, thus can be easily detected by humans and cannot create chaos at a large scale. In this paper, we propose an end to end solution to efficiently generate adversarial texts from scratch using generative models, which are not restricted to perturbing the given texts. We call it unrestricted adversarial text generation. Specifically, we train a conditional variational autoencoder (VAE) with an additional adversarial loss to guide the generation of adversarial examples. Moreover, to improve the validity of adversarial texts, we utilize discrimators and the training framework of generative adversarial networks (GANs) to make adversarial texts consistent with real data. Experimental results on sentiment analysis demonstrate the scalability and efficiency of our method. It can attack text classification models with a higher success rate than existing methods, and provide acceptable quality for humans in the meantime.

IJCAI Conference 2020 Conference Paper

Learning for Graph Matching and Related Combinatorial Optimization Problems

  • Junchi Yan
  • Shuang Yang
  • Edwin Hancock

This survey gives a selective review of recent development of machine learning (ML) for combinatorial optimization (CO), especially for graph matching. The synergy of these two well-developed areas (ML and CO) can potentially give transformative change to artificial intelligence, whose foundation relates to these two building blocks. For its representativeness and wide-applicability, this paper is more focused on the problem of weighted graph matching, especially from the learning perspective. For graph matching, we show that many learning techniques e. g. convolutional neural networks, graph neural networks, reinforcement learning can be effectively incorporated in the paradigm for extracting the node features, graph structure features, and even the matching engine. We further present outlook for the new settings for learning graph matching, and direction towards more integrated combinatorial optimization solvers with prediction models, and also the mutual embrace of traditional solver and machine learning components.

v2026.09.13