Arrow Research search

Author name cluster

Xiaodong Gu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

10 papers
1 author row

Possible papers

10

AAAI Conference 2026 Conference Paper

Anti-adversarial Learning: Desensitizing Prompts for Large Language Model

  • Xuan Li
  • Zhe Yin
  • Xiaodong Gu
  • Beijun Shen

With the widespread use of LLMs, preserving privacy in user prompts has become crucial, as prompts risk exposing private and sensitive data to cloud LLMs. Conventional techniques like homomorphic encryption (HE), secure multi-party computation, and federated learning (FL) are not well-suited to this scenario due to the lack of control over user participation in remote model interactions. In this paper, we propose PromptObfus, a novel method for desensitizing LLM prompts. The core idea of PromptObfus is "anti-adversarial" learning, which perturbs sensitive words in the prompt to obscure private information while retaining the stability of model predictions. Specifically, PromptObfus frames prompt desensitization as a masked language modeling task, replacing privacy-sensitive terms with a [MASK] token. A desensitization model is utilized to generate candidate replacements for each masked position. These candidates are subsequently selected based on gradient feedback from a surrogate model, ensuring minimal disruption to the task output. We demonstrate the effectiveness of our approach on three NLP tasks. Results show that PromptObfus effectively prevents privacy inference from remote LLMs while preserving task performance.

AAMAS Conference 2026 Conference Paper

D^3MAS: Decompose, Deduce, and Distribute for Enhanced Knowledge Sharing in Multi-Agent Systems

  • Heng Zhang
  • Yuling Shi
  • Xiaodong Gu
  • Haochen You
  • Zijian Zhang
  • Lubin Gan
  • Yilei Yuan
  • Jin Huang

Multi-agent systems powered by large language models exhibit strongcapabilitiesincollaborativeproblem-solving. However, these systems suffer from substantial knowledge redundancy. Agents duplicate efforts in retrieval and reasoning processes. This inefficiency stems from a deeper issue: current architectures lack mechanisms to ensure agents share minimal sufficient information at each operational stage. Empirical analysis reveals an average knowledge duplication rate of 47. 3% across agent communications. We propose D3MAS (Decompose, Deduce, and Distribute), a hierarchical coordination framework addressing redundancy through structural design rather than explicit optimization. The framework organizes collaboration across three coordinated layers. Task decomposition filters irrelevant sub-problems early. Collaborative reasoning captures complementary inference paths across agents. Distributed memoryprovidesaccesstonon-redundantknowledge. Theselayers coordinate through structured message passing in a unified heterogeneous graph. This cross-layer alignment ensures information remains aligned with actual task needs. Experiments on four challenging datasets show that D3MAS consistently improves reasoning accuracy by 8. 7% to 15. 6% and reduces knowledge redundancy by 46% on average.

AAMAS Conference 2026 Conference Paper

HyperAgent: Leveraging Hypergraphs for Topology Optimization in Multi-Agent Communication

  • Heng Zhang
  • Yuling Shi
  • Xiaodong Gu
  • Zijian Zhang
  • Haochen You
  • Lubin Gan
  • Yilei Yuan
  • Jin Huang

Recent advances in large language model-powered multi-agent systems have demonstrated remarkable collective intelligence through effective communication. However, existing approaches face two primary challenges: (i) Ineffective group collaboration modeling, as they rely on pairwise edge representations in graph structures, limiting their ability to capture relationships among multiple agents; and (ii) Limited task-adaptiveness in communication topology design, leading to excessive communication cost for simple tasks and insufficient coordination for complex scenarios. These issues restrict the scalability and practical deployment of adaptive collaboration frameworks. To address these challenges, we propose HyperAgent, a hypergraph-based framework that optimizes communication topologies and effectively captures group collaboration patterns using direct hyperedge representations. Unlike edge-based approaches, HyperAgent uses hyperedges to link multiple agents within the same subtask and employs hypergraph convolutional layers to achieve one-step information aggregation in collaboration groups. Additionally, it incorporates a variational autoencoder framework with sparsity regularization to dynamically adjust hypergraph topologies based on task complexity. Experiments highlight the superiority of HyperAgent in both performance and efficiency. For instance, on GSM8K, HyperAgent achieves 95. 07% accuracy while reducing token consumption by 25. 33%, demonstrating the potential of hypergraph-based optimization for multi-agent communication. ∗Corresponding author. This work is licensed under a Creative Commons Attribution International 4. 0 License. Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026), C. Amato, L. Dennis, V. Mascardi, J. Thangarajah (eds.), May 25 – 29, 2026, Paphos, Cyprus. © 2026 International Foundation for Autonomous Agents and Multiagent Systems (www. ifaamas. org). https: //doi. org/10. 65109/QTVF9552

EAAI Journal 2024 Journal Article

Co-saliency detection with two-stage co-attention mining and individual calibration

  • Zhenshan Tan
  • Xiaodong Gu
  • Qingrong Cheng

Learning-based methods have become popular in co-salient object detection (CoSOD). However, existing methods suffer from two challenging issues, including mining the inter-image co-attention and calibrating the intra-image salient objects. Moreover, the training data is insufficient. To address these challenges, we propose an end-to-end network using Two-stage Co-attention mining and Individual Calibration (TCIC) to predict the co-salient objects. Firstly, a two-stage co-attention mining architecture (TCM), including a classified co-attention module (CCM) and a focal co-attention module (FCM), is designed to model inter-image relationships. In the first stage, a CCM is applied to capture the classification interactions of multiple images, tentatively extracting the co-attention. In the second stage, we propose an FCM to adaptively suppress and aggregate multiple salient features, aiming to recalibrate the co-attention in the first stage. Secondly, considering the shape features and location information offered by the boundary features, an edge guidance module (EGM) is embedded into the individual calibration architecture (ICA) to calibrate individuals. Besides, we also adopt a co-attention transfer strategy (CTS) to keep the consistency of the co-attention during feature transfer in the decoder. Finally, TCM and ICA are integrated into a unified end-to-end framework to predict fine-grained boundary-preserving results. Besides, an image fusion algorithm (IFA) is tailored without extra pixel-level annotations for automatic generation of the composite images, aiming to supplement the training dataset. Experimental results on three prevailing testing datasets show the superiority of the proposed method in terms of various evaluation metrics.

EAAI Journal 2024 Journal Article

Enhancing accuracy, diversity, and random input compatibility in face attribute manipulation

  • Qi Guo
  • Xiaodong Gu

Recent advancements in semantic face attribute manipulation have marked significant progress, yet challenges persist regarding flexible manipulation while retaining high-accuracy reconstruction, especially given the limitations of fixed angles and layout in input facial images. To address these limitations, this paper introduces the Accurate Results, Diverse Options, and Random Input Face Attribute Manipulation Model (ADR-FACEM), a novel text-guided approach designed for nuanced and disentangled manipulation of facial attributes. This method stands out for its adaptability in attribute selection, offering a unique blend of flexibility and randomness. At the core of our proposed model lies the innovative Latent Direction Model (LDM), which leverages an adaptive nonlinear transformation trajectory. This model adeptly processes face latent codes, enabling precise manipulation of targeted attributes while preserving other facial features, all conditioned on textual descriptions. Complementing this, the Feature Distortion Alignment Model (FDAM) is intricately designed to rectify feature distortions within the image features space, thereby significantly enhancing the reconstruction quality of non-frontal images. Through comprehensive experiments, including the accuracy of facial attribute manipulation, the diversity of facial attribute manipulation options, and the inclusiveness of random unbiased input, our model ADR-FACEM demonstrates outstanding ability to maintain complex details of facial images. Quantitative comparison and qualitative analysis of nine indicators further reinforce the superiority of our method, highlighting its excellent performance in providing a wider range of choices and improvements and its compatibility with random input in the field of facial attribute manipulation.

NeurIPS Conference 2024 Conference Paper

MoGenTS: Motion Generation based on Spatial-Temporal Joint Modeling

  • Weihao Yuan
  • Yisheng He
  • Weichao Shen
  • Yuan Dong
  • Xiaodong Gu
  • Zilong Dong
  • Liefeng Bo
  • Qixing Huang

Motion generation from discrete quantization offers many advantages over continuous regression, but at the cost of inevitable approximation errors. Previous methods usually quantize the entire body pose into one code, which not only faces the difficulty in encoding all joints within one vector but also loses the spatial relationship between different joints. Differently, in this work we quantize each individual joint into one vector, which i) simplifies the quantization process as the complexity associated with a single joint is markedly lower than that of the entire pose; ii) maintains a spatial-temporal structure that preserves both the spatial relationships among joints and the temporal movement patterns; iii) yields a 2D token map, which enables the application of various 2D operations widely used in 2D images. Grounded in the 2D motion quantization, we build a spatial-temporal modeling framework, where 2D joint VQVAE, temporal-spatial 2D masking technique, and spatial-temporal 2D attention are proposed to take advantage of spatial-temporal signals among the 2D tokens. Extensive experiments demonstrate that our method significantly outperforms previous methods across different datasets, with a $26. 6\%$ decrease of FID on HumanML3D and a $29. 9\%$ decrease on KIT-ML.

NeurIPS Conference 2023 Conference Paper

GenS: Generalizable Neural Surface Reconstruction from Multi-View Images

  • Rui Peng
  • Xiaodong Gu
  • Luyang Tang
  • Shihe Shen
  • Fanqi Yu
  • Ronggang Wang

Combining the signed distance function (SDF) and differentiable volume rendering has emerged as a powerful paradigm for surface reconstruction from multi-view images without 3D supervision. However, current methods are impeded by requiring long-time per-scene optimizations and cannot generalize to new scenes. In this paper, we present GenS, an end-to-end generalizable neural surface reconstruction model. Unlike coordinate-based methods that train a separate network for each scene, we construct a generalized multi-scale volume to directly encode all scenes. Compared with existing solutions, our representation is more powerful, which can recover high-frequency details while maintaining global smoothness. Meanwhile, we introduce a multi-scale feature-metric consistency to impose the multi-view consistency in a more discriminative multi-scale feature space, which is robust to the failures of the photometric consistency. And the learnable feature can be self-enhanced to continuously improve the matching accuracy and mitigate aggregation ambiguity. Furthermore, we design a view contrast loss to force the model to be robust to those regions covered by few viewpoints through distilling the geometric prior from dense input to sparse input. Extensive experiments on popular benchmarks show that our model can generalize well to new scenes and outperform existing state-of-the-art methods even those employing ground-truth depth supervision. Code will be available at https: //github. com/prstrive/GenS.

EAAI Journal 2022 Journal Article

Dynamic-balanced double-attention fusion for image captioning

  • Changzhi Wang
  • Xiaodong Gu

Image captioning has received significant attention in the cross-modal field in which spatial and channel attentions play a crucial role. However, such attention-based approaches ignore two issues: (1) errors or noise in the channel feature map amplifies in the spatial feature map, leading to a lower model reliability; (2) image spatial feature and channel feature provide different contributions to the prediction both function words (e. g. , “in”, “out” and “on”) and notional words (e. g. , “girl”, “teddy” and “bear”). To alleviate the above issues, in this paper we propose the Dynamic-Balanced Double-Attention Fusion (DBDAF) for image captioning task that novelly exploits the attention variation and enhances the overall performance of the model. Technically, DBDAF first integrates a parallel Double Attention Network (DAN) in which channel attention is capitalized on as a supplement to the region attention, enhancing the model reliability. Then, a attention variation based Balancing Attention Fusion Mechanism (BAFM) module is devised. When predicting function words and notional words, BAFM makes a dynamic balance between channel attention and region attention based on attention variation. Moreover, to achieve the richer image description, we further devise a Doubly Stochastic Regularization (DSR) penalty and integrate it into the model loss function. Such DSR makes the model equally focus on every pixel and every channel in generating entire sentence. Extensive experiments on the three typical datasets show our DBDAF outperforms the related end-to-end leading approaches clearly. More remarkably, DBDAF achieves 1. 04% and 1. 75% improvement in terms of BLEU4 and CIDEr on the MSCOCO datasets.

AAAI Conference 2021 Conference Paper

DialogBERT: Discourse-Aware Response Generation via Learning to Recover and Rank Utterances

  • Xiaodong Gu
  • Kang Min Yoo
  • Jung-Woo Ha

Recent advances in pre-trained language models have significantly improved neural response generation. However, existing methods usually view the dialogue context as a linear sequence of tokens and learn to generate the next word through token-level self-attention. Such token-level encoding hinders the exploration of discourse-level coherence among utterances. This paper presents DialogBERT, a novel conversational response generation model that enhances previous PLM-based dialogue models. DialogBERT employs a hierarchical Transformer architecture. To efficiently capture the discourse-level coherence among utterances, we propose two training objectives, including masked utterance regression and distributed utterance order ranking in analogy to the original BERT training. Experiments on three multi-turn conversation datasets show that our approach remarkably outperforms the baselines, such as BART and DialoGPT, in terms of quantitative evaluation. The human evaluation suggests that DialogBERT generates more coherent, informative, and human-like responses than the baselines with significant margins.

IJCAI Conference 2017 Conference Paper

DeepAM: Migrate APIs with Multi-modal Sequence to Sequence Learning

  • Xiaodong Gu
  • Hongyu Zhang
  • Dongmei Zhang
  • Sunghun Kim

Computer programs written in one language are often required to be ported to other languages to support multiple devices and environments. When programs use language specific APIs (Application Programming Interfaces), it is very challenging to migrate these APIs to the corresponding APIs written in other languages. Existing approaches mine API mappings from projects that have corresponding versions in two languages. They rely on the sparse availability of bilingual projects, thus producing a limited number of API mappings. In this paper, we propose an intelligent system called DeepAM for automatically mining API mappings from a large-scale code corpus without bilingual projects. The key component of DeepAM is based on the multi-modal sequence to sequence learning architecture that aims to learn joint semantic representations of bilingual API sequences from big source code data. Experimental results indicate that DeepAM significantly increases the accuracy of API mappings as well as the number of API mappings when compared with the state-of-the-art approaches.

v2026.09.13