Arrow Research search

Author name cluster

Xiaojun Wu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
2 author rows

Possible papers

11

AAAI Conference 2026 Conference Paper

BAG: Benchmarking Anomaly Detection on Dynamic Graphs

  • Fengrui Hua
  • Yiyan Qi
  • Zikai Wei
  • Yuxing Tian
  • Chengjin Xu
  • Xiaojun Wu
  • Jia Li
  • Jian Guo

Anomaly detection in dynamic graphs is a critical area of research that focuses on identifying abnormal components within evolving graph structures that deviate significantly from typical patterns. Despite advancements in traditional temporal pattern mining and deep learning techniques, a comprehensive benchmarking framework for Dynamic Graph Anomaly Detection (DyGAD) has been lacking. To address this gap, we introduce BAG, the first comprehensive benchmark specifically designed for anomaly detection on dynamic graphs. BAG enables extensive evaluation of 25 leading DyGAD models, covering both classical approaches and advanced Dynamic Graph Neural Networks (DGNNs), across 10 diverse real-world datasets that include both synthetic and naturally occurring anomalies. The framework supports evaluations at both the edge and node levels, offering a robust tool to advance DyGAD research. Our main finding is that Continuous-time Dynamic Graph (CTDG) models demonstrate superior performance and potential in detecting anomalies in dynamic graph edges, compared to Discrete-time Dynamic Graph (DTDG) models. Furthermore, the results reveal that existing methods are less effective at detecting organic anomalies, primarily due to the presence of temporal anomalies and highly imbalanced samples. The proposed BAG benchmark significantly enhances the evaluation of DyGAD methods by improving dataset selection, metric application, and model training. Moreover, BAG supports reproducibility and further exploration in this field by integrating all models, datasets, and evaluation protocols into an open-source repository.

EAAI Journal 2026 Journal Article

Process-informed encoding and evaluation approach for scoring figure skating videos

  • Zexing Du
  • Xiaojun Wu
  • Shigang Liu

This paper exploits a process-informed encoding and evaluation method for scoring figure skating videos. Most existing works typically rely on holistic video representations for score regression or ranking, without considering the detailed semantic and motion dynamics in these long-term videos, which usually contain a set of sub-actions that need to be scored. To address these problems, we propose a motion-informed temporal parsing module (MITPM), which extracts segment-wise motion-salience features for exploring detailed information in videos. In this way, our MITPM converts entire videos to some important segments. Then, following the practical figure skating scoring process, we propose a segment-wise score decoupling (SSD) strategy, which quantifies the grades of each extracted segment explicitly. Specifically, we define a set of score prototypes to quantitatively evaluate the performance grades of each segment. Instead of performing direct regression on the whole video, our method generates the final score by combining the quantitative grades of each segments. By evaluating widely used figure skating datasets, our method achieves state-of-the-art performance. Moreover, thanks to the efficiency of our method, we achieve the best trade-off between accuracy and model complexity.

EAAI Journal 2025 Journal Article

A theme music generation model based on hybrid variational autoencoders and conditional generative adversarial networks

  • Fangzhu Jin
  • Peng Li
  • Xiaojun Wu

The musical theme is the core melodic element that runs through a work. It is the main source of recognizable features and the basis of the overall structure of the score. However, existing music generation systems often face problems with missing thematic features and monotonous structure. These systems find it difficult to capture the complex thematic changes and long-term dependencies in music, resulting in a lack of dynamic changes, artistic expression in phrasing, and structural coherence in the generated results. In this paper, we propose a thematic music generation model based on a generative adversarial network, named Thematic Music Conditional Generative Adversarial Network (TM-CGAN), which solves these challenges through three key innovations. First, we propose a thematic fusion degree calculation module based on multi-dimensional feature metrics to overcome the theoretical limitations of existing methods, which lack explicit modeling of thematic features. Subsequently, we construct a two-dimensional image representation learning framework that preserves Musical Instrument Digital Interface (MIDI) symbolic sequences and employs convolutional neural networks (CNNs) to effectively model the multi-dimensional feature dependencies inherent in musical structures. Finally, we design a hybrid variational autoencoder-conditional generative adversarial network architecture that processes latent space modeling of thematic features through variational inference mechanisms while simultaneously optimizing generation quality through adversarial training. We conducted comprehensive experiments across three diverse MIDI music datasets to validate our approach. The experimental results demonstrate that TM-CGAN significantly outperforms state-of-the-art baseline models on multiple evaluation metrics, including generated music quality, theme representativeness, and structural integrity.

NeurIPS Conference 2025 Conference Paper

Collaborating Vision, Depth, and Thermal Signals for Multi-Modal Tracking: Dataset and Algorithm

  • Xue-Feng Zhu
  • Tianyang Xu
  • Yifan Pan
  • Jinjie Gu
  • Xi Li
  • Jiwen Lu
  • Xiaojun Wu
  • Josef Kittler

Existing multi-modal object tracking approaches primarily focus on dual-modal paradigms, such as RGB-Depth or RGB-Thermal, yet remain challenged in complex scenarios due to limited input modalities. To address this gap, this work introduces a novel multi-modal tracking task that leverages three complementary modalities, including visible RGB, Depth (D), and Thermal Infrared (TIR), aiming to enhance robustness in complex scenarios. To support this task, we construct a new multi-modal tracking dataset, coined RGBDT500, which consists of 500 videos with synchronised frames across the three modalities. Each frame provides spatially aligned RGB, depth, and thermal infrared images with precise object bounding box annotations. Furthermore, we propose a novel multi-modal tracker, dubbed RDTTrack. RDTTrack integrates tri-modal information for robust tracking by leveraging a pretrained RGB-only tracking model and prompt learning techniques. In specific, RDTTrack fuses thermal infrared and depth modalities under a proposed orthogonal projection constraint, then integrates them with RGB signals as prompts for the pre-trained foundation tracking model, effectively harmonising tri-modal complementary cues. The experimental results demonstrate the effectiveness and advantages of the proposed method, showing significant improvements over existing dual-modal approaches in terms of tracking accuracy and robustness in complex scenarios. The dataset and source code are publicly available at https: //xuefeng-zhu5. github. io/RGBDT500.

AAAI Conference 2025 Conference Paper

One-Shot Reference-based Structure-Aware Image to Sketch Synthesis

  • Rui Yang
  • Honghong Yang
  • Li Zhao
  • Qin Lei
  • Mianxiong Dong
  • Kaoru Ota
  • Xiaojun Wu

Generating sketches that accurately reflect the content of reference images presents numerous challenges. Current methods either require paired training data or fail to accommodate a wider range and diversity of sketch styles. While pre-trained diffusion models have shown strong text-based control capabilities for reference-based content sketch generation, state-of-the-art methods still struggle with reference-based sketch generation for given content. The main difficulties lie in (1) balancing content preservation with style enhancement, and (2) representing content image textures at varying levels of abstraction to approximate the reference sketch style. In this paper, we propose a method (Ref2Sketch-SA) that transforms a given content image into a sketch based on a reference sketch. The core strategies include (1) using DDIM Inversion to enhance structural consistency in the sketch generation of content images; (2) injecting noise into the input image during the denoising process to produce a sketch that retains content attributes while aligning with, yet differing in texture from, the reference. Our model demonstrates superior performance across multiple evaluation metrics, including user style preference.

AAAI Conference 2025 Conference Paper

R-DTI: Drug Target Interaction Prediction Based on Second-Order Relevance Exploration

  • Yang Hua
  • Tianyang Xu
  • Xiaoning Song
  • Zhenhua Feng
  • Rui Wang
  • Wenjie Zhang
  • Xiaojun Wu

Drug Target Interaction (DTI) prediction has witnessed promising performance boosts accompanied by advanced multimodal feature extraction. However, existing approaches suffer from two main difficulties. First, the complex protein structures cannot be well represented by current protein-sequence-based feature extractors. Second, the gap between protein and drug features increases the vulnerability of the obtained classifier thus degrading the prediction robustness. To address these issues, we propose a novel R-DTI method by exploring the second-order relevance in both protein structural feature extraction and DTI prediction phases. Specifically, we construct a pre-trained structural feature extractor that mines the atomic relevance of each amino acid. Then, an inter-feature structure-preserved Riemannian network is designed to expand the existing protein extraction patterns. To improve the prediction robustness, we also develop a Riemannian classifier that uses the second-order protein-drug relevance with a unified feature space. Extensive experimental results demonstrate the merits and superiority of our R-DTI against the state-of-the-art, achieving 1.4% and 1.9% higher AUC-ROC on the BindingDB and DrugBank datasets, respectively.

EAAI Journal 2025 Journal Article

Representing-behavior-based online driving style recognition and its road verification

  • Bingxian Li
  • Yanfang Liu
  • Junwei Zhao
  • Xiangyang Xu
  • Peng Dong
  • Songbo Chen
  • Xiaojun Wu
  • Xuewu Liu

—Driving style recognition (DSR) is a critical task for intelligent vehicles, providing a means to analyze drivers’ behavioral characteristics and understand their preferences. In this study, a representing-behavior-based DSR method is proposed. It is not only capable of distinguishing various driving styles but also adept at depicting their dynamic changes during a trip. Initially, three driving behavior datasets are extracted from natural driving data. Subsequently, a clustering metric called driving style distinction is introduced to assist the k-means algorithm in searching for driving style representing behaviors (DSRBs) within these datasets. Considering the impact of input features on clustering, a genetic-algorithm-based feature selection method is implemented to optimize the DSRB search. Upon completing this task, DSRB recognition models are built through supervised learning, enabling the calculation of a novel metric called the driving style index, which describes driving styles in real time. A road test demonstrates the feasibility of the proposed method, establishing it as an effective solution for DSR problems.

NeurIPS Conference 2025 Conference Paper

Towards a General Attention Framework on Gyrovector Spaces for Matrix Manifolds

  • Rui Wang
  • Chen Hu
  • Xiaoning Song
  • Xiaojun Wu
  • Nicu Sebe
  • Ziheng Chen

Deep neural networks operating on non-Euclidean geometries have recently demonstrated impressive performance across various machine-learning applications. Several studies have extended the attention mechanism to different manifolds. However, most existing non-Euclidean attention models are tailored to specific geometries, limiting their applicability. On the other hand, recent studies show that several matrix manifolds, such as Symmetric Positive Definite (SPD), Symmetric Positive Semi-Definite (SPSD), and Grassmannian manifolds, admit gyrovector structures, which extend vector addition and scalar product into manifolds. Leveraging these properties, we propose a Gyro Attention (GyroAtt) framework over general gyrovector spaces, applicable to various matrix geometries. Empirically, we manifest GyroAtt on three gyro structures on the SPD manifold, three on the SPSD manifold, and one on the Grassmannian manifold. Extensive experiments on four electroencephalography (EEG) datasets demonstrate the effectiveness of our framework.

AAAI Conference 2024 Conference Paper

Generative-Based Fusion Mechanism for Multi-Modal Tracking

  • Zhangyong Tang
  • Tianyang Xu
  • Xiaojun Wu
  • Xue-Feng Zhu
  • Josef Kittler

Generative models (GMs) have received increasing research interest for their remarkable capacity to achieve comprehensive understanding. However, their potential application in the domain of multi-modal tracking has remained unexplored. In this context, we seek to uncover the potential of harnessing generative techniques to address the critical challenge, information fusion, in multi-modal tracking. In this paper, we delve into two prominent GM techniques, namely, Conditional Generative Adversarial Networks (CGANs) and Diffusion Models (DMs). Different from the standard fusion process where the features from each modality are directly fed into the fusion block, we combine these multi-modal features with random noise in the GM framework, effectively transforming the original training samples into harder instances. This design excels at extracting discriminative clues from the features, enhancing the ultimate tracking performance. Based on this, we conduct extensive experiments across two multi-modal tracking tasks, three baseline methods, and four challenging benchmarks. The experimental results demonstrate that the proposed generative-based fusion mechanism achieves state-of-the-art performance by setting new records on GTOT, LasHeR and RGBD1K. Code will be available at https://github.com/Zhangyong-Tang/GMMT.

EAAI Journal 2011 Journal Article

QoS multicast routing using a quantum-behaved particle swarm optimization algorithm

  • Jun Sun
  • Wei Fang
  • Xiaojun Wu
  • Zhenping Xie
  • Wenbo Xu

QoS multicast routing in networks is a very important research issue in networks and distributed systems. It is also a challenging and hard problem for high-performance networks of the next generation. Due to its NP-completeness, many heuristic methods have been employed to solve the problem. This paper proposes the modified quantum-behaved particle swarm optimization (QPSO) method for QoS multicast routing. In the proposed method, QoS multicast routing is converted into an integer programming problem with QoS constraints and is solved by the QPSO algorithm combined with loop deletion operation. The QPSO-based routing method, along with the routing algorithms based on particle swarm optimization (PSO) and genetic algorithm (GA), is tested on randomly generated network topologies for the purpose of performance evaluation. The simulation results show the efficiency of the proposed method on QoS the routing problem and its superiority to the methods based on PSO and GA.

ICRA Conference 2010 Conference Paper

DAvinCi: A cloud computing framework for service robots

  • Rajesh Arumugam
  • Vikas Reddy Enti
  • Bingbing Liu
  • Xiaojun Wu
  • Krishnamoorthy Baskaran
  • Foo Kong Foong
  • Appadorai Senthil Kumar
  • Dee Meng Kang

We propose DAvinCi, a software framework that provides the scalability and parallelism advantages of cloud computing for service robots in large environments. We have implemented such a system around the Hadoop cluster with ROS (Robotic Operating system) as the messaging framework for our robotic ecosystem. We explore the possibilities of parallelizing some of the robotics algorithms as Map/Reduce tasks in Hadoop. We implemented the FastSLAM algorithm in Map/Reduce and show how significant performance gains in execution times to build a map of a large area can be achieved with even a very small eight-node Hadoop cluster. The global map can later be shared with other robots introduced in the environment via a Software as a Service (SaaS) Model. This reduces the burden of exploration and map building for the new robot and minimizes it's need for additional sensors. Our primary goal is to develop a cloud computing environment which provides a compute cluster built with commodity hardware exposing a suite of robotic algorithms as a SaaS and share data co-operatively across the robotic ecosystem.

v2026.09.13