Arrow Research search

Author name cluster

Xiaoyang Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

20 papers
2 author rows

Possible papers

20

AAAI Conference 2026 Conference Paper

Spectrally Adaptive Channel-aware Unrolling Network for Compressed Sensing

  • Xiaoyang Wang
  • Hongping Gan

Deep Unrolling Networks (DUNs) integrate classical optimization recovery problems in Compressed Sensing (CS) with sophisticated deep learning network architectures, leading to substantial breakthroughs. However, prevailing DUNs generally face challenges concerning solidified gradient descent step size strategies, inadequate feature extraction within the iterative stage and limited information interaction between iterative stages. To overcome these obstacles, we propose SCU-Net, a channel-focused unrolling network inspired by the renowned spectral projected gradient optimization algorithm. In particular, we tailore two pivotal components, Barzilai-Borwein-gradient Descent Optimizer (BBDO) and Channel-guided Cross-attention Reconstruction Module (CCRM), to collaboratively undertake the reconstruction task. BBDO leverages a gradient calculation strategy based on BB step size to enhance data fidelity optimization, while CCRM addresses the intricate mapping issue associated with sparse induction, encompassing customized functionalities from Adaptive Channel Interaction Layer (ACIL) and Spatially Augmented Channel-aware Unit (SACU). Among them, ACIL amalgamates convolution operations and channel attention mechanisms to achieve meticulous information screening alongside efficient feature enhancement. SACU introduces dual reinforcement variables to bolster information exchange across different iterative stages, coupled with the optimization of cross-attention to facilitate the modeling of long-distance dependencies. Extensive experiments in both image CS and magnetic resonance imaging exhibit that our SCU-Net manifests superior performance, surpassing state-of-the-art methods.

EAAI Journal 2025 Journal Article

An aspect-level sentiment Graph Convolutional Network model based on transformer and frequency domain

  • Xiaoyang Wang
  • Wenfeng Liu

Aspect-based sentiment analysis (ABSA) has long been a crucial research direction in Natural Language Processing (NLP). Its significance lies in unraveling finer-grained sentiments, establishing relationships among different emotional elements, and predicting and analyzing sentiments for specific aspects within text sequences. In recent years, the use of syntactic dependencies and dependency trees to construct neural networks for sentiment analysis tasks has proven to be effective. Simultaneously, using pre-trained large models has significantly improved performance across various text sequence tasks. However, effectively leveraging syntactic dependency information and modeling relevant information in text sequences pose higher demands on the capabilities of network models. Reducing noise generated during this process and enhancing prediction accuracy are crucial challenges. In this paper, we propose an ABSA network model, called Bidirectional Encoder Representations from Transformers-Spectral Graph Convolutional Network (BSGCN), which processes text sequences using Bidirectional Encoder Representations from Transformers (BERT). Incorporating relevant information through a combination of self-attention and aspect-attention, the model employs graph convolutional network (GCN) to enhance information propagation between related nodes, thereby improving the model’s predictive capabilities. Notably, we introduce a spectral layer constructed through a Fast Fourier Transform (FFT) with random weights to reduce noise generated during the data processing. Experimental results based on datasets demonstrate that our proposed model achieves state-of-the-art performance.

AAAI Conference 2025 Conference Paper

CNC: Cross-modal Normality Constraint for Unsupervised Multi-class Anomaly Detection

  • Xiaolei Wang
  • Xiaoyang Wang
  • Huihui Bai
  • Eng Gee Lim
  • Jimin Xiao

Existing unsupervised distillation-based methods rely on the differences between encoded and decoded features to locate abnormal regions in test images. However, the decoder trained only on normal samples still reconstructs abnormal patch features well, degrading performance. This issue is particularly pronounced in unsupervised multi-class anomaly detection tasks. We attribute this behavior to ‘over-generalization’ (OG) of decoder: the significantly increasing diversity of patch patterns in multi-class training enhances the model generalization on normal patches, but also inadvertently broadens its generalization to abnormal patches. To mitigate ‘OG’, we propose a novel approach that leverages class-agnostic learnable prompts to capture common textual normality across various visual patterns, and then apply them to guide the decoded features towards a ‘normal’ textual representation, suppressing ‘over-generalization’ of the decoder on abnormal patterns. To further improve performance, we also introduce a gated mixture-of-experts module to specialize in handling diverse patch patterns and reduce mutual interference between them in multi-class training. Our method achieves competitive performance on the MVTec AD and VisA datasets, demonstrating its effectiveness.

IJCAI Conference 2025 Conference Paper

Free Lunch of Image-mask Alignment for Anomaly Image Generation and Segmentation

  • Xiangyue Li
  • Xiaoyang Wang
  • Zhibin Wan
  • Quan Zhang
  • Yupei Wu
  • Tao Deng
  • Mingjie Sun

This paper aims at generating anomalous images and their segmentation labels to address the lack of real-world anomaly samples and privacy issues. Departing from conventional approaches that use masks solely to guide the generation of anomaly images, we propose a dual-branch training strategy for the generative model. This strategy enables the simultaneous production of anomaly images and masks, with an alignment regularization loss that ensures the coherence between the generated images and their masks. During inference, only the image-generation branch is activated to produce synthetic samples for training the downstream segmentation model. Furthermore, we propose to integrate the well-trained generative model into the training of segmentation models, utilizing a generative feedback loss to refine the segmentation model's performance. Experiments show our method's IoU metrics exceed previous methods by 5. 03%, 5. 68% and 16. 63% on Real-IAD (industrial), polyp (medical), and Floor Dirty (indoor) datasets. The code is publicly accessible at https: //github. com/huan-yin/anomaly-alignment.

TMLR Journal 2025 Journal Article

Improving Single-round Active Adaptation: A Prediction Variability Perspective

  • Xiaoyang Wang
  • Yibo Jacky Zhang
  • Olawale Elijah Salaudeen
  • Mingyuan Wu
  • Hongpeng Guo
  • Chaoyang He
  • Klara Nahrstedt
  • Sanmi Koyejo

Machine learning models trained with offline data often suffer from distribution shifts in online environments and require fast adaptation to online data. The high volume of online data further stimulates the study of active adaptation approaches that achieve competitive adaptation performance by selectively annotating only 5%-10% of online data and using it to continuously train a model. Despite the reduction in data annotation cost, many prior active adaptations assume a multi-round data annotation procedure during continuous training, which hinders timely adaptation. In this work, we study a single-round active adaptation problem with a minimum data annotation turnaround time but require the selected subset of data samples to help the entire continuous training procedure until convergence. In our theoretical analysis, we find that the prediction variability of each data sample throughout the training is crucial, in addition to the conventional data diversity. The prediction variability measures how much the prediction could possibly change during the continuous training procedure. To this end, we introduce a novel approach called feature-norm scaled gradient embedding (FORGE), which incorporates prediction variability and improves the single-round active adaptation performance when combined with standard data selection strategies (e.g., k-center greedy). In addition, we provide efficient implementations to construct our FORGE embedding analytically without explicitly backpropagating gradients. Empirical results further demonstrate that our approach consistently outperforms the random selection baseline by up to 1.26% for various vision and language tasks while other competitors often underperform the random selection baseline.

IJCAI Conference 2025 Conference Paper

PCAN: A Pandemic-Compatible Attentive Neural Network for Retail Sales Forecasting

  • Fan Li
  • Guoxuan Wang
  • Huiyu Chu
  • Dawei Cheng
  • Xiaoyang Wang

The outbreak of pandemic has a huge impact on production and consumption in the business world, especially for the retail sector. As a crucial component of decision-support technology in the retail industry, sales forecasting is significant for production planning and optimizing the supply of essential goods during the pandemic. However, due to the irregular fluctuation pattern caused by uncertainty and the complex temporal correlation between multiple covariates and sales, there is still no effective approach for sales forecasting in this extreme event. To fill this gap, we propose a Pandemic-Compatible Attentive Network (PCAN) for retail sales forecasting. Specifically, to capture the irregular fluctuation patterns from the sales series, we design a fluctuation attention mechanism based on association discrepancy in the time series. Then, a parallel attention module is developed to learn the complex relationship between target sales and various dynamic influence factors in a decoupled manner. Finally, we introduce a novel rectification decoding strategy to indicate fluctuation points in prediction. By evaluating PCAN on four real-world retail food datasets from the SF Express international supply chain system, the results show that our method achieves superior performance over the existing state-of-the-art baselines. The model has been deployed in the supply chain system as a fundamental component to serve a world-leading food retailer.

IJCAI Conference 2024 Conference Paper

Hypergraph Self-supervised Learning with Sampling-efficient Signals

  • Fan Li
  • Xiaoyang Wang
  • Dawei Cheng
  • Wenjie Zhang
  • Ying Zhang
  • Xuemin Lin

Self-supervised learning (SSL) provides a promising alternative for representation learning on hypergraphs without costly labels. However, existing hypergraph SSL models are mostly based on contrastive methods with the instance-level discrimination strategy, suffering from two significant limitations: (1) They select negative samples arbitrarily, which is unreliable in deciding similar and dissimilar pairs, causing training bias. (2) They often require a large number of negative samples, resulting in expensive computational costs. To address the above issues, we propose SE-HSSL, a hypergraph SSL framework with three sampling-efficient self-supervised signals. Specifically, we introduce two sampling-free objectives leveraging the canonical correlation analysis as the node-level and group-level self-supervised signals. Additionally, we develop a novel hierarchical membership-level contrast objective motivated by the cascading overlap relationship in hypergraphs, which can further reduce membership sampling bias and improve the efficiency of sample utilization. Through comprehensive experiments on 7 real-world hypergraphs, we demonstrate the superiority of our approach over the state-of-the-art method in terms of both effectiveness and efficiency.

TMLR Journal 2024 Journal Article

Personalized Federated Learning with Spurious Features: An Adversarial Approach

  • Xiaoyang Wang
  • Han Zhao
  • Klara Nahrstedt
  • Sanmi Koyejo

One of the common approaches for personalizing federated learning is fine-tuning the global model for each local client. While this addresses some issues of statistical heterogeneity, we find that such personalization methods are vulnerable to spurious features at local agents, leading to reduced generalization performance. This work considers a setup where spurious features correlate with the label in each client's training environment, and the mixture of multiple training environments (i.e., the global environment) diminishes the spurious correlations. In other words, while the global federated learning model trained over the global environment suffers less from spurious features, the local fine-tuning step may lead to personalized models vulnerable to spurious correlations. In light of this practical and pressing challenge, we propose a novel strategy to mitigate the effect of spurious features during personalization by maintaining the adversarial transferability between the global and personalized models. Empirical results on object and action recognition tasks show that our proposed approach bounds personalized models from further exploiting spurious features while preserving the benefit of enhanced accuracy from fine-tuning.

IROS Conference 2024 Conference Paper

Self-Supervised Monocular Depth Estimation with Effective Feature Fusion and Self Distillation

  • Zhenfei Liu
  • Chengqun Song
  • Jun Cheng
  • Jiefu Luo
  • Xiaoyang Wang

Monocular depth estimation obtaining scene depth information from a single image is an important task in the field of computer vision. Constrained by the limitations of convolutional networks in conducting long-distance modeling and the underutilization of datasets, the generalization of existing models is not satisfactory. In this paper, we propose an adaptive backbone named Internal Fusion Transformer to improve generalization ability compared to convolutional backbone, like HRNet, and a Bilateral Attention module which pays more attention to low-level semantic features compared to previous fuse methods. Meanwhile, we introduce three data augmentation methods, namely cropping-resizing (cr), cropping-shuffling (cs), and mirroring (mi), for self distillation, as well as discuss their contributions to model performance improvement. Our model is trained on the KITTI dataset, and without fine-tuning, tested on NYUv2 and Make3D datasets to confirm the generalization. The experimental results illustrate the effectiveness of our design. Our model also demonstrates better performance compared to other models on the KITTI dataset.

AAAI Conference 2024 Conference Paper

SFC: Shared Feature Calibration in Weakly Supervised Semantic Segmentation

  • Xinqiao Zhao
  • Feilong Tang
  • Xiaoyang Wang
  • Jimin Xiao

Image-level weakly supervised semantic segmentation has received increasing attention due to its low annotation cost. Existing methods mainly rely on Class Activation Mapping (CAM) to obtain pseudo-labels for training semantic segmentation models. In this work, we are the first to demonstrate that long-tailed distribution in training data can cause the CAM calculated through classifier weights over-activated for head classes and under-activated for tail classes due to the shared features among head- and tail- classes. This degrades pseudo-label quality and further influences final semantic segmentation performance. To address this issue, we propose a Shared Feature Calibration (SFC) method for CAM generation. Specifically, we leverage the class prototypes which carry positive shared features and propose a Multi-Scaled Distribution-Weighted (MSDW) consistency loss for narrowing the gap between the CAMs generated through classifier weights and class prototypes during training. The MSDW loss counterbalances over-activation and under-activation by calibrating the shared features in head-/tail-class classifier weights. Experimental results show that our SFC significantly improves CAM boundaries and achieves new state-of-the-art performances. The project is available at https://github.com/Barrett-python/SFC.

YNICL Journal 2023 Journal Article

Aberrant dynamic Functional-Structural connectivity coupling of Large-scale brain networks in poststroke motor dysfunction

  • Xiaoying Liu
  • Shuting Qiu
  • Xiaoyang Wang
  • Hui Chen
  • Yuting Tang
  • Yin Qin

BACKGROUND AND PURPOSE: Stroke may lead to widespread functional and structural reorganization in the brain. Several studies have reported a potential correlation between functional network changes and structural network changes after stroke. However, it is unclear how functional-structural relationships change dynamically over the course of one resting-state fMRI scan in patients following a stroke; furthermore, we know little about their relationships with the severity of motor dysfunction. Therefore, this study aimed to investigate dynamic functional and structural connectivity (FC-SC) coupling and its relationship with motor function in subcortical stroke from the perspective of network dynamics. METHODS: Resting-state functional magnetic resonance imaging and diffusion tensor imaging were obtained from 39 S patients (19 severe and 20 moderate) and 22 healthy controls (HCs). Brain structural networks were constructed by tracking fiber tracts in diffusion tensor imaging, and structural network topology metrics were calculated using a graph-theoretic approach. Independent component analysis, the sliding window method, and k-means clustering were used to calculate dynamic functional connectivity and to estimate different dynamic connectivity states. The temporal patterns and intergroup differences of FC-SC coupling were analyzed within each state. We also calculated dynamic FC-SC coupling and its relationship with functional network efficiency. In addition, the correlation between FC-SC coupling and the Fugl-Meyer assessment scale was analyzed. RESULTS: For SC, stroke patients showed lower global efficiency than HCs (all P < 0.05), and severely affected patients had a higher characteristic path length (P = 0.003). For FC and FC-SC coupling, stroke patients predominantly showed lower local efficiency and reduced FC-SC coupling than HCs in state 2 (all P < 0.05). Furthermore, severely affected patients also showed lower local efficiency (P = 0.031) and reduced FC-SC coupling (P = 0.043) in state 3, which was markedly linked to the severity of motor dysfunction after stroke. In addition, FC-SC coupling was correlated with functional network efficiency in state 2 in moderately affected patients (r = 0.631, P = 0.004) but not significantly in severely affected patients. CONCLUSIONS: Stroke patients show abnormal dynamic FC-SC coupling characteristics, especially in individuals with severe injuries. These findings may contribute to a better understanding of the anatomical functional interactions underlying motor deficits in stroke patients and provide useful information for personalized rehabilitation strategies.

IJCAI Conference 2022 Conference Paper

CARD: Semi-supervised Semantic Segmentation via Class-agnostic Relation based Denoising

  • Xiaoyang Wang
  • Jimin Xiao
  • Bingfeng Zhang
  • Limin Yu

Recent semi-supervised semantic segmentation methods focus on mining extra supervision from unlabeled data by generating pseudo labels. However, noisy labels are inevitable in this process which prevent effective self-supervision. This paper proposes that noisy labels can be corrected based on semantic connections among features. Since a segmentation classifier produces both high and low-quality predictions, we can trace back to feature encoder to investigate how a feature in a noisy group is related to those in the confident groups. Discarding the weak predictions from the classifier, rectified predictions are assigned to the wrongly predicted features through the feature relations. The key to such an idea lies in mining reliable feature connections. With this goal, we propose a class-agnostic relation network to precisely capture semantic connections among features while ignoring their semantic categories. The feature relations enable us to perform effective noisy label corrections to boost self-training performance. Extensive experiments on PASCAL VOC and Cityscapes demonstrate the state-of-the-art performances of the proposed methods under various semi-supervised settings.

AAAI Conference 2022 Short Paper

Efficient Attribute (α,β)-Core Detection in Large Bipartite Graphs (Student Abstract)

  • Yanping Wu
  • Renjie Sun
  • Chen Chen
  • Xiaoyang Wang

In this paper, we propose a novel problem, named rational (α, β)-core detection in attribute bipartite graphs (RCD- ABG), which retrieves the connected (α, β)-core with the largest rational score. A basic greedy framework with an optimized strategy is developed and extensive experiments are conducted to evaluate the performance of the techniques.

CLeaR Conference 2022 Conference Paper

Identifying Coarse-grained Independent Causal Mechanisms with Self-supervision

  • Xiaoyang Wang
  • Klara Nahrstedt
  • Oluwasanmi O Koyejo

Among the most effective methods for uncovering high dimensional unstructured data’s generating mechanisms are techniques based on disentangling and learning independent causal mechanisms. However, to identify the disentangled model, previous methods need additional observable variables or do not provide identifiability results. In contrast, this work aims to design an identifiable generative model that approximates the underlying mechanisms from observational data using only self-supervision. Specifically, the generative model uses a degenerate mixture prior to learn mechanisms that generate or transform data. We outline sufficient conditions for an identifiable generative model up to three types of transformations that preserve a coarse-grained disentanglement. Moreover, we propose a self-supervised training method based on these identifiability conditions. We validate our approach on MNIST, FashionMNIST, and Sprites datasets, showing that the proposed method identifies disentangled models – by visualization and evaluating the downstream predictive model’s accuracy under environment shifts.

AAAI Conference 2021 Conference Paper

NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation

  • Xiaoyang Wang
  • Chen Li
  • Jianqiao Zhao
  • Dong Yu

In this paper, we propose a Chinese multi-turn topic-driven conversation dataset, NaturalConv, which allows the participants to chat anything they want as long as any element from the topic is mentioned and the topic shift is smooth. Our corpus contains 19. 9K conversations from six domains, and 400K utterances with an average turn number of 20. 1. These conversations contain in-depth discussions on related topics or widely natural transition between multiple topics. We believe either way is normal for human conversation. To facilitate the research on this corpus, we provide results of several benchmark models. Comparative results show that for this dataset, our current models are not able to provide significant improvement by introducing background knowledge/topic. Therefore, the proposed dataset should be a good benchmark for further research to evaluate the validity and naturalness of multi-turn conversation systems. Our dataset is available at https: //ai. tencent. com/ailab/nlp/dialogue/#datasets.

IJCAI Conference 2020 Conference Paper

Risk Guarantee Prediction in Networked-Loans

  • Dawei Cheng
  • Xiaoyang Wang
  • Ying Zhang
  • Liqing Zhang

The guaranteed loan is a debt obligation promise that if one corporation gets trapped in risks, its guarantors will back the loan. When more and more companies involve, they subsequently form complex networks. Detecting and predicting risk guarantee in these networked-loans is important for the loan issuer. Therefore, in this paper, we propose a dynamic graph-based attention neural network for risk guarantee relationship prediction (DGANN). In particular, each guarantee is represented as an edge in dynamic loan networks, while companies are denoted as nodes. We present an attention-based graph neural network to encode the edges that preserve the financial status as well as network structures. The experimental result shows that DGANN could significantly improve the risk prediction accuracy in both the precision and recall compared with state-of-the-art baselines. We also conduct empirical studies to uncover the risk guarantee patterns from the learned attentional network features. The result provides an alternative way for loan risk management, which may inspire more work in the future.

IJCAI Conference 2019 Conference Paper

Pivotal Relationship Identification: The K-Truss Minimization Problem

  • Weijie Zhu
  • Mengqi Zhang
  • Chen Chen
  • Xiaoyang Wang
  • Fan Zhang
  • Xuemin Lin

In a social network, the strength of relationships between users can significantly affect the stability of the network. In this paper, we use the k-truss model to measure the stability of a social network. To identify critical connections, we propose a novel problem, named k-truss minimization. Given a social network G and a budget b, it aims to find b edges for deletion which can lead to the maximum number of edge breaks in the k-truss of G. We show that the problem is NP-hard. To accelerate the computation, novel pruning rules are developed to reduce the candidate size. In addition, we propose an upper bound based strategy to further reduce the searching space. Comprehensive experiments are conducted over real social networks to demonstrate the efficiency and effectiveness of the proposed techniques.

AAAI Conference 2019 Conference Paper

STA: Spatial-Temporal Attention for Large-Scale Video-Based Person Re-Identification

  • Yang Fu
  • Xiaoyang Wang
  • Yunchao Wei
  • Thomas Huang

In this work, we propose a novel Spatial-Temporal Attention (STA) approach to tackle the large-scale person reidentification task in videos. Different from the most existing methods, which simply compute representations of video clips using frame-level aggregation (e. g. average pooling), the proposed STA adopts a more effective way for producing robust clip-level feature representation. Concretely, our STA fully exploits those discriminative parts of one target person in both spatial and temporal dimensions, which results in a 2-D attention score matrix via inter-frame regularization to measure the importances of spatial parts across different frames. Thus, a more robust clip-level feature representation can be generated according to a weighted sum operation guided by the mined 2-D attention score matrix. In this way, the challenging cases for video-based person re-identification such as pose variation and partial occlusion can be well tackled by the STA. We conduct extensive experiments on two large-scale benchmarks, i. e. MARS and DukeMTMC- VideoReID. In particular, the mAP reaches 87. 7% on MARS, which significantly outperforms the state-of-the-arts with a large margin of more than 11. 6%.

IJCAI Conference 2018 Conference Paper

Request-and-Reverify: Hierarchical Hypothesis Testing for Concept Drift Detection with Expensive Labels

  • Shujian Yu
  • Xiaoyang Wang
  • José C. Príncipe

One important assumption underlying common classification models is the stationarity of the data. However, in real-world streaming applications, the data concept indicated by the joint distribution of feature and label is not stationary but drifting over time. Concept drift detection aims to detect such drifts and adapt the model so as to mitigate any deterioration in the model's predictive performance. Unfortunately, most existing concept drift detection methods rely on a strong and over-optimistic condition that the true labels are available immediately for all already classified instances. In this paper, a novel Hierarchical Hypothesis Testing framework with Request-and-Reverify strategy is developed to detect concept drifts by requesting labels only when necessary. Two methods, namely Hierarchical Hypothesis Testing with Classification Uncertainty (HHT-CU) and Hierarchical Hypothesis Testing with Attribute-wise "Goodness-of-fit" (HHT-AG), are proposed respectively under the novel framework. In experiments with benchmark datasets, our methods demonstrate overwhelming advantages over state-of-the-art unsupervised drift detectors. More importantly, our methods even outperform DDM (the widely used supervised drift detector) when we use significantly fewer labels.

IJCAI Conference 2016 Conference Paper

Object Recognition with Hidden Attributes

  • Xiaoyang Wang
  • Qiang Ji

Attribute based object recognition performs object recognition using the semantic properties of the object. Unlike the existing approaches that treat attributes as a middle level representation and require to estimate the attributes during testing, we propose to incorporate the hidden attributes, which are the attributes used only during training to improve model learning and are not needed during testing. To achieve this goal, we develop two different approaches to incorporate hidden attributes. The first approach utilizes hidden attributes as additional information to improve the object classification model. The second approach further exploits the semantic relationships between the objects and the hidden attributes. Experiments on benchmark data sets demonstrate that both approaches can effectively improve the learning of the object classifiers over the baseline models that do not use attributes, and their combination reaches the best performance. Experiments also show that the proposed approaches outperform both state of the art methods that use attributes as middle level representation and the approaches that learn the classifiers with hidden information.

v2026.09.13