Arrow Research search

Author name cluster

Xiaowei Xu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

7 papers
1 author row

Possible papers

7

AAAI Conference 2026 Conference Paper

Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment

  • Yang Chen
  • Xiaowei Xu
  • Shuai Wang
  • Chenhui Zhu
  • Ruxue Wen
  • Xubin Li
  • Tiezheng Ge
  • Limin Wang

Normalizing Flows (NFs) are a class of generative models distinguished by a mathematically invertible architecture, where the forward pass transforms data into a latent space for density estimation, and the reverse pass generates new samples from this space. This characteristic creates an intrinsic synergy between representation learning and data generation. However, the generative quality of standard NFs is limited by poor semantic representations from log-likelihood optimization. To remedy this, we propose a novel alignment strategy that creatively leverages the invertibility of NFs: instead of regularizing the forward pass, we align the intermediate features of the generative (reverse) pass with representations from a powerful vision foundation model, demonstrating superior effectiveness over naive alignment. We also introduce a novel training-free, test-time optimization algorithm for classification, which provides a more intrinsic evaluation of the NF's embedded semantic knowledge. Comprehensive experiments demonstrate that our approach accelerates the training of NFs by over 3.3x, while simultaneously delivering significant improvements in both generative quality and classification accuracy. New state-of-the-art results for NFs are established on ImageNet 64 x 64 and 256 x 256.

EAAI Journal 2026 Journal Article

Power load forecasting based on time-frequency domain feature fusion

  • Wenhua Jiao
  • Xiao Han
  • Ce Yu
  • Xiaowei Xu
  • Qing Zhang
  • Xiang Zhang
  • Bin Wang
  • Lijuan Li

Power Load Forecasting (PLF) provides decision support for grid planning and efficiency improvement and plays a pivotal role in reducing energy consumption. Current PLF methodologies face significant challenges in effectively coordinating the complex interdependencies among multiple factors affecting power system load. To address this fundamental limitation, we propose the dual-stream architecture named Time -Frequency Interaction Network (TF-interactionNet) that synergistically integrates the time-frequency domain, enabling it to consider both autocorrelation and interaction of time series. Our approach introduces three specialized modules: 1) The Time-Mamba (T-Mamba) module employs advanced state space models to capture long-range temporal dependencies and global trend patterns in power load sequences; 2) The Frequency-Temporal Convolutional Network (F-TCN) module utilizes frequency-optimized temporal convolutional network to extract localized frequency characteristics and identify subtle load fluctuation patterns. 3) The proposed Feature Attention Fusion Network (FAFN) innovatively fuses complementary features in the time-frequency domain, retaining the key effective features while avoiding the information redundancy that occurs during fusion. Extensive experimental evaluations show that TF-interactionNet achieves state-of-the-art performance across multiple evaluation metrics. On three datasets, compared to the latest forecasting model, the proposed model reduced the mean squared error of 24-h load forecasting by 0. 1%, 0. 1%, and 1%, respectively, while the mean absolute percentage error decreased by 0. 198%, 0. 066%, and 0. 106%, respectively. It still maintains high effectiveness in predicting power load changes over longer time periods. This model improves forecasting accuracy across multiple prediction horizons, offering new insights for advancing PLF and holding sustained research value.

EAAI Journal 2025 Journal Article

Quantization-based deep diversified ensemble for medical image segmentation

  • Jiawei Zhang
  • Jialin Wang
  • Qi Wang
  • Yanchun Zhang
  • Weihong Han
  • Yangyang Mei
  • Yiyu Shi
  • Jian Zhuang

Recent advancements in fully convolutional networks (FCNs) have significantly improved medical image segmentation. Ensemble methods are often used to further enhance performance, with diversity among learners being a critical factor. However, many current approaches focus on diversifying training samples or predictions while overlooking the diversity of internal multi-scale features. This oversight can lead to high correlations among features across different learners, limiting overall effectiveness. Additionally, traditional quantization methods aim to minimize accuracy loss by maintaining a rigid quantization process. This rigidity can eliminate the randomness introduced by quantization, further reducing ensemble diversity and effectiveness. In this paper, we propose a novel approach called Quantization-based Deep Diversified Ensemble (QDD-Ens) for medical image segmentation. Our method enhances the diversity of internal features among ensemble learners through two mechanisms: deep diversified loss, which focuses on feature diversity rather than segmentation accuracy, and deep diversified quantization, which preserves beneficial randomness in quantization process. Furthermore, QDD-Ens facilitates a deeper form of ensemble learning by employing a meta-learner to integrate diversified features at multiple resolution levels from various base learners, which are diversified by two above diversify enhancement mechanisms. Extensive experiments on five public medical image segmentation datasets show that our method significantly improves segmentation accuracy and outperforms existing ensemble techniques. The source code is publicly available to support future research. (https: //github. com/JerRuy/QDD-Ens)

TCS Journal 2023 Journal Article

Many-to-many edge-disjoint paths in (n,k)-enhanced hypercube under three link-faulty hypotheses

  • Hongxi Liu
  • Mingzu Zhang
  • Xiaowei Xu

One of the most central issues in interconnection networks of parallel and distributed systems is finding edge-disjoint paths that transmit information. Finding as many as possible many-to-many edge-disjoint paths is conducive to improving the fault-tolerance of such networks. As an interconnection network topology, ( n, k ) -enhanced hypercube Q n, k ( 1 ≤ k ≤ n − 1 ), is a momentous variant of well-known hypercube. For integers 1 ≤ l ≤ n − 1 and n ≥ 2, let δ = 0 if 1 ≤ l ≤ n − k, and δ = 1 if n − k + 1 ≤ l ≤ n − 1. This paper offers a unified method to determine the minimum cardinalities of faulty links in Q n, k, whose malfunction divides this network into several connected components such that each processor has at least l + δ neighbors, each component contains no less than 2 l processors and the number of average neighbors for all processors is at least l + δ, respectively. Under these three different link-faulty assumptions, but for the condition of k = 2 and l = n − 2, the minimum cardinalities of such faulty links share the same value ( n − l − δ + 1 ) 2 l. And the value in the exceptional case is ( n − l − δ ) 2 l + 1 = 2 n − 1. In other words, we find the maximum numbers of many-to-many edge-disjoint paths of Q n, k under the above three hypotheses, which offers refined measurements for the reliability and fault-tolerance of interconnection networks.

AAAI Conference 2021 Conference Paper

C2F-FWN: Coarse-to-Fine Flow Warping Network for Spatial-Temporal Consistent Motion Transfer

  • Dongxu Wei
  • Xiaowei Xu
  • Haibin Shen
  • Kejie Huang

Human video motion transfer (HVMT) aims to synthesize videos that one person imitates other persons’ actions. Although existing GAN-based HVMT methods have achieved great success, they either fail to preserve appearance details due to the loss of spatial consistency between synthesized and exemplary images, or generate incoherent video results due to the lack of temporal consistency among video frames. In this paper, we propose Coarse-to-Fine Flow Warping Network (C2F-FWN) for spatial-temporal consistent HVMT. Particularly, C2F-FWN utilizes coarse-to-fine flow warping and Layout-Constrained Deformable Convolution (LC-DConv) to improve spatial consistency, and employs Flow Temporal Consistency (FTC) Loss to enhance temporal consistency. In addition, provided with multi-source appearance inputs, C2F-FWN can support appearance attribute editing with great flexibility and efficiency. Besides public datasets, we also collected a large-scale HVMT dataset named SoloDance for evaluation. Extensive experiments conducted on our SoloDance dataset and the iPER dataset show that our approach outperforms state-of-art HVMT methods in terms of both spatial and temporal consistency. Source code and the SoloDance dataset are available at https: //github. com/wswdx/C2F-FWN.

AAAI Conference 2019 Conference Paper

SCNN: A General Distribution Based Statistical Convolutional Neural Network with Application to Video Object Detection

  • Tianchen Wang
  • Jinjun Xiong
  • Xiaowei Xu
  • Yiyu Shi

Various convolutional neural networks (CNNs) were developed recently that achieved accuracy comparable with that of human beings in computer vision tasks such as image recognition, object detection and tracking, etc. Most of these networks, however, process one single frame of image at a time, and may not fully utilize the temporal and contextual correlation typically present in multiple channels of the same image or adjacent frames from a video, thus limiting the achievable throughput. This limitation stems from the fact that existing CNNs operate on deterministic numbers. In this paper, we propose a novel statistical convolutional neural network (SCNN), which extends existing CNN architectures but operates directly on correlated distributions rather than deterministic numbers. By introducing a parameterized canonical model to model correlated data and defining corresponding operations as required for CNN training and inference, we show that SCNN can process multiple frames of correlated images effectively, hence achieving significant speedup over existing CNN models. We use a CNN based video object detection as an example to illustrate the usefulness of the proposed SCNN as a general network model. Experimental results show that even a nonoptimized implementation of SCNN can still achieve 178% speedup over existing CNNs with slight accuracy degradation.

AAAI Conference 2018 Conference Paper

Active Lifelong Learning With “Watchdog”

  • Gan Sun
  • Yang Cong
  • Xiaowei Xu

Lifelong learning intends to learn new consecutive tasks depending on previously accumulated experiences, i. e. , knowledge library. However, the knowledge among different new coming tasks are imbalance. Therefore, in this paper, we try to mimic an effective “human cognition” strategy by actively sorting the importance of new tasks in the process of unknown-to-known and selecting to learn the important tasks with more information preferentially. To achieve this, we consider to assess the importance of the new coming task, i. e. , unknown or not, as an outlier detection issue, and design a hierarchical dictionary learning model consisting of two-level task descriptors to sparse reconstruct each task with the 0 norm constraint. The new coming tasks are sorted depending on the sparse reconstruction score in descending order, and the task with high reconstruction score will be permitted to pass, where this mechanism is called as “watchdog”. Next, the knowledge library of the lifelong learning framework encode the selected task by transferring previous knowledge, and then can also update itself with knowledge from both previously learned task and current task automatically. For model optimization, the alternating direction method is employed to solve our model and converges to a fixed point. Extensive experiments on both benchmark datasets and our own dataset demonstrate the effectiveness of our proposed model especially in task selection and dictionary learning.

v2026.09.13