Arrow Research search

Author name cluster

Xue Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

31 papers
2 author rows

Possible papers

31

JBHI Journal 2026 Journal Article

Mining Global and Local Semantics From Unlabeled Spectra for Spectral Classification

  • Wei Luo
  • Haiming Yao
  • Ang Gao
  • Tao Zhou
  • Xue Wang

Non-destructive detection methods based on molecular vibrational spectroscopy are pivotal in fields such as analytical chemistry and medical diagnostics. Recent advances have integrated deep learning with vibrational spectroscopy, significantly enhancing spectral recognition accuracy. However, these methods often rely on large annotated spectral datasets, limiting their general applicability. To address this limitation, we propose a novel approach, G lobal and L ocal S emantics M ining (GLSM), which leverages self-supervised learning to capture the global and local semantic information of unlabeled spectra, obviating the need for extensive annotated data. We devise two proxy tasks: global semantic mining and local semantic mining. The global semantic mining task is based on the premise that different views of the same spectrum can be mutually transformed, enabling the model to capture domain-invariant features across various perspectives and thereby develop a global understanding of the spectral data. This, in turn, enhances the model’s robustness to variations in peak positions. Meanwhile, the local semantic mining task posits that noisy spectra can be reconstructed into noise-free spectra, thereby facilitating the extraction of local patterns and fine-grained details, such as subtle variations in peak intensities. By combining both self-supervised tasks, our model effectively captures the global and local semantic information of the spectrum. The pre-trained model can be fine-tuned with a limited amount of labeled homologous or heterologous spectral data for semi-supervised or transfer learning-based spectral classification. Extensive experiments on three datasets in semi-supervised and transfer learning-based spectral recognition tasks comprehensively validate the effectiveness of our GLSM method, demonstrating its significant potential for real-world spectral analysis applications.

EAAI Journal 2026 Journal Article

Non-destructive Printed Circuit Board layout verification using a deterministic diffusion-guided framework

  • Xue Wang
  • Deruo Cheng
  • Chai Kiat Yeo

Non-destructive Printed Circuit Board (PCB) layout verification is critical for ensuring microelectronics reliability, yet PCB X-ray images are degraded by serious artifacts resulting from surface-mounted components due to the complex structural distortions during computed tomography (CT) imaging, such as occlusion, metallic interference, and scattering, making them challenging to mitigate. This challenge remains largely unexplored with the only related method on bare PCBs degrading seriously when applied in this context. Stochastic diffusion models work poorly on structural artifacts and rely on huge resources requiring many epochs of training. Furthermore, it is impractical to obtain noisy and the corresponding clean image pairs in real-world PCB scenarios for model training. To address these challenges, this work presents the first systematic analysis of six factors degrading non-destructive PCB X-ray imaging and a deterministic diffusion-guided layout verification framework (D2LVer) is proposed for non-destructive PCB layout verification, reducing computational cost while improving segmentation performance. D2LVer employs a deterministic reverse diffusion-guided denoised feature extraction (DDFE) in a single forward-reverse diffusion cycle by directly predicting the posterior mean as denoised features, and aggregates them with the original noisy input for segmentation. DDFE eliminates the need for extensive training, reducing computational overhead by up to 80%. The overall framework D2LVer improves dice score by up to 9. 33% over existing state-of-the-art.

AAAI Conference 2026 Conference Paper

SimDiff: Simpler Yet Better Diffusion Model for Time Series Point Forecasting

  • Hang Ding
  • Xue Wang
  • Tian Zhou
  • Tao Yao

Diffusion models have recently shown promise in time series forecasting, particularly for probabilistic predictions. However, they often fail to achieve state-of-the-art point estimation performance compared to regression-based methods. This limitation stems from difficulties in providing sufficient contextual bias to track distribution shifts and in balancing output diversity with the stability and precision required for point forecasts. Existing diffusion-based approaches mainly focus on full-distribution modeling under probabilistic frameworks, often with likelihood maximization objectives, while paying little attention to dedicated strategies for high-accuracy point estimation. Moreover, other existing point prediction diffusion methods frequently rely on pre-trained or jointly trained mature models for contextual bias, sacrificing the generative flexibility of diffusion models. To address these challenges, we propose SimDiff, a single-stage, end-to-end framework. SimDiff employs a single unified Transformer network carefully tailored to serve as both denoiser and predictor, eliminating the need for external pre-trained or jointly trained regressors. It achieves state-of-the-art point estimation performance by leveraging intrinsic output diversity and improving mean squared error accuracy through multiple inference ensembling. Key innovations, including normalization independence and the median-of-means estimator, further enhance adaptability and stability. Extensive experiments demonstrate that SimDiff significantly outperforms existing methods in time series point forecasting.

EAAI Journal 2025 Journal Article

Adversarial contrastive domain-generative learning for bacteria Raman spectrum joint denoising and cross-domain identification

  • Haiming Yao
  • Wei Luo
  • Xue Wang

Raman spectroscopy, as a label-free detection technology, has been widely utilized in the clinical diagnosis of pathogenic bacteria. However, Raman signals are naturally weak and sensitive to the condition of the acquisition process. The characteristic spectra of a bacteria can manifest varying signal-to-noise ratios and domain discrepancies under different acquisition conditions. Consequently, existing methods often face challenges when identifying unobserved acquisition conditions, i. e. , the testing acquisition conditions are unavailable during model training. In this article, a generic framework, namely, an adversarial contrastive domain-generative learning framework, is proposed for joint Raman spectroscopy denoising and cross-domain identification. The proposed method is composed of a domain generation module and a domain task module. Through adversarial learning between these two modules, it utilizes only a single available source domain spectral data to generate extended denoised domains that are semantically consistent with the source domain and extract domain-invariant representations. Experimental results show that the proposed method significantly enhances diagnostic performance, with an average recognition accuracy improvement of over +5. 0% under unknown acquisition conditions compared to existing methods. Notably, the proposed method also performs simultaneous denoising of the spectra, enhancing the signal-to-noise ratio of the original signal by an average gain of +3. 0, without requiring noise-free ground truth. These results suggest that the proposed method holds great potential as a diagnostic tool for real-world clinical cases.

NeurIPS Conference 2025 Conference Paper

DecompNet: Enhancing Time Series Forecasting Models with Implicit Decomposition

  • Donghao Luo
  • Xue Wang

In this paper, we pioneer the idea of implicit decomposition. And based on this idea, we propose a powerful decomposition-based enhancement framework, namely DecompNet. Our method converts the time series decomposition into an implicit process, where it can give a time series model the decomposition-related knowledge during inference, even though this model does not actually decompose the input time series. Thus, our DecompNet can enable a model to inherit the performance promotion brought by time series decomposition but will not introduce any additional inference costs, successfully enhancing the model performance while enjoying better efficiency. Experimentally, our DecompNet exhibits promising enhancement capability and compelling framework generality. Especially, it can also enhance the performance of the latest and state-of-the-art models, greatly pushing the performance limit of time series forecasting. Through comprehensive comparisons, DecompNet also shows excellent performance and efficiency superiority, making the decomposition-based enhancement framework surpass the well-recognized normalization-based frameworks for the first time. Code is available at this repository: https: //github. com/luodhhh/DecompNet.

IROS Conference 2025 Conference Paper

Dual-Arm Hierarchical Planning for Laboratory Automation: Vibratory Sieve Shaker Operations

  • Haoran Xiao
  • Xue Wang
  • Huimin Lu 0002
  • Zhiwen Zeng
  • Zirui Guo
  • Ziqi Ni
  • Yicong Ye
  • Wei Dai 0014

This paper addresses the challenges of automating vibratory sieve shaker operations in a materials laboratory, focusing on three critical tasks: 1) dual-arm lid manipulation in 3 cm clearance spaces, 2) bimanual handover in overlapping workspaces, and 3) obstructed powder sample container delivery with orientation constraints. These tasks present significant challenges, including inefficient sampling in narrow passages, the need for smooth trajectories to prevent spillage, and suboptimal paths generated by conventional methods. To overcome these challenges, we propose a hierarchical planning framework combining Prior-Guided Path Planning and Multi-Step Trajectory Optimization. The former uses a finite Gaussian mixture model to improve sampling efficiency in narrow passages, while the latter refines paths by shortening, simplifying, imposing joint constraints, and B-spline smoothing. Experimental results demonstrate the framework’s effectiveness: planning time is reduced by up to 80. 4%, and waypoints are decreased by 89. 4%. Furthermore, the system completes the full vibratory sieve shaker operation workflow in a physical experiment, validating its practical applicability for complex laboratory automation.

EAAI Journal 2025 Journal Article

FSGFuse: Feature synergy-guided multi-modality image fusion

  • Qiuhan Shao
  • Zheng Guan
  • Xue Wang
  • Hang Li

The purpose of infrared and visible image fusion (IVIF) is to integrate the complementary properties of the two modalities to obtain a fused image with a more comprehensive information representation. However, existing multi-modal image fusion (MMIF) methods mainly rely on auto-encoders to extract source features and achieve feature complementarity through fusion mechanisms. These methods fail to fully utilize the potential of the encoder in feature decoupling and information integration, resulting in difficulties in effectively capturing and fusing cross-modal features, which in turn introduces redundant information. To tackle the challenge, this study proposes a Feature Synergy-Guided Fusion (FSGFuse) network, which achieves IVIF by synergizing complementary features to preserve texture details and acquire thermal target information. In the pre-training stage, the feature synergism loss is utilized to guide the double-branch encoder to perceive the consistency and difference of cross-modal features, which in turn motivates the model to capture the more discriminatory feature representations in the source features. In the fusion stage, a Feature Complementary Fusion (FCF) module is designed to guide feature fusion. This module not only integrates complementary features, but also captures long-range contextual relationships between features, which facilitates cross-modal interaction while enriching the semantic information. Experiments on publicly available benchmark datasets show that the fusion performance of FSGFuse significantly outperforms existing state-of-the-art methods and positively affects the effectiveness of the downstream object detection task, with a mean average precision (mAP) improvement of 18. 57%. This result fully demonstrates its practical application value.

NeurIPS Conference 2025 Conference Paper

Less is More: Unlocking Specialization of Time Series Foundation Models via Structured Pruning

  • Lifan Zhao
  • Yanyan Shen
  • Zhaoyang Liu
  • Xue Wang
  • Jiaji Deng

Scaling laws motivate the development of Time Series Foundation Models (TSFMs) that pre-train vast parameters and achieve remarkable zero-shot forecasting performance. Surprisingly, even after fine-tuning, TSFMs cannot consistently outperform smaller, specialized models trained on full-shot downstream data. A key question is how to realize effective adaptation of TSFMs for a target forecasting task. Through empirical studies on various TSFMs, the pre-trained models often exhibit inherent sparsity and redundancy in computation, suggesting that TSFMs have learned to activate task-relevant network substructures to accommodate diverse forecasting tasks. To preserve this valuable prior knowledge, we propose a structured pruning method to regularize the subsequent fine-tuning process by focusing it on a more relevant and compact parameter space. Extensive experiments on seven TSFMs and six benchmarks demonstrate that fine-tuning a smaller, pruned TSFM significantly improves forecasting performance compared to fine-tuning original models. This ``prune-then-finetune'' paradigm often enables TSFMs to achieve state-of-the-art performance and surpass strong specialized baselines. Source code is made publicly available at \url{https: //github. com/SJTU-DMTai/Prune-then-Finetune}.

NeurIPS Conference 2025 Conference Paper

MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling

  • Yuxi Liu
  • Renjia Deng
  • Yutong He
  • Xue Wang
  • Tao Yao
  • Kun Yuan

The substantial memory demands of pre-training and fine-tuning large language models (LLMs) require memory-efficient optimization algorithms. One promising approach is layer-wise optimization, which treats each transformer block as a single layer and optimizes it sequentially, while freezing the other layers to save optimizer states and activations. Although effective, these methods ignore the varying importance of the modules within each layer, leading to suboptimal performance. Moreover, layer-wise sampling provides only limited memory savings, as at least one full layer must remain active during optimization. To overcome these limitations, we propose **M**odule-wise **I**mportance **SA**mpling (**MISA**), a novel method that divides each layer into smaller modules and assigns importance scores to each module. MISA uses a weighted random sampling mechanism to activate modules, provably reducing gradient variance compared to layer-wise sampling. Additionally, we establish an $\mathcal{O}(1/\sqrt{K})$ convergence rate under non-convex and stochastic conditions, where $K$ is the total number of training steps, and provide a detailed memory analysis showcasing MISA's superiority over existing baseline methods. Experiments on diverse learning tasks validate the effectiveness of MISA.

IJCAI Conference 2025 Conference Paper

Pre-defined Keypoints Promote Category-level Articulation Pose Estimation via Multi-Modal Alignment

  • Wenbo Xu
  • Li Zhang
  • Liu Liu
  • Yan Zhong
  • Haonan Jiang
  • Xue Wang
  • Rujing Wang

Articulations are essential in everyday interactions, yet traditional RGB-based pose estimation methods often struggle with issues such as lighting variations and shadows. To overcome these challenges, we propose a novel Pre-defined keypoint based framework for category-level articulation pose estimation via multi-modal Alignment, coined PAGE. Specifically, we first propose a customized keypoint estimation method, aiming to avoid the divergent distance pattern between heuristically generated keypoints and visible points. In addition, to reduce the mutual information redundancy between point clouds and RGB images, we design the geometry-color alignment, which fuses the features after aligning two modalities. This is followed by decoding the radius for each visible point, and applying our proposal integration scoring strategy to predict keypoints. Ultimately, the framework outputs the per-part 6D pose of the articulation. We conduct extensive experiments to evaluate PAGE across a variety of datasets, from synthetic to real-world scenarios, demonstrating its robustness and superior performance.

AAAI Conference 2025 Conference Paper

R^2-Art: Category-Level Articulation Pose Estimation from Single RGB Image via Cascade Render Strategy

  • Li Zhang
  • Haonan Jiang
  • Yukang Huo
  • Yan Zhong
  • Jianan Wang
  • Xue Wang
  • Rujing Wang
  • Liu Liu

Human life is filled with articulated objects. Previous works for estimating the pose of category-level articulated objects rely on costly 3D point clouds or RGB-D images. In this paper, our goal is to estimate category-level articulation poses from a single RGB image, where we propose R2-Art, a novel category-level Articulation pose estimation framework from a single RGB image and a cascade Render strategy. Given an RGB image as input, R2-Art estimates per-part 6D pose for the articulation. Specifically, we design parallel regression branches tailored to generate camera-to-root translation and rotation. Using the predicted joint states, we perform PC prior transformation and deformation with a joint-centric modeling approach. For further refinement, a cascade render strategy is proposed for projecting the 3D deformed prior onto the 2D mask. Extensive experiments are provided to validate our R2-Art on various datasets ranging from synthetic datasets to real-world scenarios, demonstrating the superior performance and robustness of the R2-Art. We believe that this work has the potential to be applied in many fields including robotics, embodied intelligence, and augmented reality.

NeurIPS Conference 2025 Conference Paper

RePO: Understanding Preference Learning Through ReLU-Based Optimization

  • Junkang Wu
  • Kexin Huang
  • Xue Wang
  • Jinyang Gao
  • Bolin Ding
  • Jiancan Wu
  • Xiangnan He
  • Xiang Wang

Preference learning has become a common approach in various recent methods for aligning large language models with human values. These methods optimize the preference margin between chosen and rejected responses, subject to certain constraints for avoiding over-optimization. In this paper, we report surprising empirical findings that simple ReLU activation can learn meaningful alignments even using \emph{none} of the following: (i) sigmoid-based gradient constraints, (ii) explicit regularization terms. Our experiments show that over-optimization does exist, but a threshold parameter $\gamma$ plays an essential role in preventing it by dynamically filtering training examples. We further provide theoretical analysis demonstrating that ReLU-based Preference Optimization (RePO) corresponds to the convex envelope of the 0-1 loss, establishing its fundamental soundness. Our ``RePO'' method achieves competitive or superior results compared to established preference optimization approaches. We hope this simple baseline will motivate researchers to rethink the fundamental mechanisms behind preference optimization for language model alignment.

YNICL Journal 2025 Journal Article

Structural and functional changes of Post-Stroke Depression: A multimodal magnetic resonance imaging study

  • Qiuhong Lu
  • Shunzu Lu
  • Xue Wang
  • Yanlan Huang
  • Jie Liu
  • Zhijian Liang

This study investigated changes in gray matter volume (GMV), white matter microstructure, and spontaneous brain activity in post-stroke depression (PSD) using multiple MRI techniques, including neurite orientation dispersion and density imaging (NODDI). Changes in GMV, neurite density index (NDI), orientation dispersion index (ODI), fraction of isotropic water (ISO), diffusion tensor imaging (DTI) parameters, and the amplitude of frequency fluctuations (ALFF) were assessed between PSD (n = 20), post-stroke without depression (n = 20), and normal control (n = 20) groups. Receiver operating characteristic (ROC) curve analysis was performed to test the classification performance of the variant parameters of each MRI modality, each single MRI modality and multiple MRI modality. Compared to patients with post-stroke without depression (non-PSD), those with PSD showed increased ODI and ISO in the widespread white matter, as well as increased ALFF in the left pallidum. No significant differences in the GMV or DTI parameters were observed between the two groups. Furthermore, the ODI of the right superior longitudinal fasciculus and NODDI showed the best classification performance for PSD at their respective comparison level (the areas under the ROC curves (AUC) = 0.917(0.000), 0.933(0.000)). The model of NODDI-derived parameters combined with non-diffusion MRI modality parameters (i.e., GMV and ALFF) showed better diagnostic performance than that of DTI-derived parameters. These findings suggest that PSD is associated with structural and functional abnormalities that may contribute to depressive symptoms. Additionally, NODDI showed its advantages in the description of structural alterations in emotion-related white matter pathways and classification performance in PSD.

EAAI Journal 2024 Journal Article

A hybrid complex spectral conjugate gradient learning algorithm for complex-valued data processing

  • Ke Zhang
  • Huisheng Zhang
  • Xue Wang

Complex-valued neural networks (CVNNs) have become a powerful modelling tool for complex-valued data processing. Because most of the critical points of CVNNs are saddle points, the gradient-based learning algorithms for CVNNs enjoy more chances to reach the global minima while suffering from slow convergence. To this end, we propose a hybrid complex spectral conjugate gradient learning algorithm for fast training CVNNs in this paper. The proposed algorithm combines the scaled negative gradient with a Barzilai–Borwein stepsize and an optimized conjugate term to define a new training direction, thus providing an accurate approximation of the second-order curvature of the objective function. The complex Wolfe conditions are employed to adaptively determine the optimal training stepsize. Under mild conditions, the descent property of the training direction and the convergence of the proposed algorithm are theoretically established. Simulation results on a number of benchmark complex-valued data processing problems demonstrate the efficiency of the proposed algorithm.

EAAI Journal 2024 Journal Article

Change detection on multi-sensor imagery using mixed interleaved group convolutional network

  • Kun Tan
  • Moyang Wang
  • Xue Wang
  • Jianwei Ding
  • Zhaoxian Liu
  • Chen Pan
  • Yong Mei

The difference of the spatio-spectral features of multi-sensor image causes big difficulty in change detection because of the difficulties of the feature extraction. Unlike the traditional approaches that mainly relying on manually feature design, the advances of deep learning-based methods in deep feature extraction provide new alternatives for multi-sensor imagery change detection. Specifically, the incorporation of multi-scale information from remote sensing images holds paramount importance in change detection, consistently applied in the design of various deep learning models. This study investigated a change detection approach utilizing a mixed interleaved group convolutional network (MIGCNet) on multi-sensor remote sensing imagery, with a specific focus on fine-grained kernel space and multi-scale feature analysis within convolution operations. The proposed MIGCNet, with parallel branches as the fundamental architecture, can distinguish the change information effectively by the proposed mixed interleaved group convolution (MIGC) module, which combined mixed convolution with interleaved group convolution. Meanwhile, multi-loss supervision is utilized to promote the performance of the proposed MIGCNet. Experimental results demonstrate the outperformance of the MIGCNet to handle change detection with multi-sensor images on urban area. Considering different datasets, the Overall Accuracy and Kappa Coefficient are reaching 0. 97 and 80. 67%, respectively, and the miss detection rate and the false alarm rate are as low as 0. 17 and 0. 18, respectively.

NeurIPS Conference 2024 Conference Paper

DeformableTST: Transformer for Time Series Forecasting without Over-reliance on Patching

  • Donghao Luo
  • Xue Wang

With the proposal of patching technique in time series forecasting, Transformerbased models have achieved compelling performance and gained great interest fromthe time series community. But at the same time, we observe a new problem thatthe recent Transformer-based models are overly reliant on patching to achieve idealperformance, which limits their applicability to some forecasting tasks unsuitablefor patching. In this paper, we intent to handle this emerging issue. Through divinginto the relationship between patching and full attention (the core mechanismin Transformer-based models), we further find out the reason behind this issueis that full attention relies overly on the guidance of patching to focus on theimportant time points and learn non-trivial temporal representation. Based on thisfinding, we propose DeformableTST as an effective solution to this emergingissue. Specifically, we propose deformable attention, a sparse attention mechanismthat can better focus on the important time points by itself, to get rid of the need ofpatching. And we also adopt a hierarchical structure to alleviate the efficiency issuecaused by the removal of patching. Experimentally, our DeformableTST achievesthe consistent state-of-the-art performance in a broader range of time series tasks, especially achieving promising performance in forecasting tasks unsuitable forpatching, therefore successfully reducing the reliance on patching and broadeningthe applicability of Transformer-based models. Code is available at this repository: https: //github. com/luodhhh/DeformableTST.

AAAI Conference 2024 Conference Paper

MASTER: Market-Guided Stock Transformer for Stock Price Forecasting

  • Tong Li
  • Zhaoyang Liu
  • Yanyan Shen
  • Xue Wang
  • Haokun Chen
  • Sen Huang

Stock price forecasting has remained an extremely challenging problem for many decades due to the high volatility of the stock market. Recent efforts have been devoted to modeling complex stock correlations toward joint stock price forecasting. Existing works share a common neural architecture that learns temporal patterns from individual stock series and then mixes up temporal representations to establish stock correlations. However, they only consider time-aligned stock correlations stemming from all the input stock features, which suffer from two limitations. First, stock correlations often occur momentarily and in a cross-time manner. Second, the feature effectiveness is dynamic with market variation, which affects both the stock sequential patterns and their correlations. To address the limitations, this paper introduces MASTER, a MArkert-guided Stock TransformER, which models the momentary and cross-time stock correlation and leverages market information for automatic feature selection. MASTER elegantly tackles the complex stock correlation by alternatively engaging in intra-stock and inter-stock information aggregation. Experiments show the superiority of MASTER compared with previous works and visualize the captured realistic stock correlation to provide valuable insights.

ICLR Conference 2024 Conference Paper

ModernTCN: A Modern Pure Convolution Structure for General Time Series Analysis

  • Donghao Luo 0002
  • Xue Wang

Recently, Transformer-based and MLP-based models have emerged rapidly and won dominance in time series analysis. In contrast, convolution is losing steam in time series tasks nowadays for inferior performance. This paper studies the open question of how to better use convolution in time series analysis and makes efforts to bring convolution back to the arena of time series analysis. To this end, we modernize the traditional TCN and conduct time series related modifications to make it more suitable for time series tasks. As the outcome, we propose ModernTCN and successfully solve this open question through a seldom-explored way in time series community. As a pure convolution structure, ModernTCN still achieves the consistent state-of-the-art performance on five mainstream time series analysis tasks while maintaining the efficiency advantage of convolution-based models, therefore providing a better balance of efficiency and performance than state-of-the-art Transformer-based and MLP-based models. Our study further reveals that, compared with previous convolution-based models, our ModernTCN has much larger effective receptive fields (ERFs), therefore can better unleash the potential of convolution in time series analysis. Code is available at this repository: https://github.com/luodhhh/ModernTCN.

AIIM Journal 2024 Journal Article

Non-invasive fractional flow reserve derived from reduced-order coronary model and machine learning prediction of stenosis flow resistance

  • Yili Feng
  • Ruisen Fu
  • Hao Sun
  • Xue Wang
  • Yang Yang
  • Chuanqi Wen
  • Yaodong Hao
  • Yutong Sun

Background and objective Recently, computational fluid dynamics enables the non-invasive calculation of fractional flow reserve (FFR) based on 3D coronary model, but it is time-consuming. Currently, machine learning technique has emerged as an efficient and reliable approach for prediction, which allows saving a lot of analysis time. This study aimed at developing a simplified FFR prediction model for rapid and accurate assessment of functional significance of stenosis. Methods A reduced-order lumped parameter model (LPM) of coronary system and cardiovascular system was constructed for rapidly simulating coronary flow, in which a machine learning model was embedded for accurately predicting stenosis flow resistance at a given flow from anatomical features of stenosis. Importantly, the LPM was personalized in both structures and parameters according to coronary geometries from computed tomography angiography and physiological measurements such as blood pressure and cardiac output for personalized simulations of coronary pressure and flow. Coronary lesions with invasive FFR ≤ 0. 80 were defined as hemodynamically significant. Results A total of 91 patients (93 lesions) who underwent invasive FFR were involved in FFR derived from machine learning (FFRML) calculation. Of the 93 lesions, 27 lesions (29. 0%) showed lesion-specific ischemia. The average time of FFRML simulation was about 10 min. On a per-vessel basis, the FFRML and FFR were significantly correlated (r = 0. 86, p < 0. 001). The diagnostic accuracy, sensitivity, specificity, positive predictive value and negative predictive value were 91. 4%, 92. 6%, 90. 9%, 80. 6% and 96. 8%, respectively. The area under the receiver-operating characteristic curve of FFRML was 0. 984. Conclusion In this selected cohort of patients, the FFRML improves the computational efficiency and ensures the accuracy. The favorable performance of FFRML approach greatly facilitates its potential application in detecting hemodynamically significant coronary stenosis in future routine clinical practice.

EAAI Journal 2024 Journal Article

Research on reflective clothing recognition algorithm based on combining omni-dimensional dynamic convolution and partial convolution

  • Wenbi Ma
  • Zheng Guan
  • Xue Wang
  • Zhuqing Zhang
  • Jinde Cao

Currently, in construction sites, road maintenance, airports, and other special scenarios, the process of checking whether workers are wearing reflective clothing for safety is overly reliant on manual operations, and this manual screening method is not only inefficient but also has huge labor costs. To address this problem, this paper proposes a new method for reflective clothing wear recognition. Firstly, by replacing some traditional convolutions in the neck network of the YOLOv7-tiny(You Only Look Once vertion 7 - tiny) algorithm with the ODConv(Omni-dimensional Dynamic Convolution) module, the four dimensions of the kernel space can be endowed with convolutional dynamics attributes, which improves the detection accuracy of the model. Secondly, the PConv(Partial Convolution) module is used to replace some other traditional convolutions in the neck network, aiming to ensure detection accuracy while reducing computational redundancy and memory access. Then, a new SPPC(Spatial Pyramid Pooling Curtail) module is proposed and replaces the SPPCSPC(Spatial Pyramid Pooling Cross Stage Partial Concat) module of the original neck network, which guarantees accuracy and reduces the number of model parameters at the same time. Finally, the algorithm model proposed in this paper is ported to the Jetson Nano edge computing device, which can well meet the demand for real-time detection of reflective clothing and lay the foundation for subsequent practical applications.

NeurIPS Conference 2023 Conference Paper

One Fits All: Power General Time Series Analysis by Pretrained LM

  • Tian Zhou
  • Peisong Niu
  • Xue Wang
  • Liang Sun
  • Rong Jin

Although we have witnessed great success of pre-trained models in natural language processing (NLP) and computer vision (CV), limited progress has been made for general time series analysis. Unlike NLP and CV where a unified model can be used to perform different tasks, specially designed approach still dominates in each time series analysis task such as classification, anomaly detection, forecasting, and few-shot learning. The main challenge that blocks the development of pre-trained model for time series analysis is the lack of a large amount of data for training. In this work, we address this challenge by leveraging language or CV models, pre-trained from billions of tokens, for time series analysis. Specifically, we refrain from altering the self-attention and feedforward layers of the residual blocks in the pre-trained language or image model. This model, known as the Frozen Pretrained Transformer (FPT), is evaluated through fine-tuning on all major types of tasks involving time series. Our results demonstrate that pre-trained models on natural language or images can lead to a comparable or state-of-the-art performance in all main time series analysis tasks, as illustrated in Figure1. We also found both theoretically and empirically that the self-attention module behaviors similarly to principle component analysis (PCA), an observation that helps explains how transformer bridges the domain gap and a crucial step towards understanding the universality of a pre-trained transformer. The code is publicly available at https: //anonymous. 4open. science/r/Pretrained-LM-for-TSForcasting-C561.

NeurIPS Conference 2023 Conference Paper

OneNet: Enhancing Time Series Forecasting Models under Concept Drift by Online Ensembling

  • Yifan Zhang
  • Qingsong Wen
  • Xue Wang
  • Weiqi Chen
  • Liang Sun
  • Zhang Zhang
  • Liang Wang
  • Rong Jin

Online updating of time series forecasting models aims to address the concept drifting problem by efficiently updating forecasting models based on streaming data. Many algorithms are designed for online time series forecasting, with some exploiting cross-variable dependency while others assume independence among variables. Given every data assumption has its own pros and cons in online time series modeling, we propose **On**line **e**nsembling **Net**work (**OneNet**). It dynamically updates and combines two models, with one focusing on modeling the dependency across the time dimension and the other on cross-variate dependency. Our method incorporates a reinforcement learning-based approach into the traditional online convex programming framework, allowing for the linear combination of the two models with dynamically adjusted weights. OneNet addresses the main shortcoming of classical online learning methods that tend to be slow in adapting to the concept drift. Empirical results show that OneNet reduces online forecasting error by more than $\mathbf{50}\\%$ compared to the State-Of-The-Art (SOTA) method.

NeurIPS Conference 2022 Conference Paper

FiLM: Frequency improved Legendre Memory Model for Long-term Time Series Forecasting

  • Tian Zhou
  • Ziqing Ma
  • Xue Wang
  • Qingsong Wen
  • Liang Sun
  • Tao Yao
  • Wotao Yin
  • Rong Jin

Recent studies have shown that deep learning models such as RNNs and Transformers have brought significant performance gains for long-term forecasting of time series because they effectively utilize historical information. We found, however, that there is still great room for improvement in how to preserve historical information in neural networks while avoiding overfitting to noise present in the history. Addressing this allows better utilization of the capabilities of deep learning models. To this end, we design a Frequency improved Legendre Memory model, or FiLM: it applies Legendre polynomial projections to approximate historical information, uses Fourier projection to remove noise, and adds a low-rank approximation to speed up computation. Our empirical studies show that the proposed FiLM significantly improves the accuracy of state-of-the-art models in multivariate and univariate long-term forecasting by (19. 2%, 22. 6%), respectively. We also demonstrate that the representation module developed in this work can be used as a general plugin to improve the long-term prediction performance of other deep learning modules. Code is available at https: //github. com/tianzhou2011/FiLM/.

AAAI Conference 2022 Conference Paper

Scaled ReLU Matters for Training Vision Transformers

  • Pichao Wang
  • Xue Wang
  • Hao Luo
  • Jingkai Zhou
  • Zhipeng Zhou
  • Fan Wang
  • Hao Li
  • Rong Jin

Vision transformers (ViTs) have been an alternative design paradigm to convolutional neural networks (CNNs). However, the training of ViTs is much harder than CNNs, as it is sensitive to the training parameters, such as learning rate, optimizer and warmup epoch. The reasons for training difficulty are empirically analysed in the paper Early Convolutions Help Transformers See Better, and the authors conjecture that the issue lies with the patchify-stem of ViT models. In this paper, we further investigate this problem and extend the above conclusion: only early convolutions do not help for stable training, but the scaled ReLU operation in the convolutional stem (conv-stem) matters. We verify, both theoretically and empirically, that scaled ReLU in conv-stem not only improves training stabilization, but also increases the diversity of patch tokens, thus boosting peak performance with a large margin via adding few parameters and flops. In addition, extensive experiments are conducted to demonstrate that previous ViTs are far from being well trained, further showing that ViTs have great potential to be a better substitute of CNNs.

IJCAI Conference 2021 Conference Paper

Time Series Data Augmentation for Deep Learning: A Survey

  • Qingsong Wen
  • Liang Sun
  • Fan Yang
  • Xiaomin Song
  • Jingkun Gao
  • Xue Wang
  • Huan Xu

Deep learning performs remarkably well on many time series analysis tasks recently. The superior performance of deep neural networks relies heavily on a large number of training data to avoid overfitting. However, the labeled data of many real-world time series applications may be limited such as classification in medical time series and anomaly detection in AIOps. As an effective way to enhance the size and quality of the training data, data augmentation is crucial to the successful application of deep learning models on time series data. In this paper, we systematically review different data augmentation methods for time series. We propose a taxonomy for the reviewed methods, and then provide a structured review for these methods by highlighting their strengths and limitations. We also empirically compare different data augmentation methods for different tasks including time series classification, anomaly detection, and forecasting. Finally, we discuss and highlight five future directions to provide useful research guidance.

ICML Conference 2018 Conference Paper

Minimax Concave Penalized Multi-Armed Bandit Model with High-Dimensional Convariates

  • Xue Wang
  • Mike Mingcheng Wei
  • Tao Yao

In this paper, we propose a Minimax Concave Penalized Multi-Armed Bandit (MCP-Bandit) algorithm for a decision-maker facing high-dimensional data with latent sparse structure in an online learning and decision-making process. We demonstrate that the MCP-Bandit algorithm asymptotically achieves the optimal cumulative regret in sample size T, O(log T), and further attains a tighter bound in both covariates dimension d and the number of significant covariates s, O(s^2 (s + log d). In addition, we develop a linear approximation method, the 2-step Weighted Lasso procedure, to identify the MCP estimator for the MCP-Bandit algorithm under non-i. i. d. samples. Using this procedure, the MCP estimator matches the oracle estimator with high probability. Finally, we present two experiments to benchmark our proposed the MCP-Bandit algorithm to other bandit algorithms. Both experiments demonstrate that the MCP-Bandit algorithm performs favorably over other benchmark algorithms, especially when there is a high level of data sparsity or when the sample size is not too small.

YNIMG Journal 2016 Journal Article

Functional magnetic resonance imaging of the cervical spinal cord during thermal stimulation across consecutive runs

  • Kenneth A. Weber
  • Yufen Chen
  • Xue Wang
  • Thorsten Kahnt
  • Todd B. Parrish

The spinal cord is the first site of nociceptive processing in the central nervous system and has a role in the development and perpetuation of clinical pain states. Advancements in functional magnetic resonance imaging are providing a means to non-invasively measure spinal cord function, and functional magnetic resonance imaging may provide an objective method to study spinal cord nociceptive processing in humans. In this study, we tested the validity and reliability of functional magnetic resonance imaging using a selective field-of-view gradient-echo echo-planar-imaging sequence to detect activity induced blood oxygenation level-dependent signal changes in the cervical spinal cord of healthy volunteers during warm and painful thermal stimulation across consecutive runs. At the group and subject level, the activity was localized more to the dorsal hemicord, the spatial extent and magnitude of the activity was greater for the painful stimulus than the warm stimulus, and the spatial extent and magnitude of the activity exceeded that of a control analysis. Furthermore, the spatial extent of the activity for the painful stimuli increased across the runs likely reflecting sensitization. Overall, the spatial localization of the activity varied considerably across the runs, but despite this variability, a machine-learning algorithm was able to successfully decode the stimuli in the spinal cord based on the distributed pattern of the activity. In conclusion, we were able to successfully detect and characterize cervical spinal cord activity during thermal stimulation at the group and subject level.

YNIMG Journal 2016 Journal Article

Lateralization of cervical spinal cord activity during an isometric upper extremity motor task with functional magnetic resonance imaging

  • Kenneth A. Weber
  • Yufen Chen
  • Xue Wang
  • Thorsten Kahnt
  • Todd B. Parrish

The purpose of this study was to use an isometric upper extremity motor task to detect activity induced blood oxygen level dependent signal changes in the cervical spinal cord with functional magnetic resonance imaging. Eleven healthy volunteers performed six 5minute runs of an alternating left- and right-sided isometric wrist flexion task, during which images of the cervical spinal cord were acquired with a reduced field-of-view T2*-weighted gradient-echo echo-planar-imaging sequence. Spatial normalization to a standard spinal cord template was performed, and group average activation maps were generated in a mixed-effects analysis. The task activity significantly exceeded that of the control analyses. The activity was lateralized to the hemicord ipsilateral to the task and reliable across the runs at the group and subject level. Finally, a multi-voxel pattern analysis was able to successfully decode the left and right tasks at the C6 and C7 vertebral levels.

YNIMG Journal 2014 Journal Article

White matter microstructure changes induced by motor skill learning utilizing a body machine interface

  • Xue Wang
  • Maura Casadio
  • Kenneth A. Weber
  • Ferdinando A. Mussa-Ivaldi
  • Todd B. Parrish

The purpose of this study is to identify white matter microstructure changes following bilateral upper extremity motor skill training to increase our understanding of learning-induced structural plasticity and enhance clinical strategies in physical rehabilitation. Eleven healthy subjects performed two visuo-spatial motor training tasks over 9 sessions (2–3 sessions per week). Subjects controlled a cursor with bilateral simultaneous movements of the shoulders and upper arms using a body machine interface. Before the start and within 2days of the completion of training, whole brain diffusion tensor MR imaging data were acquired. Motor training increased fractional anisotropy (FA) values in the posterior and anterior limbs of the internal capsule, the corona radiata, and the body of the corpus callosum by 4. 19% on average indicating white matter microstructure changes induced by activity-dependent modulation of axon number, axon diameter, or myelin thickness. These changes may underlie the functional reorganization associated with motor skill learning.

YNIMG Journal 2008 Journal Article

Estimating Granger causality after stimulus onset: A cautionary note

  • Xue Wang
  • Yonghong Chen
  • Mingzhou Ding

How the brain processes sensory input to produce goal-oriented behavior is not well-understood. Advanced data acquisition technology in conjunction with novel statistical methods holds the key to future progress in this area. Recent studies have applied Granger causality to multivariate population recordings such as local field potential (LFP) or electroencephalography (EEG) in event-related paradigms. The aim is to reveal the detailed time course of stimulus-elicited information transaction among various sensory and motor cortices. Presently, interdependency measures like coherence and Granger causality are calculated on ongoing brain activity obtained by removing the average event-related potential (AERP) from each trial. In this paper we point out the pitfalls of this approach in light of the inevitable occurrence of trial-to-trial variability of event-related potentials in both amplitudes and latencies. Numerical simulations and experimental examples are used to illustrate the ideas. Special emphasis is placed on the important role played by single trial analysis of event-related potentials in experimentally establishing the main conclusion.

v2026.09.13