Arrow Research search

Author name cluster

Xin Wen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

16 papers
2 author rows

Possible papers

16

EAAI Journal 2026 Journal Article

A consistency-driven pseudo-labeling framework for robust functional connectivity modeling in neuropsychiatric disorder diagnosis

  • Xin Wen
  • Shijie Guo
  • Li Dong
  • Xiaobo Liu
  • Wenbo Ning
  • Jie Shi
  • Songhua Liu
  • Cheng Luo

The incidence of neuropsychiatric disorders such as Autism Spectrum Disorder (ASD), Attention Deficit Hyperactivity Disorder (ADHD), and Major Depressive Disorder (MDD) continues to rise. Deep learning-based computer-aided diagnosis (CAD) has emerged as a promising approach to alleviate the increasing burden on neuroimaging-based clinical resources. However, neuroimaging modalities such as functional magnetic resonance imaging (fMRI) involve complex spatiotemporal characteristics, making their representations susceptible to various types of noise and interference, which in turn hampers the effectiveness of CAD. To address this challenge, we propose a pseudo-label consistency-driven framework for functional connectivity (FC) reconstruction and discriminative modeling (PL-FCDM), aiming to enhance both the representational quality and discriminative power of FC features. Specifically, two complementary pseudo-labeling models are developed to independently capture discriminative features from the temporal domain (time series) and spatial domain (dynamic functional connectivity), enabling pseudo label prediction from distinct modalities. Then a consistency-based filtering strategy is applied to construct high-confidence reconstructed functional connectivity. These graphs are subsequently fed into a classification model comprising a Feature Optimization Autoencoder and a Depthwise Separable Convolutional Neural Network for efficient identification of neuropsychiatric disorders. Extensive experiments conducted on four publicly available multi-site datasets—ABIDE I, ABIDE II, ADHD-200, and REST-meta-MDD demonstrate that the proposed method achieves classification accuracies of 76. 14%, 74. 37%, 72. 89%, and 71. 15%, respectively. These results consistently outperform several state-of-the-art approaches, validating the effectiveness and robustness of the proposed framework in feature refinement and multi-disorder recognition.

JBHI Journal 2026 Journal Article

Diagnosis of Major Depressive Disorder Based on Multi-Granularity Brain Networks Fusion

  • Mengni Zhou
  • Rongkun Mi
  • Ang Zhao
  • Xin Wen
  • Yan Niu
  • Xubin Wu
  • Yanqing Dong
  • Yaru Xu

Major Depressive Disorder (MDD) is a common mental disorder, and making an early and accurate diagnosis is crucial for effective treatment. Functional Connectivity Network (FCN) constructed based on functional Magnetic Resonance Imaging (fMRI) have demonstrated the potential to reveal the mechanisms underlying brain abnormalities. Deep learning has been widely employed to extract features from FCN, but existing methods typically operate directly on the network, failing to fully exploit their deep information. Although graph coarsening techniques offer certain advantages in extracting the brain’s complex structure, they may also result in the loss of critical information. To address this issue, we propose the Multi-Granularity Brain Networks Fusion (MGBNF) framework. MGBNF models brain networks through multi-granularity analysis and constructs combinatorial modules to enhance feature extraction. Finally, the Constrained Attention Pooling (CAP) mechanism is employed to achieve the effective integration of multi-channel features. In the feature extraction stage, the parameter sharing mechanism is introduced and applied to multiple channels to capture similar connectivity patterns between different channels while reducing the number of parameters. We validate the effectiveness of the MGBNF model on multiple classification tasks and various brain atlases. The results demonstrate that MGBNF outperforms baseline models in terms of classification performance. Ablation experiments further validate its effectiveness. In addition, we conducted a thorough analysis of the variability of different subtypes of MDD by multiple classification tasks, and the results support further clinical applications.

JBHI Journal 2026 Journal Article

SSDiff: A Contrast-Free Virtual LGE Generator for Acute Myocardial Infarction with Joint Segmentation via Diffusion Model

  • Jing Qi
  • Xiuzheng Yue
  • Miao Hu
  • Xin Wen
  • Yinyin Chen
  • Hang Jin
  • Chengyan Wang
  • Tao Li

Myocardial infarction (MI) remains a major cause of death and disability. Although late gadolinium enhancement (LGE) cardiac MRI is the reference for assessing myocardial viability, it requires contrast injection, complex protocols, and added cost. Prior virtual LGE approaches-mostly GAN-based-mainly use cine or T1 mapping and ignore T2-weighted short-tau inversion recovery (T2-STIR), which is highly sensitive to edema in acute MI. They also typically require manual post-hoc delineation of infarcts. We propose SSDiff ( S ynthesis joint S egmentation Diff usion), a multitask conditional diffusion framework that synthesizes contrast-free virtual LGE from routine cine + T2-STIR for acute infarct assessment and simultaneously segments myocardium, ventricular blood pool, and infarct. SSDiff introduces a feature-disentangled attention module that isolates sequence-specific cues to steer the diffusion process, and a cross-fusion module that aligns synthesis and segmentation decoders for mutual optimization. Evaluated on a multi-center, multi-vendor dataset of 409 subjects (2, 177 aligned cine-T2-STIR-LGE triplets), SSDiff yields significant gains in synthetic image quality and downstream segmentation accuracy over strong baselines. Beyond serving as a clinically feasible alternative when LGE is unavailable or contraindicated, SSDiff also generates paired image-mask samples that augment LGE-scarce training, highlighting its practical utility and translational potential. Code is available at: https://github.com/QijingGJ/SSDiff.

EAAI Journal 2025 Journal Article

A framework for super-resolution of side-scan sonar images: Combination of variational Bayes and regional feature selection

  • Xin Wen
  • Chensheng Cheng
  • Lu Li
  • Feihu Zhang
  • Guang Pan

Side-scan sonar is widely used in ocean exploration due to its broad search range and strong identification capabilities. However, the inherent characteristics of acoustic images often result in poor image quality, negatively impacting subsequent downstream tasks’ accuracy. Image super-resolution (SR) technology based on deep learning technology is employed to address this issue. Despite this, existing SR models face two main challenges when applied to side-scan sonar images: (1) less data in side-scan sonar images causes the model overfitting problem; (2) less effective features in side-scan sonar images cause lower efficiency. To overcome these challenges, this paper proposes a deep learning framework that integrates a Bayesian structure with region-based feature selection. First, we introduce a rolling region selection method to extract key features of interest from side-scan sonar images, enhancing efficiency without compromising quality. Additionally, we replace traditional Convolutional Neural Networks (CNN) with Variational Bayes Convolutional Neural Networks (VB-CNN) to perform the SR task, improving generalization on small datasets and mitigating the risk of overfitting. Experiments conducted on the Side-Scan Sonar Visual Object Classes (SSS-VOC) dataset and other datasets demonstrate our proposed approach’s effectiveness through both qualitative and quantitative comparisons.

YNIMG Journal 2025 Journal Article

A practical measure of integrated information reveals alpha-band activity and the posterior cortex as neural correlates of arousal

  • Xin Wen
  • Yu Chang
  • Sijie Li
  • Jing Wang
  • Xiaoli Li
  • Duan Li
  • Changwei Wei
  • Zhenhu Liang

The search for neurophysiological markers of consciousness and their neural substrates remains a focal point in neuroscience research. The integrated information theory (IIT) provides a promising quantitative framework for consciousness assessment, but computational limitations of existing Φ estimation methods hinder an in-depth understanding of large-scale cortical integration. Here, we proposed a new measure, Φ c o p u l a, by incorporating the Gaussian copula approach for estimating integrated information. Simulation analysis demonstrated that Φ c o p u l a significantly outperformed common estimators, maintaining the lowest bias and mean squared error (MSE) even in non-Gaussian high-dimensional systems. We applied Φ c o p u l a to electroencephalographic data across different arousal states: awake, propofol-induced unresponsiveness, and non-rapid eye movement (NREM) sleep. Results revealed that alpha-band Φ c o p u l a significantly decreased during both propofol anesthesia (p < 0. 001) and sleep (p < 0. 014) states. Moreover, classification analysis demonstrated that Φ c o p u l a -based classifiers achieved superior accuracy in distinguishing arousal states compared to functional connectivity and network efficiency measures (p < 0. 030 for anesthesia; p < 0. 043 for sleep). Among the functional networks, the dorsal attention network (DAN) and default mode network (DMN) contributed most to Φ c o p u l a. Among the anatomical brain regions, the cingulate and posterior cortices showed the greatest contributions. Our findings suggest that Φ c o p u l a is a practical and effective metric for quantifying integrated information, with substantial potential for monitoring arousal levels in clinical and experimental settings. The posterior cortex, especially the posterior cingulate cortex (PCC), shows the greatest contribution to arousal-related information integration, revealing its critical role in consciousness.

ICRA Conference 2025 Conference Paper

Generalizing Motion Planners with Mixture of Experts for Autonomous Driving

  • Qiao Sun 0001
  • Huimin Wang
  • Jiahao Zhan
  • Fan Nie
  • Xin Wen
  • Leimeng Xu
  • Kun Zhan
  • Peng Jia 0007

Large real-world driving datasets have sparked significant research into various aspects of learning-based motion planners for autonomous driving. These include data augmentation, model architecture, reward design, training strategies, and planner pipelines. In this paper, we review and benchmark previous methods. Experiments show that many of these approaches have limited generalization abilities in planning performance due to overly complex designs or training paradigms. Experiments further reveal that as models are appropriately scaled, many designs become redundant. Therefore, we introduce StateTransformer-2 (STR2), a scalable, decoder-only motion planner. STR2uses a Vision Transformer (ViT) encoder and a mix-of-experts (MoE) causal transformer architecture. The MoE backbone addresses modality collapse and reward balancing by expert routing during training. Extensive experiments on the NuPlan dataset show that our method generalizes better than previous approaches across different test sets and closed-loop simulations. We evaluate its scalability on billions of real-world urban driving scenarios, demonstrating consistent accuracy improvements as both data and model size grow.

NeurIPS Conference 2025 Conference Paper

Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation

  • Zheng Anlin
  • Xin Wen
  • Xuanyang Zhang
  • Chuofan Ma
  • Tiancai Wang
  • Gang Yu
  • Xiangyu Zhang
  • Xiaojuan Qi

In this work, we present a novel direction to build an image tokenizer directly on top of a frozen vision foundation model, which is a largely underexplored area. Specifically, we employ a frozen vision foundation model as the encoder of our tokenizer. To enhance its effectiveness, we introduce two key components: (1) a region-adaptive quantization framework that reduces redundancy in the pre-trained features on regular 2D grids, and (2) a semantic reconstruction objective that aligns the tokenizer’s outputs with the foundation model’s representations to preserve semantic fidelity. Based on these designs, our proposed image tokenizer, \textbf{\ours}, achieves substantial improvements in image reconstruction and generation quality, while also enhancing token efficiency. It further boosts autoregressive (AR) generation---achieving a gFID of \textbf{1. 36} on ImageNet benchmarks, while accelerating model convergence by \textbf{three times}, and enabling high-fidelity class-conditional synthesis without the need for classifier-free guidance (CFG). The code is available at \href{https: //github. com/CVMI-Lab/VFMTok}{https: //github. com/CVMI-Lab/VFMTok}.

EAAI Journal 2024 Journal Article

Multi-scale context feature and cross-attention network-enabled system and software-based for pavement crack detection

  • Xin Wen
  • Shuo Li
  • Hao Yu
  • Yu He

Pavement crack detection continues to be a stubborn problem given the interference of various factors in the actual pavement and the complex topological structure of asphalt pavement. Among all the obstacles, the bottleneck of pavement crack detection lies in the difficulty of segmenting the cracks in the pavement images whose edges are blurred. This paper proposes a multi-scale context feature and cross-attention based on convolutional neural network for accurate and robust pavement crack segmentation. The multi-scale context feature module is built in different deep networks to extract rich crack feature information. Subsequently, in order to effectively promote the seamless integration of features at different levels, we deploy cross-attention modules to each branch. After that, we add deep supervision to each branch to accelerate training. Finally, we integrate the outputs of each branch to obtain the final output diagram. The comparative experiments on various pavement datasets show that the method has better robustness. At the same time, this paper designs a complete system of pavement crack detection (PCD) and develops corresponding engineering application software. The PCD system can record the real-time pavement image data to the edge server, and the client can also monitor the real-time pavement images from the edge server through the HTTP protocol.

NeurIPS Conference 2024 Conference Paper

What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable Insights

  • Xin Wen
  • Bingchen Zhao
  • Yilun Chen
  • Jiangmiao Pang
  • Xiaojuan Qi

Severe data imbalance naturally exists among web-scale vision-language datasets. Despite this, we find CLIP pre-trained thereupon exhibits notable robustness to the data imbalance compared to supervised learning, and demonstrates significant effectiveness in learning generalizable representations. With an aim to investigate the reasons behind this finding, we conduct controlled experiments to study various underlying factors, and reveal that CLIP's pretext task forms a dynamic classification problem wherein only a subset of classes is present in training. This isolates the bias from dominant classes and implicitly balances the learning signal. Furthermore, the robustness and discriminability of CLIP improve with more descriptive language supervision, larger data scale, and broader open-world concepts, which are inaccessible to supervised learning. Our study not only uncovers the mechanisms behind CLIP's generalizability beyond data imbalance but also provides transferable insights for the research community. The findings are validated in both supervised and self-supervised learning, enabling models trained on imbalanced data to achieve CLIP-level performance on diverse recognition tasks. Code and data are available at: https: //github. com/CVMI-Lab/clip-beyond-tail.

EAAI Journal 2023 Journal Article

Adoption of energy consumption in urban mobility considering digital carbon footprint: A two-phase interval-valued Fermatean fuzzy dominance methodology

  • Jeevaraj S.
  • Ilgin Gokasar
  • Muhammet Deveci
  • Dursun Delen
  • Bilal Bahaa Zaidan
  • Xin Wen
  • Wen-Long Shang
  • Gang Kou

Interval-valued Fermatean fuzzy sets play a significant role in modelling decision-making problems with incomplete information more accurately than intuitionistic fuzzy sets. Various decision-making methods have been introduced for the different classes IFSs. In this study, we aim to introduce a novel two-phase interval-valued Fermatean fuzzy dominance method which suits the decision-making problems modelled under the IVFFS environment well and study its applications in the adoption of energy consumption in Urban mobility considering digital carbon footprint. The proposed method considers the importance and performance of one alternative with respect to all others, which is not the case with many available decision-making algorithms introduced in the literature. Transportation is one of the most significant sources of global greenhouse gas (GHG) emissions. Numerous potential remedies are proposed to reduce the quantity of GHG generated by transportation activities, including regulatory measures and public transit digitalization initiatives. Decision-makers, however, should consider the digital carbon footprint of such projects. This study proposes three alternatives for reducing GHG emissions from transportation activities: incremental adoption of digital technologies to reduce energy consumption and greenhouse gases, disruptive digitalization technologies in urban mobility, and redesign of urban mobility using regulatory approaches and economic instruments. The proposed novel two-phase interval-valued Fermatean fuzzy dominance method will be utilized to rank these alternative projects in order of advantage. First, the problem is converted into a multi-criterion group decision-making problem. Then a novel two-phase interval-valued Fermatean fuzzy dominance method is designed and developed to rank the alternatives. The importance and advantage of the proposed two-phase method over other existing methods are discussed by using sensitivity and comparative analysis. The results indicate that rethinking urban mobility through governmental policies and economic tools is the least advantageous choice, while incremental adoption of digital technologies is the most advantageous.

NeurIPS Conference 2023 Conference Paper

CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection

  • Chuofan Ma
  • Yi Jiang
  • Xin Wen
  • Zehuan Yuan
  • Xiaojuan Qi

Deriving reliable region-word alignment from image-text pairs is critical to learnobject-level vision-language representations for open-vocabulary object detection. Existing methods typically rely on pre-trained or self-trained vision-languagemodels for alignment, which are prone to limitations in localization accuracy orgeneralization capabilities. In this paper, we propose CoDet, a novel approachthat overcomes the reliance on pre-aligned vision-language space by reformulatingregion-word alignment as a co-occurring object discovery problem. Intuitively, bygrouping images that mention a shared concept in their captions, objects corresponding to the shared concept shall exhibit high co-occurrence among the group. CoDet then leverages visual similarities to discover the co-occurring objects andalign them with the shared concept. Extensive experiments demonstrate that CoDethas superior performances and compelling scalability in open-vocabulary detection, e. g. , by scaling up the visual backbone, CoDet achieves 37. 0 $AP^m_{novel}$ and 44. 7 $AP^m_{all}$ on OV-LVIS, surpassing the previous SoTA by 4. 2 $AP^m_{novel}$ and 9. 8 $AP^m_{all}$. Code is available at https: //github. com/CVMI-Lab/CoDet.

AAAI Conference 2023 Conference Paper

KT-Net: Knowledge Transfer for Unpaired 3D Shape Completion

  • Zhen Cao
  • Wenxiao Zhang
  • Xin Wen
  • Zhen Dong
  • Yu-Shen Liu
  • Xiongwu Xiao
  • Bisheng Yang

Unpaired 3D object completion aims to predict a complete 3D shape from an incomplete input without knowing the correspondence between the complete and incomplete shapes. In this paper, we propose the novel KTNet to solve this task from the new perspective of knowledge transfer. KTNet elaborates a teacher-assistant-student network to establish multiple knowledge transfer processes. Specifically, the teacher network takes complete shape as input and learns the knowledge of complete shape. The student network takes the incomplete one as input and restores the corresponding complete shape. And the assistant modules not only help to transfer the knowledge of complete shape from the teacher to the student, but also judge the learning effect of the student network. As a result, KTNet makes use of a more comprehensive understanding to establish the geometric correspondence between complete and incomplete shapes in a perspective of knowledge transfer, which enables more detailed geometric inference for generating high-quality complete shapes. We conduct comprehensive experiments on several datasets, and the results show that our method outperforms previous methods of unpaired point cloud completion by a large margin. Code is available at https://github.com/a4152684/KT-Net.

NeurIPS Conference 2022 Conference Paper

Self-Supervised Visual Representation Learning with Semantic Grouping

  • Xin Wen
  • Bingchen Zhao
  • Anlin Zheng
  • Xiangyu Zhang
  • Xiaojuan Qi

In this paper, we tackle the problem of learning visual representations from unlabeled scene-centric data. Existing works have demonstrated the potential of utilizing the underlying complex structure within scene-centric data; still, they commonly rely on hand-crafted objectness priors or specialized pretext tasks to build a learning framework, which may harm generalizability. Instead, we propose contrastive learning from data-driven semantic slots, namely SlotCon, for joint semantic grouping and representation learning. The semantic grouping is performed by assigning pixels to a set of learnable prototypes, which can adapt to each sample by attentive pooling over the feature and form new slots. Based on the learned data-dependent slots, a contrastive objective is employed for representation learning, which enhances the discriminability of features, and conversely facilitates grouping semantically coherent pixels together. Compared with previous efforts, by simultaneously optimizing the two coupled objectives of semantic grouping and contrastive learning, our approach bypasses the disadvantages of hand-crafted priors and is able to learn object/group-level representations from scene-centric images. Experiments show our approach effectively decomposes complex scenes into semantic groups for feature learning and significantly benefits downstream tasks, including object detection, instance segmentation, and semantic segmentation. Code is available at: https: //github. com/CVMI-Lab/SlotCon.

YNIMG Journal 2021 Journal Article

Age-dependent cross frequency coupling features from children to adults during general anesthesia

  • Zhenhu Liang
  • Na Ren
  • Xin Wen
  • Haiwen Li
  • Hang Guo
  • Yaqun Ma
  • Zheng Li
  • Xiaoli Li

BACKGROUND: The frequency coupling characteristics in electroencephalogram (EEG) induced by anesthetics have been well studied in adults, but the investigation of the age-dependent cross frequency coupling features from children to adults is still lacking. METHODS: We analyzed EEG signals recorded from pediatric to adult patients (n = 131), separated into six age groups: <1 year (n = 15), 1-3 years (n = 23), 3-6 years (n = 19), 6-12 years (n = 18), 12-18 years (n = 16), and 18-45 years (n = 40). Age related EEG power and cross frequency coupling analysis (phase amplitude coupling (PAC) and quadratic phase coupling) of data from maintenance of a surgical state of anesthesia (MOSSA) was conducted. Also, for patients of ages less than 6 years, we analyzed the performance of cross frequency coupling derived indices in distinguishing the states of wakefulness, MOSSA, and recovery of consciousness (ROC). RESULTS: (1) During MOSSA, EEG power substantially increased with age from infancy to 3-6 years then decreased with age in the theta-gamma frequency bands. The infant group (<1 year) had the highest slow oscillation (SO) power among all age groups. (2) The distinct PAC pattern is absent in patients less than 1 year of age both in SO-alpha and delta-alpha frequency band coupling during propofol induced unconsciousness. The modulation index between delta and alpha oscillations in MOSSA increased with age. (3) Wavelet bicoherence derived indices reach their peaks in the 3-6 years group and then decrease with age growth. (4) The Diag_En index (normalized entropy of the diagonal bicoherence entries of the bicoherence matrix) performed the best at distinguishing different states for ages less than 6 years (p<0.05). CONCLUSIONS: The combination of propofol induction and sevoflurane maintenance exhibited age-dependent EEG power spectra, PAC, and bicoherence, likely related to brain development. These observations suggest new rules for infant and child brain state monitoring during general anesthesia are needed.

YNIMG Journal 2021 Journal Article

WeBrain: A web-based brainformatics platform of computational ecosystem for EEG big data analysis

  • Li Dong
  • Jianfu Li
  • Qiunan Zou
  • Yufan Zhang
  • Lingling Zhao
  • Xin Wen
  • Jinnan Gong
  • Fali Li

The current evolution of 'cloud neuroscience' leads to more efforts with the large-scale EEG applications, by using EEG pipelines to handle the rapidly accumulating EEG data. However, there are a few specific cloud platforms that seek to address the cloud computational challenges of EEG big data analysis to benefit the EEG community. In response to the challenges, a WeBrain cloud platform (https://webrain.uestc.edu.cn/) is designed as a web-based brainformatics platform and computational ecosystem to enable large-scale EEG data storage, exploration and analysis using cloud high-performance computing (HPC) facilities. WeBrain connects researchers from different fields to EEG and multimodal tools that have become the norm in the field and the cloud processing power required to handle those large EEG datasets. This platform provides an easy-to-use system for novice users (even no computer programming skills) and provides satisfactory maintainability, sustainability and flexibility for IT administrators and tool developers. A range of resources are also available on https://webrain.uestc.edu.cn/, including documents, manuals, example datasets related to WeBrain, and collected links to open EEG datasets and tools. It is not necessary for users or administrators to install any software or system, and all that is needed is a modern web browser, which reduces the technical expertise required to use or manage WeBrain. The WeBrain platform is sponsored and driven by the China-Canada-Cuba international brain cooperation project (CCC-Axis, http://ccc-axis.org/), and we hope that WeBrain will be a promising cloud brainformatics platform for exploring brain information in large-scale EEG applications in the EEG community.

JBHI Journal 2020 Journal Article

A Feasible Feature Extraction Method for Atrial Fibrillation Detection From BCG

  • Xin Wen
  • Yanqi Huang
  • Xiaomei Wu
  • Biyong Zhang

Atrial fibrillation (AF) is the most frequently occurring form of arrhythmia, which induces multiple fatal diseases and impairs the quality of life in patients; thus, the study of the diagnostic methods for detecting AF is clinically important. Here, we present a feature extraction method for the detection of AF using a ballistocardiogram (BCG), which is based on a physiological signal database collected by a non-contact sensor. The BCG signals, including both with AF and sinus rhythm (SR), were collected from 37 subjects during overnight sleep (approximately 8 h). The signals were split into 2915 1-min segments (AF: 1494, SR: 1421) without overlap and labeled as AF and SR. BCG signals were transformed into BCG energy signals in order to highlight the features of AF and SR BCG signals; and four new data sequences representing different characteristics of the BCG energy signals were generated. The mean value, variance, skewness, and kurtosis of the four data sequences were calculated and 16 features were extracted for each segment. Five machine learning algorithms were used for classification. The results of this study show that the support vector machine performed the best among the five tested classifiers and achieved sensitivity, precision, and accuracy of 0. 968, 0. 928, and 0. 945, respectively. These results indicate that the proposed feature extraction method can be well applied to AF and SR classification and may lay foundations for the development of systems for long-term home cardiac monitoring and AF screening.

v2026.09.13