Arrow Research search

Author name cluster

Xiaolei Xu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
2 author rows

Possible papers

6

EAAI Journal 2025 Journal Article

Eviformer: An uncertainty fault diagnosis framework guided by evidential deep learning

  • Jingjie Luo
  • Fucai Li
  • Xiaolei Xu
  • Wenqiang Zhao
  • Dongqing Zhang

Unpredictable signals are commonly encountered during equipment operation, and existing deep learning-based fault diagnosis methods often fail to accurately evaluate the uncertainty of diagnostic results, limiting the model's capacity to respond to unexpected signals. To address this challenge, this study introduces a modified Transformer model based on deep evidential learning—Eviformer. The proposed model first employs a Swin-Transformer (ST)-based distribution projector, which preserves the advantages of ST in extracting features from vibration signals and simultaneously projects these signals directly into a Dirichlet distribution with second-order probabilities. Furthermore, by incorporating novel evidence correction terms and a constraint factor to reconstruct the evidence constraint loss, more precise uncertainty quantification in diagnostic predictions is achieved. Using a gear-bearing vibration dataset, comparative experiments were conducted across various scenarios, including out-of-distribution gear faults, faults in unmonitored components, noise interference, and variable speed conditions. The results demonstrate that the proposed method can promptly issue uncertainty-based warnings when encountering vibration signals that significantly differ from the training set distribution, thereby offering essential support for maintenance decisions.

ICLR Conference 2025 Conference Paper

Frame-Voyager: Learning to Query Frames for Video Large Language Models

  • Sicheng Yu
  • Chengkai Jin
  • Huanyu Wang
  • Zhenghao Chen
  • Sheng Jin
  • Zhongrong Zuo
  • Xiaolei Xu
  • Zhenbang Sun

Video Large Language Models (Video-LLMs) have made remarkable progress in video understanding tasks. However, they are constrained by the maximum length of input tokens, making it impractical to input entire videos. Existing frame selection approaches, such as uniform frame sampling and text-frame retrieval, fail to account for the information density variations in the videos or the complex instructions in the tasks, leading to sub-optimal performance. In this paper, we propose Frame-Voyager that learns to query informative frame combinations, based on the given textual queries in the task. To train Frame-Voyager, we introduce a new data collection and labeling pipeline, by ranking frame combinations using a pre-trained Video-LLM. Given a video of M frames, we traverse its T-frame combinations, feed them into a Video-LLM, and rank them based on Video-LLM's prediction losses. Using this ranking as supervision, we train Frame-Voyager to query the frame combinations with lower losses. In experiments, we evaluate Frame-Voyager on four Video Question Answering benchmarks by plugging it into two different Video-LLMs. The experimental results demonstrate that Frame-Voyager achieves impressive results in all settings, highlighting its potential as a plug-and-play solution for Video-LLMs.

TMLR Journal 2025 Journal Article

TempFlex: Advancing MLLMs with Temporal Perception and Natively Scalable Resolution Encoding

  • Zhanyu Wang
  • Chen Tang
  • Haoyu He
  • Kuan Feng
  • Chao Wang
  • Bingni Zhang
  • Xiaolei Xu
  • SHEN WANG

Multimodal large language models (MLLMs) have made significant progress across vision-language tasks, yet many designs still suffer from two core limitations. (i) Excessive visual tokens and broken global context: Tiled Patch Encoding fragments high-resolution images, leading to token overload and disrupting global attention modeling. (ii) Lack of temporal reasoning: Most models process video as independent frames using static image encoders, failing to capture temporal dynamics. We present TempFlex-VL, a token-efficient and temporally aware MLLM that addresses both issues through lightweight architectural enhancements. First, we introduce a resolution-agnostic visual encoder that directly processes full images without tiling, preserving global context while substantially reducing visual tokens. Second, we propose Temporal Fiber Fusion (TFF), a plug-and-play module with three complementary pathways: (1) a dynamic local-convolution branch for fine-grained motion, (2) a gated memory accumulator for long-term dependencies, and (3) a periodic encoder for modeling cyclic patterns. These signals are softly fused, enabling the model to adapt to diverse temporal structures without overfitting. To support large-scale video-language pretraining, we curate TempFlex-2M, a high-quality synthetic video–text corpus generated in a single stage via GPT-4o with direct visual prompting. We instantiate TempFlex-VL using two different language backbones, Gemma3-4B and Qwen3-4B, demonstrating the generality of our design across architectures. Both variants achieve state-of-the-art or competitive results on a wide range of image and video benchmarks while markedly improving token efficiency. Code is publicly available at: https://github.com/wang-zhanyu/TempFlex.

YNICL Journal 2021 Journal Article

Disorder- and emotional context-specific neurofunctional alterations during inhibitory control in generalized anxiety and major depressive disorder

  • Congcong Liu
  • Jing Dai
  • Yuanshu Chen
  • Ziyu Qi
  • Fei Xin
  • Qian Zhuang
  • Xinqi Zhou
  • Feng Zhou

Major Depressive Disorder (MDD) and Generalized Anxiety Disorder (GAD) are highly debilitating and often co-morbid disorders. The disorders exhibit partly overlapping dysregulations on the behavioral and neurofunctional level. The determination of disorder-specific behavioral and neurofunctional dysregulations may therefore promote neuro-mechanistic and diagnostic specificity. In order to determine disorder-specific alterations in the domain of emotion-cognition interactions the present study examined emotional context-specific inhibitory control in treatment-naïve MDD (n = 37) and GAD (n = 35) patients and healthy controls (n = 35). On the behavioral level MDD but not GAD exhibited impaired inhibitory control irrespective of emotional context. On the neural level, MDD-specific attenuated recruitment of inferior/medial parietal, posterior frontal, and mid-cingulate regions during inhibitory control were found during the negative context. GAD exhibited a stronger engagement of the left dorsolateral prefrontal cortex relative to MDD. Overall the findings from the present study suggest disorder- and emotional context-specific behavioral and neurofunctional inhibitory control dysregulations in major depression and may point to a depression-specific neuropathological and diagnostic marker.

YNIMG Journal 2021 Journal Article

Segregating domain-general from emotional context-specific inhibitory control systems - ventral striatum and orbitofrontal cortex serve as emotion-cognition integration hubs

  • Qian Zhuang
  • Lei Xu
  • Feng Zhou
  • Shuxia Yao
  • Xiaoxiao Zheng
  • Xinqi Zhou
  • Jialin Li
  • Xiaolei Xu

Inhibitory control hierarchically regulates cognitive and emotional systems in the service of adaptive goal-directed behavior across changing task demands and environments. While previous studies convergently determined the contribution of prefrontal-striatal systems to general inhibitory control, findings on the specific circuits that mediate emotional context-specific impact on inhibitory control remained inconclusive. Against this background we combined an evaluated emotional Go/No Go task with fMRI in a large cohort of subjects (N=250) to segregate brain systems and circuits that mediate domain-general from emotion-specific inhibitory control. Particularly during a positive emotional context, behavioral results showed a lower accuracy for No Go trials and a faster response time for Go trials. While the dorsal striatum and lateral frontal regions were involved in inhibitory control irrespective of emotional context, activity in the ventral striatum (VS) and medial orbitofrontal cortex (mOFC) varied as a function of emotional context. On the voxel-wise whole-brain network level, limbic and striatal systems generally exhibited highest changes in global brain connectivity during inhibitory control, while global brain connectivity of the left mOFC was less decreased during emotional contexts. Functional connectivity analyses moreover revealed that negative coupling between the VS with inferior frontal gyrus (IFG)/insula and mOFC varied as a function of emotional context. Together these findings indicate separable domain- general as well as emotional context-specific inhibitory brain systems which specifically encompass the VS and its connections with frontal regions.

YNIMG Journal 2016 Journal Article

Voluntary control of anterior insula and its functional connections is feedback-independent and increases pain empathy

  • Shuxia Yao
  • Benjamin Becker
  • Yayuan Geng
  • Zhiying Zhao
  • Xiaolei Xu
  • Weihua Zhao
  • Peng Ren
  • Keith M. Kendrick

Real-time functional magnetic resonance imaging (rtfMRI)-assisted neurofeedback (NF) training allows subjects to acquire volitional control over regional brain activity. Emerging evidence suggests its potential clinical utility as an effective non-invasive treatment approach in mental disorders. The therapeutic potential of rtfMRI-NF training depends critically upon whether: (1) acquired self-regulation produces functionally relevant changes at behavioral and brain network levels and (2) training effects can be maintained in the absence of feedback. To address these key questions, the present study combined rtfMRI-NF training for acquiring volitional anterior insula (AI) regulation with a sham-controlled between-subject design. The functional relevance of acquired AI control was assessed using both behavioral (pain empathy) and neural (activity, functional connectivity) indices. Maintenance of training effects in the absence of feedback was assessed two days later. During successful acquisition of volitional AI up-regulation subjects exhibited stronger empathic responses, increased AI-prefrontal coupling in circuits involved in learning and emotion regulation and increased resting state connectivity within AI-centered empathy networks. At follow-up both self-regulation and increased connectivity in empathy networks were fully maintained, although without further increases in empathy ratings. Overall these findings support the potential clinical application of rtfMRI-NF for inducing functionally relevant and lasting changes in emotional brain circuitry.

v2026.09.13