Arrow Research search

Author name cluster

Xian Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

7 papers
2 author rows

Possible papers

7

AAAI Conference 2026 Conference Paper

Equivariant Atomic and Lattice Modeling Using Geometric Deep Learning for Crystal Structure Optimization

  • Ziduo Yang
  • Yi-Ming Zhao
  • Xian Wang
  • Wei Zhuo
  • Xiaoqing Liu
  • Lei Shen

Structure optimization, which yields the relaxed structure (minimum‑energy state), is essential for reliable materials property calculations, yet traditional ab initio approaches such as density‑functional theory (DFT) are computationally intensive. Machine learning (ML) has emerged to alleviate this bottleneck but suffers from two major limitations: (i) existing models operate mainly on atoms, leaving lattice vectors implicit despite their critical role in structural optimization; and (ii) they often rely on multi-stage, non-end-to-end workflows that are prone to error accumulation. Here, we present E³Relax, an end-to-end equivariant graph neural network that maps an unrelaxed crystal directly to its relaxed structure. E³Relax promotes both atoms and lattice vectors to graph nodes endowed with dual scalar–vector features, enabling unified and symmetry‑preserving modeling of atomic displacements and lattice deformations. A layer‑wise supervision strategy forces every network depth to make a physically meaningful refinement, mimicking the incremental convergence of DFT while preserving a fully end‑to‑end pipeline. We evaluate E³Relax on four benchmark datasets and demonstrate that it achieves remarkable accuracy and efficiency. Through DFT validations, we show that the structures predicted by E³Relax are energetically favorable, making them suitable as high-quality initial configurations to accelerate DFT calculations.

EAAI Journal 2026 Journal Article

Multikernel correntropy transfer robust dictionary learning and its application in bearing fault diagnosis

  • Chuliang Liu
  • Zhiwei Luo
  • Zhonghe Huang
  • Yanwei Sang
  • Xian Wang

Detecting subtle fault signatures in vibration signals, masked by intense, non-Gaussian noise, poses a major challenge for the early diagnosis of bearing faults. This paper presents a novel multikernel correntropy transfer robust dictionary learning (MKC-TRDL) framework designed to address the challenges in bearing fault diagnosis. MKC-TRDL incorporates a multikernel correntropy-based data fidelity term, specifically crafted to minimize the impact of outliers, thereby ensuring more robust fault feature extraction. Furthermore, a transfer regularization term is introduced to guide the target dictionary to remain closely aligned with the source dictionary, striking an effective balance between preserving general signal features and adapting to the specific operating conditions of the bearing. This approach significantly enhances the robustness of the approach and its capacity to perform reliably in dynamic and noisy environments. Simulations and experimental results show that the MKC-TRDL method effectively extracts early bearing fault features, particularly in the presence of strong complex noise.

AIIM Journal 2026 Journal Article

PreLora: A fine-tuning approach with low-rank matrix decomposition and prefix tuning for pre-hospital emergency text classification

  • Feng Tian
  • Xian Wang
  • Saicong Lu
  • Jiaxuan Gu
  • Shitao Zhou
  • Penghui Li
  • Zhen Wang
  • Zengjun Jin

Objective With expanding applications of artificial intelligence technology in the medical field, Large Language Models (LLMs) have achieved substantial success in medical text processing. However, there remain a number of challenges in effectively adapting to specific tasks, such as pre-hospital emergency text classification. Methods We propose a novel fine-tuning method PreLora, which combines prefix tuning with matrix low-rank decomposition. First, this approach incorporates task-specific prompts based on multi-layer perceptron (MLP) encoder into the input. Then, it inject trainable rank-decomposed matrices into every layer of the transformer architecture to compress model parameters, reduce the number of parameters, and capture correlations among the input. To validate its efficacy, we carried out a comparative validation on a pre-hospital emergency text dataset. Results Comparison results indicated that the model fine-tuned with PreLora outperformed the baseline models without fine-tuning, achieving a performance improvement of 45. 4%–75. 4%. Moreover, PreLora ranked first among all fine-tuning methods across each LLM evaluated. An in-depth performance analysis was further conducted on 21 ICD-10 categories with distinct semantic features. The results revealed a negative correlation between model performance and semantic similarity of ICD-10 categories: the low similarity groups performed better, while the high similarity groups performed worse. Notably, PreLora consistently maintained robust performance, with a smaller performance decline in high-similarity categories compared to other fine-tuning methods. In the classifying complex cases with high semantic similarity, PreLora still showed superior adaptability, improving by 68. 6%–95. 8% compared to the baseline model and 0. 4%–8. 4% compared to other fine-tuning methods. Conclusion This study demonstrates PreLora is an effective fine-tuning method to process pre-hospital emergency text classification. It has the potential to expand to other mainstream models for adapting specific tasks in the medical field.

ICRA Conference 2025 Conference Paper

Dashing for the Golden Snitch: Multi-Drone Time-Optimal Motion Planning with Multi-Agent Reinforcement Learning

  • Xian Wang
  • Jin Zhou
  • Yuanli Feng
  • Jiahao Mei
  • Jiming Chen 0001
  • Shuo Li

Recent innovations in autonomous drones have facilitated time-optimal flight in single-drone configurations, and enhanced maneuverability in multi-drone systems by applying optimal control and learning-based methods. However, few studies have achieved time-optimal motion planning for multi-drone systems, particularly during highly agile maneuvers or in dynamic scenarios. This paper presents a decentralized policy network using multi-agent reinforcement learning for time-optimal multi-drone flight. To strike a balance between flight efficiency and collision avoidance, we introduce a soft collision-free mechanism inspired by optimization-based methods. By customizing PPO in a centralized training, decentralized execution (CTDE) fashion, we unlock higher efficiency and stability in training while ensuring lightweight implementation. Extensive simulations show that, despite slight performance tradeoffs compared to single-drone systems, our multi-drone approach maintains near-time-optimal performance with a low collision rate. Real-world experiments validate our method, with two quadrotors using the same network as in simulation achieving a maximum speed of 13. 65 m/s and a maximum body rate of 13. 4 rad/s in a 5. 5 m × 5. 5 m × 2. 0 m space across various tracks, relying entirely on onboard computation [video 3 3 https://youtu.be/KACuFMtGGpo][code 4 4 https://github.com/KafuuChikai/Dashing-for-the-Golden-Snitch-Multi-Drone-RL].

NeurIPS Conference 2025 Conference Paper

MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query

  • Wei Chow
  • Yuan Gao
  • Linfeng Li
  • Xian Wang
  • Qi Xu
  • Hang Song
  • Lingdong Kong
  • Ran Zhou

Semantic retrieval is crucial for modern applications yet remains underexplored in current research. Existing datasets are limited to single languages, single images, or singular retrieval conditions, often failing to fully exploit the expressive capacity of visual information as evidenced by maintained performance when images are replaced with captions. However, practical retrieval scenarios frequently involve interleaved multi-condition queries with multiple images. Hence, this paper introduces MERIT, the first multilingual dataset for interleaved multi-condition semantic retrieval, comprising 320, 000 queries with 135, 000 products in 5 languages, covering 7 distinct product categories. Extensive experiments on MERIT identify existing models's critical limitation: focusing solely on global semantic information while neglecting specific conditional elements in queries. Consequently, we propose Coral, a novel fine-tuning framework that adapts pre-trained MLLMs by integrating embedding reconstruction to preserve fine-grained conditional elements and contrastive learning to extract comprehensive global semantics. Experiments demonstrate that Coral achieves a 45. 9% performance improvement over conventional approaches on MERIT, with strong generalization capabilities validated across 8 established retrieval benchmarks. Collectively, our contributions—a novel dataset, identification of critical limitations in existing approaches, and an innovative fine-tuning framework—establish a foundation for future research in interleaved multi-condition semantic retrieval. Data & Code: MERIT-2025. github. io

NeurIPS Conference 2025 Conference Paper

Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models

  • zhentao he
  • Can Zhang
  • Ziheng Wu
  • Zhenghao Chen
  • Yufei Zhan
  • Yifan Li
  • Zhao Zhang
  • Xian Wang

Recent advancements in multimodal large language models (MLLMs) have enhanced document understanding by integrating textual and visual information. However, existing models exhibit incompleteness within their paradigm in real-world scenarios, particularly under visual degradation (e. g. , blur, occlusion, low contrast). In such conditions, the current response paradigm often fails to adequately perceive visual degradation and ambiguity, leading to overreliance on linguistic priors or misaligned visual-textual reasoning. This difficulty in recognizing uncertainty frequently results in the generation of hallucinatory content, especially when a precise answer is not feasible. To better demonstrate and analyze this phenomenon and problem, we propose KIE-HVQA, the first benchmark dedicated to evaluating OCR hallucination in degraded document understanding. This dataset includes test samples spanning identity cards, invoices, and prescriptions, with simulated real-world degradations and pixel-level annotations for OCR reliability. This setup allows for evaluating models' capacity, under degraded input, to distinguish reliable visual information and answer accordingly, thereby highlighting the challenge of avoiding hallucination on uncertain data. To achieve vision-faithful reasoning and thereby avoid the aforementioned issues, we further introduce a Group Relative Policy Optimization (GRPO)-based framework featuring a novel reward mechanism. By incorporating a self-awareness of visual uncertainty and an analysis method that initiates refusal to answer to increase task difficulty within our supervised fine-tuning and reinforcement learning framework, we successfully mitigated hallucinations in ambiguous regions. Experiments on Qwen2. 5-VL demonstrate that our 7B-parameter model achieves a ~28% absolute improvement in hallucination-free accuracy over GPT-4o on KIE-HVQA and there is no significant performance drop in standard tasks, highlighting both effectiveness and robustness. This work advances the development of reliable MLLMs for real-world document analysis by addressing critical challenges in visual-linguistic alignment under degradation.

EAAI Journal 2023 Journal Article

Yaw system restart strategy optimization of wind turbines in mountain wind farms based on operational data mining and multi-objective optimization

  • Jialu Han
  • Xian Wang
  • Xuebing Yang
  • Qihui Ling
  • Wei Liu

The wind resources of mountain wind farms are affected by more complex terrain than are found at flat wind farms. Wind turbine (WT) failures caused by frequent operation of the yaw system occur often in mountain wind farms, resulting in significantly reduced economic benefits over the WT life cycle. Considering the problem of frequent yaw motions of WTs in mountain wind farms, this study introduces multi-objective optimization theory into the yaw system restart strategy optimization of WTs for the first time and proposes an optimized yaw system restart strategy with wind speed segmentation in the wind speed region below the rated wind speed. The yaw behavior of mountain wind farms in southern China is analyzed from multiple perspectives; a segmentation scheme of wind speed intervals for yaw control was determined from the results. The Non-dominated Sorting Genetic Algorithm-II (NSGA-II) is used to optimize the control parameters for each sub-interval with the goals of minimizing yaw operation time and maximizing energy generation. A set of Pareto solutions is obtained and evaluated with the Technique for Order of Preference by Similarity to an Ideal Solution (TOPSIS) method to produce the optimal solution. The simulation results indicated that after the WT adopts the optimized yaw system restart strategy in the wind speed region below the rated wind speed, the number of yaw motions is decreased by 82. 9%, with only a slight decrease in energy generation. The optimization method proposed in this study can significantly reduce yaw operation time with minimal loss of energy generation, which provides a new approach to balancing the relationship between reducing the WT failure rate and ensuring power generation efficiency and is expected to significantly improve the economic benefits to wind farms.

v2026.09.27