Arrow Research search

Author name cluster

Xuan Wu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

14 papers
2 author rows

Possible papers

14

AAAI Conference 2026 Conference Paper

Efficient Few-Step Solution Generation via Discrete Flow Matching for Combinatorial Optimization

  • Yuanshu Li
  • Di Wang
  • Wei Du
  • Xuan Wu
  • Peng Zhao
  • Yubin Xiao
  • You Zhou

Combinatorial optimization problems (COPs) are fundamental to many real-world applications where efficiently producing high-quality solutions is critical. Recent advances in diffusion-based non-autoregressive models have reformulated solving COPs as a generative process, achieving promising results. However, almost all of these methods still suffer from accumulated errors and high inference costs due to the multi-step stochastic denoising process. To address these issues, we propose EFLOCO, an efficient discrete flow matching method for solving COPs, learning structured and deterministic solution trajectories. EFLOCO replaces noise-driven updates with smooth and guided transitions, thereby improves inference stability and quality. Furthermore, we introduce an adaptive time-step scheduler that makes more efforts in critical transition regions, yielding strong performance under few-step constraints. Experiments on standard Traveling Salesman Problems (TSPs) and Asymmetric TSPs (ATSPs) show that our method consistently outperforms both learning-based and heuristic baselines in terms of solution quality and inference speed.

IJCAI Conference 2025 Conference Paper

DGL: Dynamic Global-Local Information Aggregation for Scalable VRP Generalization with Self-Improvement Learning

  • Yubin Xiao
  • Yuesong Wu
  • Rui Cao
  • Di Wang
  • Zhiguang Cao
  • Xuan Wu
  • Peng Zhao
  • Yuanshu Li

The Vehicle Routing Problem (VRP) is a critical combinatorial optimization problem with wide-reaching real-world applications, particularly in logistics, transportation. While neural network-based VRP solvers have shown impressive results on test instances similar to training data, their performance often degrades when faced with varying scales and unseen distributions, limiting their practical applicability. To overcome these limitations, we introduce DGL (Dynamic Global-Local Information Aggregation), a novel model that combines global and local information to effectively solve VRPs. DGL dynamically adjusts local node selections within a localized range, capturing local invariance across problems of different scales and distributions, thereby enhancing generalization. At the same time, DGL integrates global context into the decision-making process, providing richer information for more informed decisions. Additionally, we propose a replacement-based self-improvement learning framework that leverages data augmentation and random replacement techniques, further enhancing DGL's robustness. Extensive experiments on synthetic datasets, benchmark datasets, and real-world country map instances demonstrate that DGL achieves state-of-the-art performance, particularly in generalizing to large-scale VRPs and real-world scenarios. These results showcase DGL's effectiveness in solving complex, realistic optimization challenges and highlight its potential for practical applications.

IROS Conference 2025 Conference Paper

L-SNI: A Language-Driven Semantic Navigation System for Inspection Tasks

  • Jiawang Ma
  • Weichen Guo
  • Xuan Wu
  • Zinan Zhuang
  • Rongxiang Zeng
  • Yongliang Shi
  • Gang Ma 0008

For inspection robots to achieve generalizability, stability, and ease of use, it is crucial that they understand natural language commands and navigate accurately to specified target objects. We propose L-SNI, a semantic navigation system adapted for inspection tasks, offering generalizability, robust stability, and practical ease of use. In the perception phase, L-SNI constructs a precise geometric depth map of the environment using LiDAR, while RGB images are employed to extract object categories, which are then combined with depth data to generate a semantic map. To enable the large language model (LLM) to interpret the environment, L-SNI encodes the 3D semantic map into a plain text representation. During single-task execution, L-SNI decodes human commands into inspection primitives using an LLM constrained by system initial prompts. These inspection primitives guide the robot’s low-level planner for task execution. To address the challenge of traditional 3D LiDAR localization and navigation systems in accurately positioning the robot around target objects during inspection tasks, we propose a target cost gradient to assist in optimizing the robot’s target point selection and attitude control in maps with semantic information. Upon reaching the target, L-SNI uses a visual language model (VLM) to describe the scene, which is simplified by the LLM into a user-friendly response. Through testing on 18 indoor scenes from the Matterport 3D dataset, L-SNI achieves a 46. 9% improvement in Success Rate (SR) and a 58. 3% increase in Success weighted by Path Length (SPL) over existing state-of-the-art (SOTA) solutions, while also demonstrating superior target image understanding. Moreover, it can be easily deployed on real-world robots without complex initialization.

ICLR Conference 2025 Conference Paper

Youku Dense Caption: A Large-scale Chinese Video Dense Caption Dataset and Benchmarks

  • Zixuan Xiong
  • Guangwei Xu
  • Wenkai Zhang
  • Yuan Miao
  • Xuan Wu
  • LinHai
  • Ruijie Guo
  • Hai-Tao Zheng

With the explosive growth of video content, video captions have emerged as a crucial tool for video comprehension, significantly enhancing the ability to understand and retrieve information from videos. However, most publicly available dense video captioning datasets are in English, resulting in a scarcity of large-scale and high-quality Chinese dense video captioning datasets. To address this gap within the Chinese community and to promote the advancement of Chinese multi-modal models, we develop the first, large-scale, and high-quality Chinese dense video captioning dataset, named Youku Dense Caption. This dataset is sourced from Youku, a prominent Chinese video-sharing website. Youku Dense Caption includes 31,466 complete short videos annotated by 311,921 Chinese captions. To the best of our knowledge, it is currently the largest publicly available dataset for fine-grained Chinese video descriptions. Additionally, we establish several benchmarks for Chinese video-language tasks based on the Youku Dense Caption, including retrieval, grounding, and generation tasks. Extensive experiments and evaluations are conducted on existing state-of-the-art multi-modal models, demonstrating the dataset's utility and the potential for further research.

EAAI Journal 2024 Journal Article

A behavior three-way decision approach under interval-valued triangular fuzzy numbers with application to the selection of additive manufacturing composites

  • Guoquan Xie
  • Wanying Zhu
  • Jiangyang Xiang
  • Tao Li
  • Xuan Wu
  • Yong Peng
  • Honghao Zhang
  • Kui Wang

Additive manufacturing composites, also recognized as three-dimensional (3D) printing composites, are highly anticipated for their potential to replace industrial materials due to the availability of multiple printing processes and optional materials. However, research gaps exist in cognitive deficiencies and psychological behaviors of decision-makers, as well as experimental error effects caused by material testing, resulting in material selection as a challenging issue. Therefore, this study proposes a novel behavior three-way decision model under the interval-valued triangular fuzzy number (IVTFN) to settle the selection issue of 3D printing composites. The research contributions are summarized as follows. First, the IVTFN is presented to account for the impacts of cognitive deficiency and experimental errors, based on which the concepts of information entropy and fuzzy measure are further developed to conduct the criterion weights. In addition, by integrating the prospect theory and regret theory, a framework for constructing the behavioral decision matrix is presented. Moreover, a novel behavior three-way decision model with the perspectives of objective and preference is proposed to classify the decision region. This study presents a comprehensive methodology integrating the three-way decision model and multi-criteria decision-making method to achieve both alternative ranking and alternative classifying. Finally, a research case of 3D printing composites reinforced by continuous hybrid fibers is adopted to illustrate the validity of the methodology. Comparative analysis and sensitivity analysis are also performed. This study offers valuable insights and tools for systematically tackling the 3D printing composite material selection issues.

ICRA Conference 2024 Conference Paper

An Onboard Framework for Staircases Modeling Based on Point Clouds

  • Chun Qing
  • Rongxiang Zeng
  • Xuan Wu
  • Yongliang Shi
  • Gan Ma

The detection of traversable regions on staircases and the physical modeling constitutes pivotal aspects of the mobility of legged robots. This paper presents an onboard framework tailored to the detection of traversable regions and the modeling of physical attributes of staircases by point cloud data. To mitigate the influence of illumination variations and the overfitting due to the dataset diversity, a series of data augmentations are introduced to enhance the training of the fundamental network. A curvature suppression cross-entropy(CSCE) loss is proposed to reduce the ambiguity of prediction on the boundary between traversable and non-traversable regions. Moreover, a measurement correction based on the pose estimation of stairs is introduced to calibrate the output of raw modeling that is influenced by tilted perspectives. Lastly, we collect a dataset pertaining to staircases and introduce new evaluation criteria. Through a series of rigorous experiments conducted on this dataset, we substantiate the superior accuracy and generalization capabilities of our proposed method. Codes, models, and datasets will be available at https://github.com/szturobotics/Stair-detection-and-modeling-project.

AAAI Conference 2024 Conference Paper

Distilling Autoregressive Models to Obtain High-Performance Non-autoregressive Solvers for Vehicle Routing Problems with Faster Inference Speed

  • Yubin Xiao
  • Di Wang
  • Boyang Li
  • Mingzhao Wang
  • Xuan Wu
  • Changliang Zhou
  • You Zhou

Neural construction models have shown promising performance for Vehicle Routing Problems (VRPs) by adopting either the Autoregressive (AR) or Non-Autoregressive (NAR) learning approach. While AR models produce high-quality solutions, they generally have a high inference latency due to their sequential generation nature. Conversely, NAR models generate solutions in parallel with a low inference latency but generally exhibit inferior performance. In this paper, we propose a generic Guided Non-Autoregressive Knowledge Distillation (GNARKD) method to obtain high-performance NAR models having a low inference latency. GNARKD removes the constraint of sequential generation in AR models while preserving the learned pivotal components in the network architecture to obtain the corresponding NAR models through knowledge distillation. We evaluate GNARKD by applying it to three widely adopted AR models to obtain NAR VRP solvers for both synthesized and real-world instances. The experimental results demonstrate that GNARKD significantly reduces the inference time (4-5 times faster) with acceptable performance drop (2-3%). To the best of our knowledge, this study is first-of-its-kind to obtain NAR VRP solvers from AR ones through knowledge distillation.

ICRA Conference 2024 Conference Paper

Research on bionic foldable wing for flapping wing micro air vehicle

  • Shengjie Xiao
  • Kai Hu 0004
  • Yuhong Sun
  • Yun Wang
  • Bo Qin
  • Huichao Deng
  • Xuan Wu
  • Xilun Ding

This paper presents a bionic foldable wing that imitates the hind wing of ladybirds. Based on the folding mechanism of the hind wing of ladybirds and the theory of origami, the motion model of the bionic foldable wing is established, yield the motion law of the crease angles and the variation relationship between the panels are obtained. Bionic foldable wings utilise shape memory alloy to drive wings to fold, and embedded torsion springs to release energy to realize the function of wing unfolding. In the experiments of the vehicle equipped with foldable wings, the lift and attitude torque of bionic foldable wings are measured by the F/T sensor. The experimental results indicated that its aerodynamic performance is basically close to that of our optimized non-foldable wings. Moreover, the vehicle with foldable wings has been able to overcome gravity to achieve flight, which provides a novel concept for the research on flapping wing.

EAAI Journal 2023 Journal Article

A hybrid multi-stage decision-making method with probabilistic interval-valued hesitant fuzzy set for 3D printed composite material selection

  • Guoquan Xie
  • Kui Wang
  • Xuan Wu
  • Jin Wang
  • Tao Li
  • Yong Peng
  • Honghao Zhang

The 3D printed composite material selection is of great interest due to its extensive application prospect and can be considered as a challenging multiple-criteria decision making (MCDM) issue. The hesitation and uncertainty of experts are difficult to measure, and the high degree of interaction among criteria is often overlooked in the decision process. In addition, composites are required to serve in harsh environments for various mechanical and industrial fields, resulting in the degradation of mechanical properties. In this study, a hybrid multi-stage decision-making method is developed to conduct 3D printed composite material selection in harsh environments. The theory of probabilistic interval-valued hesitant fuzzy set (PIVHFS) is proposed to characterize the decision-making information of experts, which can effectively quantify the assessment in uncertain environments. An integrated method that combines Choquet fuzzy integral and Shapley value is proposed to obtain the weight vector of criteria, which can reflect the mutual influence between criteria and their overall importance. The final decision-making result and the optimal alternative can be calculated by the PIVHFS-based Tomada de Decisão Interativa Multicritério and Technique for order preference by similarity to an ideal solution (TODIM-TOPSIS) method. An empirical application, i. e. , 3D printed composites in the background of automotive chassis, is applied to validate the application of the proposed method. Comparative analysis, sensitivity analysis, and managerial implications are also conducted to illustrate the validity of the method. This paper provides a valuable tool for addressing the material selection issue of 3D printed composites from a multi-criteria perspective.

NeurIPS Conference 2021 Conference Paper

Coresets for Clustering with Missing Values

  • Vladimir Braverman
  • Shaofeng Jiang
  • Robert Krauthgamer
  • Xuan Wu

We provide the first coreset for clustering points in $\mathbb{R}^d$ that have multiple missing values (coordinates). Previous coreset constructions only allow one missing coordinate. The challenge in this setting is that objective functions, like \kMeans, are evaluated only on the set of available (non-missing) coordinates, which varies across points. Recall that an $\epsilon$-coreset of a large dataset is a small proxy, usually a reweighted subset of points, that $(1+\epsilon)$-approximates the clustering objective for every possible center set. Our coresets for $k$-Means and $k$-Median clustering have size $(jk)^{O(\min(j, k))} (\epsilon^{-1} d \log n)^2$, where $n$ is the number of data points, $d$ is the dimension and $j$ is the maximum number of missing coordinates for each data point. We further design an algorithm to construct these coresets in near-linear time, and consequently improve a recent quadratic-time PTAS for $k$-Means with missing values [Eiben et al. , SODA 2021] to near-linear time. We validate our coreset construction, which is based on importance sampling and is easy to implement, on various real data sets. Our coreset exhibits a flexible tradeoff between coreset size and accuracy, and generally outperforms the uniform-sampling baseline. Furthermore, it significantly speeds up a Lloyd's-style heuristic for $k$-Means with missing values.

AAAI Conference 2020 Conference Paper

Multi-View Partial Multi-Label Learning with Graph-Based Disambiguation

  • Ze-Sen Chen
  • Xuan Wu
  • Qing-Guo Chen
  • Yao Hu
  • Min-Ling Zhang

In multi-view multi-label learning (MVML), each training example is represented by different feature vectors and associated with multiple labels simultaneously. Nonetheless, the labeling quality of training examples is tend to be affected by annotation noises. In this paper, the problem of multi-view partial multi-label learning (MVPML) is studied, where the set of associated labels are assumed to be candidate ones and only partially valid. To solve the MVPML problem, a two-stage graph-based disambiguation approach is proposed. Firstly, the ground-truth labels of each training example are estimated by disambiguating the candidate labels with fused similarity graph. After that, the predictive model for each label is learned from embedding features generated from disambiguation-guided clustering analysis. Extensive experimental studies clearly validate the effectiveness of the proposed approach in solving the MVPML problem.

AAAI Conference 2019 Conference Paper

CAFE: Adaptive VDI Workload Prediction with Multi-Grained Features

  • Yao Zhang
  • Wen-Ping Fan
  • Xuan Wu
  • Hua Chen
  • Bin-Yang Li
  • Min-Ling Zhang

Virtual desktop infrastructure (VDI) is a virtualization technology that hosts desktop operating system on centralized server in a data center of private or public cloud. Effective resource management is of crucial importance for VDI customers, where maintaining sufficient virtual machines helps guarantee satisfactory user experience while turning off spare virtual machines helps save running cost. Generally, existing techniques work in passive manner by either driving available capacity reactively or configuring management schedules manually. In this paper, a novel proactive resource management approach is proposed which aims to predict VDI pool workload adaptively by utilizing CoArse to Fine historical dEscriptive (CAFE) features. Specifically, aggregate session count from pool end users serves as the basis for workload measurement and predictive model induction. Extensive experiments on real VDI customers data sets clearly validate the effectiveness of multi-grained features for VDI workload prediction. Furthermore, practical insights identified in our VDI data analytics are also discussed.

IJCAI Conference 2019 Conference Paper

Multi-View Multi-Label Learning with View-Specific Information Extraction

  • Xuan Wu
  • Qing-Guo Chen
  • Yao Hu
  • Dengbao Wang
  • Xiaodong Chang
  • Xiaobo Wang
  • Min-Ling Zhang

Multi-view multi-label learning serves an important framework to learn from objects with diverse representations and rich semantics. Existing multi-view multi-label learning techniques focus on exploiting shared subspace for fusing multi-view representations, where helpful view-specific information for discriminative modeling is usually ignored. In this paper, a novel multi-view multi-label learning approach named SIMM is proposed which leverages shared subspace exploitation and view-specific information extraction. For shared subspace exploitation, SIMM jointly minimizes confusion adversarial loss and multi-label loss to utilize shared information from all views. For view-specific information extraction, SIMM enforces an orthogonal constraint w. r. t. the shared subspace to utilize view-specific discriminative information. Extensive experiments on real-world data sets clearly show the favorable performance of SIMM against other state-of-the-art multi-view multi-label learning approaches.

IJCAI Conference 2018 Conference Paper

Towards Enabling Binary Decomposition for Partial Label Learning

  • Xuan Wu
  • Min-Ling Zhang

The task of partial label (PL) learning is to learn a multi-class classifier from training examples each associated with a set of candidate labels, among which only one corresponds to the ground-truth label. It is well known that for inducing multi-class predictive model, the most straightforward solution is binary decomposition which works by either one-vs-rest or one-vs-one strategy. Nonetheless, the ground-truth label for each PL training example is concealed in its candidate label set and thus not accessible to the learning algorithm, binary decomposition cannot be directly applied under partial label learning scenario. In this paper, a novel approach is proposed to solving partial label learning problem by adapting the popular one-vs-one decomposition strategy. Specifically, one binary classifier is derived for each pair of class labels, where PL training examples with distinct relevancy to the label pair are used to generate the corresponding binary training set. After that, one binary classifier is further derived for each class label by stacking over predictions of existing binary classifiers to improve generalization. Experimental studies on both artificial and real-world PL data sets clearly validate the effectiveness of the proposed binary decomposition approach w. r. t state-of-the-art partial label learning techniques.

v2026.09.13