Arrow Research search

Author name cluster

Xianping Fu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

15 papers
2 author rows

Possible papers

15

AAAI Conference 2026 Conference Paper

Conditional Prompt Learning via Degradation Perception for Underwater Image Enhancement

  • Mingze Yao
  • Zhiying Jiang
  • Xianping Fu
  • Huibing Wang

Underwater Image Enhancement (UIE) focuses on improving visual quality from various underwater scenes. Existing methods simplistically treat various degradations as homogeneous, disregarding their intrinsic connections and causing models to blindly learn, resulting in conflicting optimization goals and visual distortions. To address above limitations, we propose a Conditional Prompt Learning via Degradation Perception (CPLDP) model, which employs conditional prompt as degradation perception priors and guides underwater image enhancement. Specifically, we show that the natural language prompts not only promote distinguishing different degraded images, but also aid in exploring more details with semantic information. Therefore, our method generates five key degradation prompts (green/blue/green-blue color casts, uneven illumination and haze) with conditional prompt learning. Subsequently, considering the intrinsic relationships among different degradations, we employ degradation perceptions as priors and fine-tune the learning strategy to enhance underwater images. During training, an adaptive loss function with multi-degradations is designed, allowing it to effectively handle the task conflicts among multiple underwater degradations. Additionally, we conduct a human visual-based underwater dataset with various degradation types by subjective statistics. Extensive experiments on both full-reference and non-reference datasets demonstrate that our CPLDP can achieve better visual results and outperforms state-of-the-art UIE methods across various degradation scenarios.

AAAI Conference 2026 Conference Paper

Instance-Guided Scene Adaptation for Unsupervised Person Search

  • Linfeng Qi
  • Huibing Wang
  • Jinjia Peng
  • Xianping Fu
  • Jiqing Zhang

Unsupervised Domain Adaptation (UDA) is a challenging task in person search. It adapts a well-trained model from a labeled source domain to an unlabeled target domain for privacy and efficiency. Currently, most of the state-of-the-art UDA person search methods adopt multi-scale feature alignment techniques to learn domain-invariant representations. However, person search is a multi-granularity task, and such an indiscriminate method of bridging the differences between domains misleads the identity learning process, which significantly limits the model's performance. In this paper, we propose an Instance-Guided Scene Adaptation (IGSA) framework by eradicating scene disparities and focusing the tasks on instances, effectively eliminating the contradiction between person search and domain adaptation. In IGSA, a Scene-Aware Bidirectional Filter (SABF) is designed to divide the image features into background and foreground to perform bidirectional modulations, thereby achieving simultaneous scene elimination and instance enhancement. To further improve the reliability of identity learning, we also propose an Instance Consistency Contrastive Learning (ICCL) method. By performing cross-epoch updates on the instance-level memory bank and re-initializing the cluster-level memory bank, the problem of inconsistent training across epochs caused by instance identity drift can be alleviated. Through the above designs, our method can achieve state-of-the-art performance on two benchmark datasets, with 82.1% mAP and 83.8% top-1 on the CUHK-SYSU dataset and 41.1% mAP and 82.3% top-1 on the PRW dataset, which is even better than some supervised methods.

IJCAI Conference 2025 Conference Paper

Consensus-Guided Incomplete Multi-view Clustering via Cross-view Affinities Learning

  • Qian Liu
  • Huibing Wang
  • Jinjia Peng
  • Yawei Chen
  • Mingze Yao
  • Xianping Fu
  • Yang Wang

Incomplete multi-view clustering (IMC) has garnered substantial attention due to its capacity to handle unlabeled data. Existing methods predominantly explore pairwise consistency between every two views. However, such consistency is highly susceptible to missing samples and outliers within a certain view and thus deviates from the true clustering distribution. Moreover, dual-view interaction neglects the collaboration effects of multiple views, making it challenging to capture the holistic characteristics across views. In response to these issues, we propose a novel Consensus-Guided Incomplete Multi-view Clustering via Cross-view Affinities Learning (CAL). Specifically, CAL reconstructs views with available instances to mine sample-wise affinities and harness comprehensive content information within views. Subsequently, to extract clean structural information, CAL imposes a structured sparse constraint on the representation tensor to eliminate biased errors. Furthermore, by integrating the consensus representation into a representation tensor, CAL can employ high-order interaction of multiple views to depict the semantic correlation between views while acquiring a unified structural graph across multiple views. Extensive experiments on seven benchmark datasets demonstrate that CAL outperforms some state-of-the-art methods in clustering performance. The code is available at https: //github. com/whbdmu/CAL.

IROS Conference 2025 Conference Paper

Data-Driven MPC for Attitude Control of Autonomous Underwater Robot

  • Tianzhu Gao
  • Yudong Luo
  • Na Zhao 0008
  • Jianda Wang
  • Yuanchu Yan
  • Xianping Fu
  • Xi Luo
  • Yantao Shen 0001

High maneuverability is essential to the autonomous operation of underwater robots. To achieve real-time maneuvering motion, the control strategy must take into account nonlinear hydrodynamic effects, which are extremely difficult to accurately capture during motion and therefore a balance must be struck between accuracy and real-time computational efficiency. Therefore, this paper proposes a data-driven approach to model the dynamics of the underwater robot using Sparse Identification of Nonlinear Dynamics (SINDy). Compared with existing works, our method does not require any physical prior knowledge and only uses a short period of onboard sensor data. Subsequently, the learned dynamic model is incorporated into a model predictive controller (MPC) to enable precise attitude control. Finally, the proposed method is implemented on our developed fully vectored propulsion underwater robot, and a series of attitude tracking experiments are conducted in an indoor water tank. Experimental results reveal that our approach significantly improves the model accuracy and reduces the attitude tracking errors by over 79% at a control frequency of 20 Hz, which proves the effectiveness and real-time performance of the method.

EAAI Journal 2025 Journal Article

Detail-focused and polarization-guided multi-modality fusion for underwater image clarity enhancing

  • Mingze Yao
  • Huibing Wang
  • Yudong Li
  • Wenzhe Liu
  • Xianping Fu

Underwater optics imaging typically suffer from the impurities scattering and light absorption in dynamic and complex underwater environment, which significantly effect the clarity and visibility of images. Existing cutting-edge Underwater Image Enhancement (UIE) methods mostly focus on color correction and contrast enhancement neglect the textual and detail information of objects, leading to imbalance exposure and edge features missing. To overcome these problems, we propose a novel detail-focused and polarization guided multi-modality fusion network (DFPG-Net), for enhancing underwater images. Unlike the previous methods, we first construct a Detail-Focused Convolution (DFC) block for extracting features from underwater multimodal images, which integrates difference convolutions to capture prior and edge information. Meanwhile, polarization information is introduced with a Multi-scale Polarization Guided (MPG) fusion module, which intends to maintain and enhance the texture and details information from the degree of polarization information and angle of polarization information obtained from the polarized image. Additionally, a parallel progressive attention network is designed to explore and combine the valuable and discriminative information in feature learning stage. Extensive experiments on the constructed underwater dataset validate the effectiveness and superior performance of the proposed DFPG-Net, which against state-of-the-art methods in both machine evaluation metrics and visual perception.

ICRA Conference 2025 Conference Paper

Efficient Cross-Boundary Grasping in Stacked Clutter with Single-Visual Mapping Multi-Step

  • Yudong Luo
  • Tong Wang
  • Feiyu Xie
  • Na Zhao 0008
  • Xianping Fu
  • Yantao Shen 0001

In logistics applications, the vision-based technology for grasping target objects in the air is relatively mature. However, when operating across the air and water such as grasping marine products from the water, the visual information collected by the camera will be disturbed by ripples and bubbles on the water surface, resulting in low grasping efficiency. Therefore, we introduce a grasping strategy based on single-visual mapping for multi-step (SVMMS) strategy to achieve cross-medium operations involving stacked objects. Specifically, we design a multifunctional integrated Deep Q-learning-based network model to extract visual features from the scene to effectively detect stacked objects and outputs their hierarchical relationships. Moreover, we quantify the underlying relationship between motion logic during action execution and changes in RGB-D during action execution to help the robot achieve efficient and collision-free operations. Our approach also incorporates a time-series design with prioritized experience replay to globally optimize the action sequence. Additionally, we propose a novel sim2real method by combining domain randomization to address the difference in object sizes between the simulation and the real world. Extensive experiments in both simulation and physical environments show that SVMMS-Grasp significantly outperforms existing methods in terms of task success rate, stability, and operational efficiency.

EAAI Journal 2024 Journal Article

A depth map stitching framework based on salient region matching

  • Zetian Mi
  • Haixia Qi
  • Jiaxin Chen
  • Yang Yu
  • Yujia Wang
  • Huibing Wang
  • Xianping Fu

Depth map stitching technology is urgently needed in varying fields such as panoramic three-dimensional reconstruction, underwater mapping, autonomous driving, and robot collision avoidance control. Existing image stitching research mainly focuses on three-primary-color images and heavily relies on feature detection quality, while depth map stitching is rarely studied due to the sparse features and low resolution. To address the above challenges, this paper proposes a color image-guided depth map stitching framework consisting of two stages: coarse depth map stitching and depth correction. In the first stage, a color-image-guided coarse depth map stitching network is proposed to accurately estimate the homography matrix, which can further eliminate parallax artifacts in the stitched depth map. In the second stage, a transformer-based depth correction network is designed, which combines saliency detection and regional matching strategies for depth correction, in order to solve the inconsistency in depth value introduced by point-to-point matching and improve the speed of depth correction to some extent. Extensive experiments demonstrate that the proposed method effectively solves the problem of disparity artifacts and mismatched depth values of the same target, and has excellent performance in depth map stitching.

IJCAI Conference 2024 Conference Paper

Fast One-Stage Unsupervised Domain Adaptive Person Search

  • Tianxiang Cui
  • Huibing Wang
  • Jinjia Peng
  • Ruoxi Deng
  • Xianping Fu
  • Yang Wang

Unsupervised person search aims to localize a particular target person from a gallery set of scene images without annotations, which is extremely challenging due to the unexpected variations of the unlabeled domains. However, most existing methods dedicate to developing multi-stage models to adapt domain variations while using clustering for iterative model training, which inevitably increase model complexity. To address this issue, we propose a Fast One-stage Unsupervised person Search (FOUS) which complementaryly integrates domain adaption with label adaption within an end-to-end manner without iterative clustering. To minimize the domain discrepancy, FOUS introduced an Attention-based Domain Alignment Module (ADAM) which can not only align various domains for both detection and ReID tasks but also construct an attention mechanism to reduce the adverse impacts of low-quality candidates resulting from unsupervised detection. Moreover, to avoid the redundant iterative clustering mode, FOUS adopts a prototype-guided labeling method which minimizes redundant correlation computations for partial samples and assigns noisy coarse label groups efficiently. The coarse label groups will be continuously refined via label-flexible training network with an adaptive selection strategy. With the adapted domains and labels, FOUS can achieve the state-of-the-art (SOTA) performance on two benchmark datasets, CUHK-SYSU and PRW. The code is available at https: //github. com/whbdmu/FOUS.

ICRA Conference 2024 Conference Paper

Model Predictive Control for an Autonomous Underwater Robot with Fully Vectored Propulsion

  • Tianzhu Gao
  • Yudong Luo
  • Chao Lv
  • Weirong Luo
  • Xianping Fu
  • Na Zhao 0008
  • Xi Luo
  • Yantao Shen 0001

Due to the low motion efficiency and maneuver-ability of underwater robots with six degrees of freedom, it is challenging for them to respond quickly to the attitude requirements during underwater autonomous manipulation. This paper presents a novel autonomous underwater robot with fully vectored propulsion and a model predictive control method to achieve more agile and efficient movements autonomously. In detail, we first design a robot with eight vector-distributed thruster layouts for fully vectored propulsion and construct the software architecture based on the robot operating system (ROS). Then, we establish the hydrodynamic model by adopting the Fossen approach and construct a 13-dimensional system state-space equation, which is discretized using the explicit fourth-order Runge-Kutta method. To achieve autonomous manipulation, model predictive control is employed along with physical constraints of the custom-built robot to enable real-time prediction and optimization of the robot’s states for control purposes. Finally, numerical simulations and experiments of the Point-to-Point Motion are conducted to test the robot’s performance. Experimental results reveal that the average error of each direction is 0. 0027 m, 0. 0031 m, and 0. 0368 m in the x-axis, y-axis, and z-axis, respectively, and 0. 8502°, 2. 1941°, 0. 2408° corresponding to three attitude angles, which verify the performance of employing MPC to control an autonomous underwater robot with fully vectored propulsion.

IJCAI Conference 2024 Conference Paper

Scene-Adaptive Person Search via Bilateral Modulations

  • Yimin Jiang
  • Huibing Wang
  • Jinjia Peng
  • Xianping Fu
  • Yang Wang

Person search aims to localize specific a target person from a gallery set of images with various scenes. As the scene of moving pedestrian changes, the captured person image inevitably bring in lots of background noise and foreground noise on the person feature, which are completely unrelated to the person identity, leading to severe performance degeneration. To address this issue, we present a Scene-Adaptive Person Search (SEAS) model by introducing bilateral modulations to simultaneously eliminate scene noise and maintain a consistent person representation to adapt to various scenes. In SEAS, a Background Modulation Network (BMN) is designed to encode the feature extracted from the detected bounding box into a multi-granularity embedding, which reduces the input of background noise from multiple levels with norm-aware. Additionally, to mitigate the effect of foreground noise on the person feature, SEAS introduces a Foreground Modulation Network (FMN) to compute the clutter reduction offset for the person embedding based on the feature map of the scene image. By bilateral modulations on both background and foreground within an end-to-end manner, SEAS obtains consistent feature representations without scene noise. SEAS can achieve state-of-the-art (SOTA) performance on two benchmark datasets, CUHK-SYSU with 97. 1% mAP and PRW with 60. 5% mAP. The code is available at https: //github. com/whbdmu/SEAS.

EAAI Journal 2023 Journal Article

Unsupervised underwater image enhancement via content-style representation disentanglement

  • Pengli Zhu
  • Yancheng Liu
  • Yuanquan Wen
  • Minyi Xu
  • Xianping Fu
  • Siyuan Liu

The absorption and scattering properties of the water medium cause various types of distortion in underwater images, which seriously affects the accuracy and effectiveness of subsequent processing. The application of supervised learning algorithms in underwater image enhancement is limited by the difficulty of obtaining a large number of underwater paired images in practical applications. As a solution, we propose an unsupervised representation disentanglement based underwater image enhancement method (URD-UIE). URD-UIE disentangles content information (e. g. , texture, semantics) and style information (e. g. , chromatic aberration, blur, noise, and clarity) from underwater images and then employs the disentangled information to generate the target distortion-free image. Our proposed method URD-UIE adopts an unsupervised cycle-consistent adversarial translation architecture and combines multiple loss functions to impose specific constraints on the output results of each module to ensure the structural consistency of underwater images before and after enhancement. The experimental results demonstrate that the URD-UIE technique effectively enhances the quality of underwater images when training with unpaired data, resulting in a significant improvement in the performance of the standard model for underwater object detection and semantic segmentation.

EAAI Journal 2022 Journal Article

Learning latent features with local channel drop network for vehicle re-identification

  • Xianping Fu
  • Jinjia Peng
  • Guangqi Jiang
  • Huibing Wang

Vehicle re-identification targets to find the target vehicle images in a large dataset which is composed of vehicle images from multiple non-overlapping cameras. Due to the various illumination, viewpoints and resolutions, it is challenging to find the right vehicle images accurately. Most existing works put emphasis on learning strong features by exploiting the attention parts in vehicle images, which leads to some small important cues being suppressed by these significant parts. Hence, a local channel drop network (LCDNet) is proposed in this paper, which focuses on seeking the latent features by releasing the constraint of most attentive features. Specially, besides the normal local feature learning network, LCDNet consists of an attentive local feature learning branch that drops some regions to promote learning the attentive features of local regions. Besides, the batch ranking loss is introduced to split the samples into two groups in a batch and regularize them by enforcing a margin, which ensures the model to learn meaningful features to distinct vehicles. Moreover, to further calculate the similarity of various images, the paper proposes a multi-distance based ranking method to achieve more accurate results. Experiments on several benchmark datasets validate the effectiveness of the proposed method.

EAAI Journal 2020 Journal Article

Purifying real images with an attention-guided style transfer network for gaze estimation

  • Xianping Fu
  • YuXiao Yan
  • Yang Yan
  • Jinjia Peng
  • Huibing Wang

Recently, the progress of learning-by-synthesis has proposed a training model for synthetic images, which can effectively reduce the cost of human and material resources. Image synthesis has been widely accepted as a cost effective way to learn models because it provides training sets that are large, diverse and accurately labeled. However, the realism of the synthetic image is not enough, this affects generalization on naturalistic test image. In an attempt to address this issue, previous methods learn a model to improve the realism of synthetic image. Different from previous methods, we take the first step towards purifying the real image to weaken the influence of light and convert the distribution of an outdoor naturalistic image through a real-time style transfer task to that of indoor synthetic image. In this paper, we first introduce the segmentation masks to construct Red, Green, and Blue-mask (RGB-mask) pairs as inputs, then we design an attention-guided style transfer network to learn style features separately from the attention and background regions, learn content features from full and attention regions. Moreover, we propose a novel region-level task-guided loss to restrain the features learnt from style and content. Experiments were performed using a mixed research (qualitative and quantitative) method to demonstrate the possibility of purifying real images in complex directions. We evaluate the proposed method on three public datasets, including Labeled pupils in the wild (LPW), Common Objects in COntext (COCO) and MPIIGaze. Extensive experimental results show that the proposed method is effective and achieves the state-of-the-art results.

IJCAI Conference 2020 Conference Paper

Unsupervised Vehicle Re-identification with Progressive Adaptation

  • Jinjia Peng
  • Yang Wang
  • Huibing Wang
  • Zhao Zhang
  • Xianping Fu
  • Meng Wang

Vehicle re-identification (reID) aims at identifying vehicles across different non-overlapping cameras views. The existing methods heavily relied on well-labeled datasets for ideal performance, which inevitably causes fateful drop due to the severe domain bias between the training domain and the real-world scenes; worse still, these approaches required full annotations, which is labor-consuming. To tackle these challenges, we propose a novel Progressive Adaptation Learning method for vehicle reID, named PAL, which infers from the abundant data without annotations. For PAL, a data adaptation module is employed for source domain, which generates the images with similar data distribution to unlabeled target domain as “pseudo target samples”. These pseudo samples are combined with the unlabeled samples that are selected by a dynamic sampling strategy to make training faster. We further proposed a weighted label smoothing (WLS) loss, which considers the similarity between samples with different clusters to balance the confidence of pseudo labels. Comprehensive experimental results validate the advantages of PAL on both VehicleID and VeRi-776 dataset.

EAAI Journal 2019 Journal Article

Purifying naturalistic images through a real-time style transfer semantics network

  • Tongtong Zhao
  • YuXiao Yan
  • Ibrahim Shehi Shehu
  • Xianping Fu
  • Huibing Wang

Recently, the progress of learning-by-synthesis has proposed a training model for synthetic images, which can effectively reduce the cost of human and material resources. However, due to the different distribution of synthetic images compared to real images, the desired performance cannot still be achieved. Real images consist of multiple forms of light orientation, while synthetic images consist of a uniform light orientation. These features are considered to be characteristic of outdoor and indoor scenes, respectively. To solve this problem, the previous method learned a model to improve the realism of the synthetic image. Different from the previous methods, this paper takes the first step to purify real images. Through the style transfer task, the distribution of outdoor real images is converted into indoor synthetic images, thereby reducing the influence of light. Therefore, this paper proposes a real-time style transfer network that preserves image content information (e. g. , gaze direction, pupil center position) of an input image (real image) while inferring style information (e. g. , image color structure, semantic features) of style image (synthetic image). In addition, the network accelerates the convergence speed of the model and adapts to multi-scale images. Experiments were performed using mixed studies (qualitative and quantitative) methods to demonstrate the possibility of purifying real images in complex directions. Qualitatively, it compares the proposed method with the available methods in a series of indoor and outdoor scenarios of the LPW dataset. In quantitative terms, it evaluates the purified image by training a gaze estimation model on the cross data set. The results show a significant improvement over the baseline method compared to the raw real image.

v2026.09.13