Arrow Research search

Author name cluster

Huibing Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

22 papers
1 author row

Possible papers

22

AAAI Conference 2026 Conference Paper

Conditional Prompt Learning via Degradation Perception for Underwater Image Enhancement

  • Mingze Yao
  • Zhiying Jiang
  • Xianping Fu
  • Huibing Wang

Underwater Image Enhancement (UIE) focuses on improving visual quality from various underwater scenes. Existing methods simplistically treat various degradations as homogeneous, disregarding their intrinsic connections and causing models to blindly learn, resulting in conflicting optimization goals and visual distortions. To address above limitations, we propose a Conditional Prompt Learning via Degradation Perception (CPLDP) model, which employs conditional prompt as degradation perception priors and guides underwater image enhancement. Specifically, we show that the natural language prompts not only promote distinguishing different degraded images, but also aid in exploring more details with semantic information. Therefore, our method generates five key degradation prompts (green/blue/green-blue color casts, uneven illumination and haze) with conditional prompt learning. Subsequently, considering the intrinsic relationships among different degradations, we employ degradation perceptions as priors and fine-tune the learning strategy to enhance underwater images. During training, an adaptive loss function with multi-degradations is designed, allowing it to effectively handle the task conflicts among multiple underwater degradations. Additionally, we conduct a human visual-based underwater dataset with various degradation types by subjective statistics. Extensive experiments on both full-reference and non-reference datasets demonstrate that our CPLDP can achieve better visual results and outperforms state-of-the-art UIE methods across various degradation scenarios.

AAAI Conference 2026 Conference Paper

Instance-Guided Scene Adaptation for Unsupervised Person Search

  • Linfeng Qi
  • Huibing Wang
  • Jinjia Peng
  • Xianping Fu
  • Jiqing Zhang

Unsupervised Domain Adaptation (UDA) is a challenging task in person search. It adapts a well-trained model from a labeled source domain to an unlabeled target domain for privacy and efficiency. Currently, most of the state-of-the-art UDA person search methods adopt multi-scale feature alignment techniques to learn domain-invariant representations. However, person search is a multi-granularity task, and such an indiscriminate method of bridging the differences between domains misleads the identity learning process, which significantly limits the model's performance. In this paper, we propose an Instance-Guided Scene Adaptation (IGSA) framework by eradicating scene disparities and focusing the tasks on instances, effectively eliminating the contradiction between person search and domain adaptation. In IGSA, a Scene-Aware Bidirectional Filter (SABF) is designed to divide the image features into background and foreground to perform bidirectional modulations, thereby achieving simultaneous scene elimination and instance enhancement. To further improve the reliability of identity learning, we also propose an Instance Consistency Contrastive Learning (ICCL) method. By performing cross-epoch updates on the instance-level memory bank and re-initializing the cluster-level memory bank, the problem of inconsistent training across epochs caused by instance identity drift can be alleviated. Through the above designs, our method can achieve state-of-the-art performance on two benchmark datasets, with 82.1% mAP and 83.8% top-1 on the CUHK-SYSU dataset and 41.1% mAP and 82.3% top-1 on the PRW dataset, which is even better than some supervised methods.

AAAI Conference 2026 Conference Paper

Localization-Anchored Instance Discrimination for Domain Adaptive Person Search

  • Linfeng Qi
  • Huibing Wang
  • Jinjia Peng
  • Jiqing Zhang

Domain-adaptive person search (DAPS) aims to transfer pedestrian detection and re-identification capabilities from a labeled source domain to an unlabeled target domain, yet faces critical challenges from domain shift: semantic confusion among overlapping instances, over-reliance on shallow features for look-alike targets, and poor discriminability of small-scale instances. To address these issues, we propose the Localization-Anchored Instance Discrimination (LAID) framework, which leverages spatial relationships between bounding boxes as auxiliary signals to enhance instance identity learning. LAID integrates three complementary strategies: 1) Cost-Aware Instance Matching (CAIM) uses IoU-based global optimal assignment to align current detections with historical identities, reducing overlap-induced misassociations; 2) Dual-Scope Contrastive Learning (DSCL) combines spatial separation constraints (for geometrically distant pairs) with global contrastive learning, prompting the model to learn deep discriminative features beyond superficial similarities; 3) Task-Sensitivity Alignment (TSA) aligns confidence distributions of detection and ReID heads via KL divergence, ensuring consistent pseudo-label generation. Extensive experiments on CUHK-SYSU and PRW datasets demonstrate that LAID outperforms state-of-the-art DAPS methods, validating its effectiveness in mitigating domain shift and narrowing the performance gap between supervised and domain-adaptive person search.

AAAI Conference 2026 Conference Paper

Prompting Adversarial Transferability via Path Flatness Attack

  • Zeze Tao
  • Jinjia Peng
  • Huibing Wang

Deep neural networks are susceptible to adversarial examples, which induce incorrect predictions through imperceptible perturbations. Transfer-based attacks create adversarial examples for surrogate models and transfer these examples to target models under black-box scenarios. Recent studies have established a strong correlation between the geometric properties of loss landscapes and the transferability of adversarial examples, demonstrating that flatter loss surfaces consistently yield superior transferability. However, we identify that these methods fail to account for the loss landscape flatness along the path from the current point to local minima, resulting in poor transferability. To address this, this paper constructs a novel Path Flatness Attack (PFA) method to significantly enhance the transferability of adversarial examples. Specifically, this paper proposes a novel path flatness indicator that not only evaluates the flatness in local minima regions but also explicitly quantifies the loss surface geometry along the trajectory from the current point to the minimum. Furthermore, we incorporate the path flatness indicator into the attack process, integrating penalties over low-loss points along the path while maximizing the loss function, thereby explicitly flattening the loss landscape. Extensive experiments demonstrate that PFA consistently achieves state-of-the-art attack performance across all experimental settings.

AAAI Conference 2025 Conference Paper

Anchor Learning with Potential Cluster Constraints for Multi-view Clustering

  • Yawei Chen
  • Huibing Wang
  • Jinjia Peng
  • Yang Wang

Anchor-based multi-view clustering has received extensive attention due to its efficient performance. Existing methods only focus on how to dynamically learn anchors from the original data and simultaneously construct anchor graphs describing the relationships between samples and perform clustering, while ignoring the reality of anchors, i.e., high-quality anchors should be generated uniformly from different clusters of data rather than scattered outside the clusters. To deal with this problem, we propose a noval method termed Anchor Learning with Potential Cluster Constraints for Multi-view Clustering (ALPC) method. Specifically, ALPC first establishes a shared latent semantic module to constrain anchors to be generated from specific clusters, and subsequently, ALPC improves the representativeness and discriminability of anchors by adapting the anchor graph to capture the common clustering center of mass from samples and anchors, respectively. Finally, ALPC combines anchor learning and graph construction into a unified framework for collaborative learning and mutual optimization to improve the clustering performance. Extensive experiments demonstrate the effectiveness of our proposed method compared to some state-of-the-art MVC methods.

AAAI Conference 2025 Conference Paper

CDE-Learning: Camera Deviation Elimination Learning for Unsupervised Person Re-identification

  • Jinjia Peng
  • Songyu Zhang
  • Huibing Wang

Unsupervised Person Re-identification (Re-ID) aims to identify the same person shot from non-overlapping cameras without any annotated data. In this task, attributes such as contrast, saturation, and resolution of the camera cause the deviation in target features. Since the camera label is readily available, they are employed to achieve the constraints across cameras and smooth the deviations during the model training phase. However, features from the same camera are prone to generating false positives due to the identical camera properties, which induce camera deviations on pseudo-label assignment. To address this problem, this paper proposes a novel camera-unbiased method named Camera Deviation Elimination Learning (CDE-Learning). In CDE-Learning, the Camera Deviation Compensation (CDC) module is designed to align data distributions from disparate cameras to decouple camera information from identity information during the pseudo-label allocation. Our Camera Deviation Balancing (CDB) module integrates different camera constraints in a united loss and adjusts camera constraints by constructing contrastive pairs between intra-camera and inter-camera. After explicit constraints, the Camera Attribution Auxiliary (CAA) task predicts whether a pair of images originates from the same camera to implicitly enhance the capacity to distinguish the camera deviation. We demonstrated the superior performance of the proposed CDE-Learning on benchmark datasets.

IJCAI Conference 2025 Conference Paper

Consensus-Guided Incomplete Multi-view Clustering via Cross-view Affinities Learning

  • Qian Liu
  • Huibing Wang
  • Jinjia Peng
  • Yawei Chen
  • Mingze Yao
  • Xianping Fu
  • Yang Wang

Incomplete multi-view clustering (IMC) has garnered substantial attention due to its capacity to handle unlabeled data. Existing methods predominantly explore pairwise consistency between every two views. However, such consistency is highly susceptible to missing samples and outliers within a certain view and thus deviates from the true clustering distribution. Moreover, dual-view interaction neglects the collaboration effects of multiple views, making it challenging to capture the holistic characteristics across views. In response to these issues, we propose a novel Consensus-Guided Incomplete Multi-view Clustering via Cross-view Affinities Learning (CAL). Specifically, CAL reconstructs views with available instances to mine sample-wise affinities and harness comprehensive content information within views. Subsequently, to extract clean structural information, CAL imposes a structured sparse constraint on the representation tensor to eliminate biased errors. Furthermore, by integrating the consensus representation into a representation tensor, CAL can employ high-order interaction of multiple views to depict the semantic correlation between views while acquiring a unified structural graph across multiple views. Extensive experiments on seven benchmark datasets demonstrate that CAL outperforms some state-of-the-art methods in clustering performance. The code is available at https: //github. com/whbdmu/CAL.

EAAI Journal 2025 Journal Article

Detail-focused and polarization-guided multi-modality fusion for underwater image clarity enhancing

  • Mingze Yao
  • Huibing Wang
  • Yudong Li
  • Wenzhe Liu
  • Xianping Fu

Underwater optics imaging typically suffer from the impurities scattering and light absorption in dynamic and complex underwater environment, which significantly effect the clarity and visibility of images. Existing cutting-edge Underwater Image Enhancement (UIE) methods mostly focus on color correction and contrast enhancement neglect the textual and detail information of objects, leading to imbalance exposure and edge features missing. To overcome these problems, we propose a novel detail-focused and polarization guided multi-modality fusion network (DFPG-Net), for enhancing underwater images. Unlike the previous methods, we first construct a Detail-Focused Convolution (DFC) block for extracting features from underwater multimodal images, which integrates difference convolutions to capture prior and edge information. Meanwhile, polarization information is introduced with a Multi-scale Polarization Guided (MPG) fusion module, which intends to maintain and enhance the texture and details information from the degree of polarization information and angle of polarization information obtained from the polarized image. Additionally, a parallel progressive attention network is designed to explore and combine the valuable and discriminative information in feature learning stage. Extensive experiments on the constructed underwater dataset validate the effectiveness and superior performance of the proposed DFPG-Net, which against state-of-the-art methods in both machine evaluation metrics and visual perception.

EAAI Journal 2025 Journal Article

Region-guided spatial feature aggregation network for vehicle re-identification

  • Yanzhen Xiong
  • Jinjia Peng
  • Zeze Tao
  • Huibing Wang

In the context of the advancement of smart city management, re-identification technology has emerged as an area of particular interest and research in the field of artificial intelligence, especially vehicle re-identification (re-ID), which aims to identify target vehicles in multiple non-overlapping fields of view. Most existing methods rely on fine-grained cues in the salient regions. Although impressive results have been achieved, these methods typically require additional auxiliary networks to localize the salient regions containing fine-grained cues. Meanwhile, changes in state such as illumination, viewpoint and occlusion can affect the position of the salient regions. To solve the above problems, this paper proposes a Region-guided Spatial Feature Aggregation Network (RSFAN) for vehicle re-ID, which forces the model to learn the latent information in the minor salient regions. Firstly, a Regional Localization (RL) module is proposed to automatically locate the salient regions without additional auxiliary networks. In addition, to mitigate the misguidance caused by the inaccurate salient regions, a Spatial Feature Aggregation (SFA) module is designed to weaken and enhance the expression of the salient and minor salient regions, respectively. Meanwhile, to enhance the diversity of the minor salient region-related information, a Cross-level Channel Attention (CCA) module is designed to implement cross-level interactions through the channel attention mechanism across different levels. Finally, to constrain the distributional differences between the salient regions and minor salient regions feature, a Distributional Variance (DV) loss is proposed. The extensive experiments show that the RSFAN has a good performance on VeRi-776, VehicleID, VeRi-Wild and Market1501 datasets.

NeurIPS Conference 2025 Conference Paper

Spatiotemporal Consensus with Scene Prior for Unsupervised Domain Adaptive Person Search

  • Yimin Jiang
  • Huibing Wang
  • Jinjia Peng

Person Search aims to locate query persons in gallery scene images, but faces severe performance degradation under domain shifts. Unsupervised domain adaptation transfers knowledge from the labeled source domain to the unlabeled target domain and iteratively rectifies the pseudo-labels. However, the pseudo-labels are inevitably contaminated by the source-biased model, which misleads the training process. This, in turn, reduces the quality of the pseudo-labels themselves and ultimately affects the search performance. In this paper, we propose a Spatiotemporal Consensus with Scene Prior (STCSP) framework that effectively eliminates the interference of noise on pseudo-labels, establishes positive feedback, and thus gradually bridging the domain gap. Firstly, STCSP uses a Spatiotemporal Consensus pipeline to suppress the noise from being mixed into the pseudo-labels. Secondly, leveraging the scene prior, STCSP employs our designed Iterative Bilateral Extremum Matching method to prevent the occurrence of some incorrect pseudo-labels. Thirdly, we propose a Scene Prior Contrastive Learning module, which encourages the model to directly acquire the scene prior knowledge from the target domain, thereby mitigating the generation of noise. By suppressing noise contamination, avoiding noise occurrence and mitigating noise generation, our framework achieves state-of-the-art performance on two benchmark datasets, PRW with 50. 2% mAP and CUHK-SYSU with 87. 0% mAP.

IJCAI Conference 2025 Conference Paper

Stabilizing Holistic Semantics in Diffusion Bridge for Image Inpainting

  • Jinjia Peng
  • Mengkai Li
  • Huibing Wang

Image inpainting aims to restore the original image from a damaged version. Recently, a special type of diffusion bridge model has achieved promising performance by directly mapping the degradation process and restoring corrupted images through the corresponding reverse process. However, due to the lack of explicit semantic priors during the denoising process, the inpainted results typically exhibit inferior context-stability and semantic consistency. To this end, this paper proposes a novel Global Structure-Guided Diffusion Bridge framework (GSGDiff), which incorporates an additional structure restorer to stabilize the generation of holistic semantics. Specifically, to acquire richer semantic structure priors, this paper proposes a posterior sampling approach that captures semantically global and consistent structures at each timestep, efficiently integrating them into the texture generation through the corresponding guidance module. Additionally, considering the characteristics of diffusion models with low denoising levels at larger timesteps, this paper proposes a semantic fusion schedule to avoid noise interference by reducing the weight of ineffective guided semantics in the early stages. By applying the proposed posterior sampling to the texture denoising process, GSGDiff can achieve more stable and superior inpainting results over competitive baselines. Experiments on Places2, Paris Street View and CelebA-HQ datasets validate the efficacy of the proposed method.

AAAI Conference 2025 Conference Paper

Unsupervised Domain Adaptive Person Search via Dual Self-Calibration

  • Linfeng Qi
  • Huibing Wang
  • Jiqing Zhang
  • Jinjia Peng
  • Yang Wang

Unsupervised Domain Adaptive (UDA) person search focuses on employing the model trained on a labeled source domain dataset to a target domain dataset without any additional annotations. Most effective UDA person search methods typically utilize the ground truth of the source domain and pseudo-labels derived from clustering during the training process for domain adaptation. However, the performance of these approaches will be significantly restricted by the disrupting pseudo-labels resulting from inter-domain disparities. In this paper, we propose a Dual Self-Calibration (DSCA) framework for UDA person search that effectively eliminates the interference of noisy pseudo-labels by considering both the image-level and instance-level features perspectives. Specifically, we first present a simple yet effective Perception-Driven Adaptive Filter (PDAF) to adaptively predict a dynamic filter threshold based on input features. This threshold assists in eliminating noisy pseudo-boxes and other background interference, allowing our approach to focus on foreground targets and avoid indiscriminate domain adaptation. Besides, we further propose a Cluster Proxy Representation (CPR) module to enhance the update strategy of cluster representation, which mitigates the pollution of clusters from misidentified instances and effectively streamlines the training process for unlabeled target domains. With the above design, our method can achieve state-of-the-art (SOTA) performance on two benchmark datasets, with 80.2% mAP and 81.7% top-1 on the CUHK-SYSU dataset, with 39.9% mAP and 81.6% top-1 on the PRW dataset, which is comparable to or even exceeds the performance of some fully supervised methods.

EAAI Journal 2024 Journal Article

A depth map stitching framework based on salient region matching

  • Zetian Mi
  • Haixia Qi
  • Jiaxin Chen
  • Yang Yu
  • Yujia Wang
  • Huibing Wang
  • Xianping Fu

Depth map stitching technology is urgently needed in varying fields such as panoramic three-dimensional reconstruction, underwater mapping, autonomous driving, and robot collision avoidance control. Existing image stitching research mainly focuses on three-primary-color images and heavily relies on feature detection quality, while depth map stitching is rarely studied due to the sparse features and low resolution. To address the above challenges, this paper proposes a color image-guided depth map stitching framework consisting of two stages: coarse depth map stitching and depth correction. In the first stage, a color-image-guided coarse depth map stitching network is proposed to accurately estimate the homography matrix, which can further eliminate parallax artifacts in the stitched depth map. In the second stage, a transformer-based depth correction network is designed, which combines saliency detection and regional matching strategies for depth correction, in order to solve the inconsistency in depth value introduced by point-to-point matching and improve the speed of depth correction to some extent. Extensive experiments demonstrate that the proposed method effectively solves the problem of disparity artifacts and mismatched depth values of the same target, and has excellent performance in depth map stitching.

IJCAI Conference 2024 Conference Paper

Fast One-Stage Unsupervised Domain Adaptive Person Search

  • Tianxiang Cui
  • Huibing Wang
  • Jinjia Peng
  • Ruoxi Deng
  • Xianping Fu
  • Yang Wang

Unsupervised person search aims to localize a particular target person from a gallery set of scene images without annotations, which is extremely challenging due to the unexpected variations of the unlabeled domains. However, most existing methods dedicate to developing multi-stage models to adapt domain variations while using clustering for iterative model training, which inevitably increase model complexity. To address this issue, we propose a Fast One-stage Unsupervised person Search (FOUS) which complementaryly integrates domain adaption with label adaption within an end-to-end manner without iterative clustering. To minimize the domain discrepancy, FOUS introduced an Attention-based Domain Alignment Module (ADAM) which can not only align various domains for both detection and ReID tasks but also construct an attention mechanism to reduce the adverse impacts of low-quality candidates resulting from unsupervised detection. Moreover, to avoid the redundant iterative clustering mode, FOUS adopts a prototype-guided labeling method which minimizes redundant correlation computations for partial samples and assigns noisy coarse label groups efficiently. The coarse label groups will be continuously refined via label-flexible training network with an adaptive selection strategy. With the adapted domains and labels, FOUS can achieve the state-of-the-art (SOTA) performance on two benchmark datasets, CUHK-SYSU and PRW. The code is available at https: //github. com/whbdmu/FOUS.

IJCAI Conference 2024 Conference Paper

Scene-Adaptive Person Search via Bilateral Modulations

  • Yimin Jiang
  • Huibing Wang
  • Jinjia Peng
  • Xianping Fu
  • Yang Wang

Person search aims to localize specific a target person from a gallery set of images with various scenes. As the scene of moving pedestrian changes, the captured person image inevitably bring in lots of background noise and foreground noise on the person feature, which are completely unrelated to the person identity, leading to severe performance degeneration. To address this issue, we present a Scene-Adaptive Person Search (SEAS) model by introducing bilateral modulations to simultaneously eliminate scene noise and maintain a consistent person representation to adapt to various scenes. In SEAS, a Background Modulation Network (BMN) is designed to encode the feature extracted from the detected bounding box into a multi-granularity embedding, which reduces the input of background noise from multiple levels with norm-aware. Additionally, to mitigate the effect of foreground noise on the person feature, SEAS introduces a Foreground Modulation Network (FMN) to compute the clutter reduction offset for the person embedding based on the feature map of the scene image. By bilateral modulations on both background and foreground within an end-to-end manner, SEAS obtains consistent feature representations without scene noise. SEAS can achieve state-of-the-art (SOTA) performance on two benchmark datasets, CUHK-SYSU with 97. 1% mAP and PRW with 60. 5% mAP. The code is available at https: //github. com/whbdmu/SEAS.

EAAI Journal 2023 Journal Article

Hybrid partial-constrained learning with orthogonality regularization for unsupervised person re-identification

  • Jiazuo Yu
  • Jinjia Peng
  • Kai Li
  • Huibing Wang

Person re-identification (re-ID) aims at determining whether there is a specific person in image sets or videos via computer vision technology. State-of-the-art unsupervised re-ID methods extract image features through CNNs-based networks and store these extracted features in memory for identity matching. However, extracted global features of these methods ignore the problem of information redundancy and the influence of the constraints between the internal features. To overcome these problems, a Hybrid Partial-constrained Learning (HPcL) network with orthogonality regularization is proposed to learn a discriminative visual representation by generating hybrid features. Specifically, the hybrid features are generated by our designed Dynamic Fusion Module (DFM) to initialize the memory dictionary and match the identity, which can constrain each part of the features extracted by our proposed Multi-Scale (M-S) module and learn robust visual representations. In addition, a new orthogonal regularization method is introduced to constrain orthogonality of the kernel weights and features, which reduces the correlations among features. Extensive experimental results on Market-1501, DukeMTMC-reID, PersonX, and MSMT17 datasets demonstrate that our method is effective and superior to the state-of-the-art methods.

EAAI Journal 2022 Journal Article

Adaptive multi-view multiple-means clustering via subspace reconstruction

  • Wenzhe Liu
  • Luyao Liu
  • Yong Zhang
  • Huibing Wang
  • Lin Feng

Clustering is a notable research topic, but it is still challenging when facing massive multi-view data from different ways or multiple feature extractors. The crucial problem is how to promote cooperative learning between views via subspace reconstruction while preserving the underlying geometric structure of data. Moreover, most existing methods habitually utilize K-means to achieve the final results, which is not conducive to dealing with intricate non-convex patterns in multi-perspective data. Based on the above consideration, in this paper, we present a fresh multi-view clustering approach called Adaptive Multi-view Multiple-Means Clustering via Subspace Reconstruction(AM 2 CSR). AM 2 CSR aims to simultaneously capture compatible, complementary, geometric, and discrimination information among multiple views. Subsequently, a low-rank restriction is forced on the low-dimensional representation to reduce redundancy, and K-Multiple-Means(KMM) is adopted as the clustering technique to achieve satisfying results. Additionally, an effective iteration updating method with a convergence guarantee is applied to settle the optimization matter of AM 2 CSR. Extensive empirical experiments on eight benchmark datasets exhibit the superiority of AM 2 CSR.

EAAI Journal 2022 Journal Article

Learning latent features with local channel drop network for vehicle re-identification

  • Xianping Fu
  • Jinjia Peng
  • Guangqi Jiang
  • Huibing Wang

Vehicle re-identification targets to find the target vehicle images in a large dataset which is composed of vehicle images from multiple non-overlapping cameras. Due to the various illumination, viewpoints and resolutions, it is challenging to find the right vehicle images accurately. Most existing works put emphasis on learning strong features by exploiting the attention parts in vehicle images, which leads to some small important cues being suppressed by these significant parts. Hence, a local channel drop network (LCDNet) is proposed in this paper, which focuses on seeking the latent features by releasing the constraint of most attentive features. Specially, besides the normal local feature learning network, LCDNet consists of an attentive local feature learning branch that drops some regions to promote learning the attentive features of local regions. Besides, the batch ranking loss is introduced to split the samples into two groups in a batch and regularize them by enforcing a margin, which ensures the model to learn meaningful features to distinct vehicles. Moreover, to further calculate the similarity of various images, the paper proposes a multi-distance based ranking method to achieve more accurate results. Experiments on several benchmark datasets validate the effectiveness of the proposed method.

EAAI Journal 2021 Journal Article

Multi-view Low-rank Preserving Embedding: A novel method for multi-view representation

  • Xiangzhu Meng
  • Lin Feng
  • Huibing Wang

In recent years, we have witnessed a surge of interest in multi-view representation learning. When facing multiple views that are highly related but sightly different from each other, most existing multi-view methods might fail to fully explore multi-view information. Additionally, pairwise correlations among multiple views often vary drastically, which makes multi-view representation challenging. Therefore, how to learn appropriate representation from multi-view information is still an open but challenging problem. To handle this issue, this paper proposes a novel multi-view learning method, named Multi-view Low-rank Preserving Embedding (MvLPE). It integrates all views into a common latent space, termed as centroid view, by minimizing the disagreement between centroid view and each view, which encourages different views to mutually learn from each other. Unlike existing methods with explicit weight definition, the proposed method could automatically allocate an ideal weight for each view according to its contribution. Besides, MvLPE could maintain its low-rank reconstruction structure for each view while integrating all views into centroid view. Since there is no closed-form solution for MvLPE, an effective algorithm based on iterative alternating strategy is provided to obtain the solution. The experiments on six benchmark datasets validate the effectiveness of the proposed method, which achieves superior performance over its counterparts.

EAAI Journal 2020 Journal Article

Purifying real images with an attention-guided style transfer network for gaze estimation

  • Xianping Fu
  • YuXiao Yan
  • Yang Yan
  • Jinjia Peng
  • Huibing Wang

Recently, the progress of learning-by-synthesis has proposed a training model for synthetic images, which can effectively reduce the cost of human and material resources. Image synthesis has been widely accepted as a cost effective way to learn models because it provides training sets that are large, diverse and accurately labeled. However, the realism of the synthetic image is not enough, this affects generalization on naturalistic test image. In an attempt to address this issue, previous methods learn a model to improve the realism of synthetic image. Different from previous methods, we take the first step towards purifying the real image to weaken the influence of light and convert the distribution of an outdoor naturalistic image through a real-time style transfer task to that of indoor synthetic image. In this paper, we first introduce the segmentation masks to construct Red, Green, and Blue-mask (RGB-mask) pairs as inputs, then we design an attention-guided style transfer network to learn style features separately from the attention and background regions, learn content features from full and attention regions. Moreover, we propose a novel region-level task-guided loss to restrain the features learnt from style and content. Experiments were performed using a mixed research (qualitative and quantitative) method to demonstrate the possibility of purifying real images in complex directions. We evaluate the proposed method on three public datasets, including Labeled pupils in the wild (LPW), Common Objects in COntext (COCO) and MPIIGaze. Extensive experimental results show that the proposed method is effective and achieves the state-of-the-art results.

IJCAI Conference 2020 Conference Paper

Unsupervised Vehicle Re-identification with Progressive Adaptation

  • Jinjia Peng
  • Yang Wang
  • Huibing Wang
  • Zhao Zhang
  • Xianping Fu
  • Meng Wang

Vehicle re-identification (reID) aims at identifying vehicles across different non-overlapping cameras views. The existing methods heavily relied on well-labeled datasets for ideal performance, which inevitably causes fateful drop due to the severe domain bias between the training domain and the real-world scenes; worse still, these approaches required full annotations, which is labor-consuming. To tackle these challenges, we propose a novel Progressive Adaptation Learning method for vehicle reID, named PAL, which infers from the abundant data without annotations. For PAL, a data adaptation module is employed for source domain, which generates the images with similar data distribution to unlabeled target domain as “pseudo target samples”. These pseudo samples are combined with the unlabeled samples that are selected by a dynamic sampling strategy to make training faster. We further proposed a weighted label smoothing (WLS) loss, which considers the similarity between samples with different clusters to balance the confidence of pseudo labels. Comprehensive experimental results validate the advantages of PAL on both VehicleID and VeRi-776 dataset.

EAAI Journal 2019 Journal Article

Purifying naturalistic images through a real-time style transfer semantics network

  • Tongtong Zhao
  • YuXiao Yan
  • Ibrahim Shehi Shehu
  • Xianping Fu
  • Huibing Wang

Recently, the progress of learning-by-synthesis has proposed a training model for synthetic images, which can effectively reduce the cost of human and material resources. However, due to the different distribution of synthetic images compared to real images, the desired performance cannot still be achieved. Real images consist of multiple forms of light orientation, while synthetic images consist of a uniform light orientation. These features are considered to be characteristic of outdoor and indoor scenes, respectively. To solve this problem, the previous method learned a model to improve the realism of the synthetic image. Different from the previous methods, this paper takes the first step to purify real images. Through the style transfer task, the distribution of outdoor real images is converted into indoor synthetic images, thereby reducing the influence of light. Therefore, this paper proposes a real-time style transfer network that preserves image content information (e. g. , gaze direction, pupil center position) of an input image (real image) while inferring style information (e. g. , image color structure, semantic features) of style image (synthetic image). In addition, the network accelerates the convergence speed of the model and adapts to multi-scale images. Experiments were performed using mixed studies (qualitative and quantitative) methods to demonstrate the possibility of purifying real images in complex directions. Qualitatively, it compares the proposed method with the available methods in a series of indoor and outdoor scenarios of the LPW dataset. In quantitative terms, it evaluates the purified image by training a gaze estimation model on the cross data set. The results show a significant improvement over the baseline method compared to the raw real image.

v2026.09.13