Arrow Research search

Author name cluster

Chenggang Yan

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

23 papers
1 author row

Possible papers

23

AAAI Conference 2026 Conference Paper

2D-CrossScan Mamba: Enhancing State Space Models with Spatially Consistent Multi-Path 2D Information Propagation

  • Longlong Yu
  • Wenxi Li
  • Yaoqi Sun
  • Hang Xu
  • Chenggang Yan
  • Yuchen Guo

Despite recent progress in adapting State Space Models such as Mamba to vision tasks, their intrinsic 1D scanning mechanism imposes limitations when applied to inherently 2D-structured data like images. Existing adaptations, including VMamba and 2DMamba, either suffer from inconsistency between scanning order and spatial locality or restrict inter-patch communication to singular paths, hindering effective information propagation. In this paper, we propose 2D-CrossScan, a novel 2D-compatible scan framework that enables spatially consistent, multi-path hidden state propagation by integrating modified state equations over two-dimensional neighborhoods. Furthermore, we mitigate redundant information accumulation due to overlapping paths via cross-directional subtraction. To fully align with the 2D spatial structure, we introduce a multi-directional scanning strategy that starts simultaneously from all four corners of the image, enabling diverse propagation paths and better feature integration. Our approach maintains efficiency, requiring only minimal architectural changes to existing Mamba variants. Experimental results demonstrate substantial improvements in multiple visual tasks, including object detection and semantic segmentation on PANDA and COCO datasets. Compared to baseline SSM-based methods, 2D-CrossScan consistently yields better spatial representations, as confirmed by extensive effective receptive field visualizations and attention analyses. These results highlight the importance of geometry-aware state propagation and validate 2D-CrossScan as a simple yet powerful extension to SSMs for vision.

EAAI Journal 2026 Journal Article

Adaptive eco-cooperative adaptive cruise control for heterogeneous Vehicle platoons using online identification-informed deep reinforcement learning

  • Jianing Sun
  • Chenhao Xiong
  • Chunjie Zhai
  • Yuyuan Li
  • Xiongding Liu
  • Chuqiao Chen
  • Chenggang Yan
  • Yahong Chen

With the advancement of autonomous driving technologies, deep reinforcement learning (DRL) has increasingly played a pivotal role in the control of intelligent connected vehicle (ICV) platoons. Traditional DRL-based platoon control methods often overlook the dynamic discrepancies between vehicles (such as weight and powertrain delays), which results in inadequate adaptability of the control systems. To address this issue, this paper introduces an innovative heterogeneous platoon control method that combines the Adaptive Forgetting Factor Least Squares (AFFLS) algorithm with Proximal Policy Optimization (PPO). The proposed approach leverages the AFFLS algorithm to perform real-time identification of vehicle dynamic parameters, enabling the control system to dynamically adjust to various vehicles. This enhances the platoon stability, passenger comfort, communication delay robustness, and fuel efficiency. Simulation results demonstrate the practical advantages of this approach: maintaining the average inter-vehicle distance error within 0. 01 m, reducing the stabilization time to approximately 3 s, and lowering fuel consumption. Compared to four baseline methods, our method exhibits superior adaptability and robustness. This study offers a novel perspective for the collaborative control of heterogeneous platoons in real-world applications, achieving more efficient and safer vehicle control in dynamic traffic environments.

EAAI Journal 2026 Journal Article

Feature boosting and scale-aware network with multi-modal information for underwater salient object detection

  • Tingyu Wang
  • Junzhe Lu
  • Bin Wan
  • Rongfeng Lu
  • Yaoqi Sun
  • Duanpo Wu
  • Yanbin Liu
  • Chenggang Yan

Underwater salient object detection (USOD) plays an important role in marine engineering applications. Due to the poor image quality caused by the complex underwater environment, USOD remains a challenging task. Existing methods typically fuse features extracted from underwater images and depth maps to exploit more salient cues. However, they neglect the degradation and noise corruption inherent in underwater multi-modal data. Moreover, these methods pay limited attention to the scale variations of underwater objects. To address these issues, we propose a feature boosting and scale-aware network (FBS-Net), which includes two processes: underwater feature boosting and scale-aware iterative fusion decoding. The underwater feature boosting process employs a scene contrast perception module (SCPM) to improve underwater Red–Green–Blue (RGB) image feature reliability through comparative analysis of the differences between RGB images and underwater enhanced images, and a frequency domain decoupling fusion module (FDFM) to reconstruct high-quality scene representations from noisy multi-modal inputs. In the scale-aware iterative fusion decoding process, a scale-aware iterative fusion decoder (SIFD) is introduced to dynamically process multi-scale information while suppressing noise and conflicting features through multiple refinement iterations. Extensive experiments on the USOD10K, COD10K, and USOD datasets demonstrate that our proposed method outperforms 16 state-of-the-art methods. Our code and results are available at https: //github. com/llllxxx2333/FBSNet.

AAAI Conference 2026 Conference Paper

Forgetting Knowledge Localization and Isolation for Continual Forgetting of Pre-trained Vision Models

  • Zhiwen Yang
  • Jiehua Zhang
  • Chenggang Yan
  • Yuhan Gao
  • Zongpeng Li
  • Xichun Sheng
  • Liang Li

Continual forgetting task aims to continuously remove multiple target knowledge subsets from pre-trained models while maintaining the integrity of remaining knowledge. Existing methods suffer from both incomplete forgetting of target knowledge and unintended forgetting of indistinguishable remaining knowledge. To address these challenges, we propose the forgetting knowledge localization and isolation for continual forgetting in pre-trained vision models which precisely forgets target knowledge while reducing over-forgetting of remaining knowledge. To achieve precise forgetting, we first propose the forgetting knowledge layer localization to explore layers in the model which are more related to forgetting knowledge. Then, we design the forgetting knowledge parameter isolation to isolate the parameters sensitive to forgetting knowledge in these selected layers, mitigating over-forgetting of remaining knowledge. Finally, we fine-tune these isolated parameters and freeze the remaining parameters to achieve efficient forgetting while maintaining high performance on retained datasets. Extensive experimental results demonstrate that our method achieves superior performance over state-of-the-art methods across multiple continual forgetting tasks.

AAAI Conference 2026 Conference Paper

Temporal Calibrating and Distilling for Scene-Text Aware Text-Video Retrieval

  • Zhiqian Zhao
  • Liang Li
  • Lei Shen
  • Xichun Sheng
  • Yaoqi Sun
  • Fang Kang
  • Chenggang Yan

Existing text-video retrieval methods mainly focus on singlemodal video content (i.e., visual entities), often overlooking heterogeneous scene text that is ubiquitous in human environments. Although scene text in videos provides finegrained semantics for cross-modal retrieval, effectively utilizing it presents two key challenges: (1) Temporally dense scene text disrupts sync with sparse video frames, obstructing video understanding;(2) Redundant scene text and irrelevant video frames hinder the learning of discriminative temporal clues for retrieval. To address them, we propose a temporal scene-text calibrating and distilling (TCD) network for textvideo retrieval. Specifically, we first design a window-OCR captioner that aggregates dense scene text into OCR captions to facilitate feature interaction. Next, we devise a heterogeneous semantics calibration module that leverages scene text as a self-supervised signal to temporally align window-level OCR captions and frame-level video features. Further, we introduce a context-guided temporal clue distillation module to learn the complementary and relevant details between scene text and video modalities, thereby obtaining discriminative temporal clues for retrieval. Extensive experiments show that our TCD achieves state-of-the-art performance on three scene-text related benchmarks.

IJCAI Conference 2025 Conference Paper

K-Buffers: A Plug-in Method for Enhancing Neural Fields with Multiple Buffers

  • Haofan Ren
  • Zunjie Zhu
  • Xiang Chen
  • Ming Lu
  • Rongfeng Lu
  • Chenggang Yan

Neural fields are now the central focus of research in 3D vision and computer graphics. Existing methods mainly focus on various scene representations, such as neural points and 3D Gaussians. However, few works have studied the rendering process to enhance the neural fields. In this work, we propose a plug-in method named K-Buffers that leverages multiple buffers to improve the rendering performance. Our method first renders K buffers from scene representations and constructs K pixel-wise feature maps. Then, We introduce a K-Feature Fusion Network (KFN) to merge the K pixel-wise feature maps. Finally, we adopt a feature decoder to generate the rendering image. We also introduce an acceleration strategy to improve rendering speed and quality. We apply our method to well-known radiance field baselines, including neural point fields and 3D Gaussian Splatting (3DGS). Extensive experiments demonstrate that our method effectively enhances the rendering performance of neural point fields and 3DGS.

AAAI Conference 2025 Conference Paper

Region-aware Difference Distilling with Attribute-guided Contrastive Regularization for Change Captioning

  • Rong Li
  • Liang Li
  • Jiehua Zhang
  • Qiang Zhao
  • Hongkui Wang
  • Chenggang Yan

Change captioning aims to describe the differences between two similar images using natural language, significantly aiding in understanding and monitoring changes. This challenging task requires a fine-grained understanding of subtle changes while resisting disturbances like viewpoint shifts and illumination variations. Existing methods often rely solely on global difference features and lack comprehensive alignment of linguistic and visual information, leading to overlooking fine-grained details and generating semantic hallucinated sentences. To address these limitations, we propose the region-aware difference distilling (RDD) network with attribute-guided contrastive regularization (ACR). The RDD uses global difference features to progressively distill regional difference features using learnable vectors, allowing for more precise identification of changed regions. The ACR enhances comprehensive alignment between linguistic and visual information by formulating Nouns-to-Objects (N2O) and Verbs-to-Actions (V2A) alignment losses to regularize the regional difference features. Promising results on three datasets demonstrate that our method outperforms the state-of-the-art change captioning methods.

AAAI Conference 2025 Conference Paper

SdalsNet: Self-Distilled Attention Localization and Shift Network for Unsupervised Camouflaged Object Detection

  • Peiyao Shou
  • Yixiu Liu
  • Wei Wang
  • Yaoqi Sun
  • Zhigao Zheng
  • Shangdong Zhu
  • Chenggang Yan

Unsupervised camouflaged object detection (UCOD) poses significant challenges, primarily attributed to the absence of human labels. Existing UCOD methodologies, leveraging attention mechanisms, often struggle to achieve precise localization of camouflaged objects. To overcome this limitation, we introduce a groundbreaking fully unsupervised algorithm for attention-guided camouflaged object localization, shift, and inference, termed the self-distilled attention localization and shift network (SdalsNet). In this study, we formulate an attention localization methodology aimed at accurately identifying the central coordinate of the camouflaged object. Furthermore, we propose four distinct loss functions tailored to refine the precision of attentional positioning. These loss functions effectively constrain the distances between three types of class tokens, facilitating seamless attentional shifting across the input sample. Additionally, we design a sophisticated prediction inference technique to reconstruct the binary output of an attention map, thereby providing a comprehensive understanding of the detected camouflaged objects. Experimental results on four challenging COD benchmark datasets corroborate the effectiveness of our proposed approach, demonstrating notable superiority over state-of-the-art methods.

EAAI Journal 2024 Journal Article

ADNet: Anti-noise dual-branch network for road defect detection

  • Bin Wan
  • Xiaofei Zhou
  • Yaoqi Sun
  • Tingyu Wang
  • Chengtao Lv
  • Shuai Wang
  • Haibing Yin
  • Chenggang Yan

This paper addresses the issue of noise interference in road defect detection, caused by various environmental factors or acquisition equipment. In this article, we add three different levels of salt & pepper noise to the road defect dataset and propose a novel anti-noise dual-branch network (ADNet). The proposed ADNet leverages two backbone networks equipped with the dual-branch interaction (DI) modules to learn the defect information from noise and clear images for improving noise immunity. Then, the weighted feature representation (WFR) module is designed to extract more context-aware cues from the multi-level feature. Additionally, the region perception unit is proposed, where channel-spatial attention optimization (CSAO) module extracts more defect region information by utilizing the attention mechanism and multi-scale refinement (MR) optimizes the boundary information with the U-Net structure. Extensive experimental results demonstrate that the proposed method outperforms state-of-the-art methods, making it a promising solution for detecting road defects in noisy environments.

AAAI Conference 2024 Conference Paper

Coupled Confusion Correction: Learning from Crowds with Sparse Annotations

  • Hansong Zhang
  • Shikun Li
  • Dan Zeng
  • Chenggang Yan
  • Shiming Ge

As the size of the datasets getting larger, accurately annotating such datasets is becoming more impractical due to the expensiveness on both time and economy. Therefore, crowd-sourcing has been widely adopted to alleviate the cost of collecting labels, which also inevitably introduces label noise and eventually degrades the performance of the model. To learn from crowd-sourcing annotations, modeling the expertise of each annotator is a common but challenging paradigm, because the annotations collected by crowd-sourcing are usually highly-sparse. To alleviate this problem, we propose Coupled Confusion Correction (CCC), where two models are simultaneously trained to correct the confusion matrices learned by each other. Via bi-level optimization, the confusion matrices learned by one model can be corrected by the distilled data from the other. Moreover, we cluster the ``annotator groups'' who share similar expertise so that their confusion matrices could be corrected together. In this way, the expertise of the annotators, especially of those who provide seldom labels, could be better captured. Remarkably, we point out that the annotation sparsity not only means the average number of labels is low, but also there are always some annotators who provide very few labels, which is neglected by previous works when constructing synthetic crowd-sourcing annotations. Based on that, we propose to use Beta distribution to control the generation of the crowd-sourcing labels so that the synthetic annotations could be more consistent with the real-world ones. Extensive experiments are conducted on two types of synthetic datasets and three real-world datasets, the results of which demonstrate that CCC significantly outperforms state-of-the-art approaches. Source codes are available at: https://github.com/Hansong-Zhang/CCC.

NeurIPS Conference 2024 Conference Paper

CURE4Rec: A Benchmark for Recommendation Unlearning with Deeper Influence

  • Chaochao Chen
  • Jiaming Zhang
  • Yizhao Zhang
  • Li Zhang
  • Lingjuan Lyu
  • Yuyuan Li
  • Biao Gong
  • Chenggang Yan

With increasing privacy concerns in artificial intelligence, regulations have mandated the right to be forgotten, granting individuals the right to withdraw their data from models. Machine unlearning has emerged as a potential solution to enable selective forgetting in models, particularly in recommender systems where historical data contains sensitive user information. Despite recent advances in recommendation unlearning, evaluating unlearning methods comprehensively remains challenging due to the absence of a unified evaluation framework and overlooked aspects of deeper influence, e. g. , fairness. To address these gaps, we propose CURE4Rec, the first comprehensive benchmark for recommendation unlearning evaluation. CURE4Rec covers four aspects, i. e. , unlearning Completeness, recommendation Utility, unleaRning efficiency, and recommendation fairnEss, under three data selection strategies, i. e. , core data, edge data, and random data. Specifically, we consider the deeper influence of unlearning on recommendation fairness and robustness towards data with varying impact levels. We construct multiple datasets with CURE4Rec evaluation and conduct extensive experiments on existing recommendation unlearning methods. Our code is released at https: //github. com/xiye7lai/CURE4Rec.

AAAI Conference 2024 Conference Paper

Quad Bayer Joint Demosaicing and Denoising Based on Dual Encoder Network with Joint Residual Learning

  • Bolun Zheng
  • Haoran Li
  • Quan Chen
  • Tingyu Wang
  • Xiaofei Zhou
  • Zhenghui Hu
  • Chenggang Yan

The recent imaging technology Quad Bayer CFA brings better imaging PSNR and higher visual quality compared to traditional Bayer CFA, but also serious challenges for demosaicing and denoising during the ISP pipeline. In this paper, we propose a novel dual encoder network, namely DRNet, to achieve joint demosaicing and denoising for Quad Bayer CFA. The dual encoders are carefully designed in that one is mainly constructed by a joint residual block to jointly estimate the residuals for demosaicing and denoising separately. In contrast, the other one is started with a pixel modulation block which is specially designed to match the characteristics of Quad Bayer pattern for better feature extraction. We demonstrate the effectiveness of each proposed component through detailed ablation investigations. The comparison results on public benchmarks illustrate that our DRNet achieves an apparent performance gain~(0.38dB to the 2nd best) from the state-of-the-art method and balances performance and efficiency well. The experiments on real-world images show that the proposed method could enhance the reconstruction quality from the native ISP algorithm.

AAAI Conference 2023 Conference Paper

Improving Dynamic HDR Imaging with Fusion Transformer

  • Rufeng Chen
  • Bolun Zheng
  • Hua Zhang
  • Quan Chen
  • Chenggang Yan
  • Gregory Slabaugh
  • Shanxin Yuan

Reconstructing a High Dynamic Range (HDR) image from several Low Dynamic Range (LDR) images with different exposures is a challenging task, especially in the presence of camera and object motion. Though existing models using convolutional neural networks (CNNs) have made great progress, challenges still exist, e.g., ghosting artifacts. Transformers, originating from the field of natural language processing, have shown success in computer vision tasks, due to their ability to address a large receptive field even within a single layer. In this paper, we propose a transformer model for HDR imaging. Our pipeline includes three steps: alignment, fusion, and reconstruction. The key component is the HDR transformer module. Through experiments and ablation studies, we demonstrate that our model outperforms the state-of-the-art by large margins on several popular public datasets.

EAAI Journal 2023 Journal Article

SMINet:Semantics-aware multi-level feature interaction network for surface defect detection

  • Bin Wan
  • Xiaofei Zhou
  • Yaoqi Sun
  • Zunjie Zhu
  • Haibing Yin
  • Ji Hu
  • Jiyong Zhang
  • Chenggang Yan

To boost the product quality, numerous saliency-based surface defect detection methods have been devoted to the areas of industrial production, construction consumable, road construction. However, the existing salient object detection (SOD) methods not only consume a significant amount of computing resources but also fail to meet the detection efficiency requirements of enterprises. Therefore, this paper proposes a lightweight semantics-aware multi-level feature interaction network (SMINet), to address the above issues. In the encoder phase, we integrate multiple adjacent level features in the cross-layer feature fusion (CFF) module to alleviate the discrepancy between multi-scale features. In the decoder phase, we first employ the semantic-aware feature extraction (SFE) module to mine the location cues embedded in the high-level features. Afterwards, we introduce the detail-aware context attention (DCA) module based on the attention mechanism to recover more spatial details. Extensive experiments on four surface defect datasets validate that our SMINet outperforms the existing state-of-the-art methods.

IJCAI Conference 2023 Conference Paper

Sph2Pob: Boosting Object Detection on Spherical Images with Planar Oriented Boxes Methods

  • Xinyuan Liu
  • Hang Xu
  • Bin Chen
  • Qiang Zhao
  • Yike Ma
  • Chenggang Yan
  • Feng Dai

Object detection on panoramic/spherical images has been developed rapidly in the past few years, where IoU-calculator is a fundamental part of various detector components, i. e. Label Assignment, Loss and NMS. Due to the low efficiency and non-differentiability of spherical Unbiased IoU, spherical approximate IoU methods have been proposed recently. We find that the key of these approximate methods is to map spherical boxes to planar boxes. However, there exists two problems in these methods: (1) they do not eliminate the influence of panoramic image distortion; (2) they break the original pose between bounding boxes. They lead to the low accuracy of these methods. Taking the two problems into account, we propose a new sphere-plane boxes transform, called Sph2Pob. Based on the Sph2Pob, we propose (1) an differentiable IoU, Sph2Pob-IoU, for spherical boxes with low time-cost and high accuracy and (2) an agent Loss, Sph2Pob-Loss, for spherical detection with high flexibility and expansibility. Extensive experiments verify the effectiveness and generality of our approaches, and Sph2Pob-IoU and Sph2Pob-Loss together boost the performance of spherical detectors. The source code is available at https: //github. com/AntXinyuan/sph2pob.

AAAI Conference 2022 Conference Paper

Unbiased IoU for Spherical Image Object Detection

  • Feng Dai
  • Bin Chen
  • Hang Xu
  • Yike Ma
  • Xiaodong Li
  • Bailan Feng
  • Peng Yuan
  • Chenggang Yan

As one of the fundamental components of object detection, intersection-over-union (IoU) calculations between two bounding boxes play an important role in samples selection, NMS operation and evaluation of object detection algorithms. This procedure is well-defined and solved for planar images, while it is challenging for spherical ones. Some existing methods utilize planar bounding boxes to represent spherical objects. However, they are biased due to the distortions of spherical objects. Others use spherical rectangles as unbiased representations, but they adopt excessive approximate algorithms when computing the IoU. In this paper, we propose an unbiased IoU as a novel evaluation criterion for spherical image object detection, which is based on the unbiased representations and utilize unbiased analytical method for IoU calculation. This is the first time that the absolutely accurate IoU calculation is applied to the evaluation criterion, thus object detection algorithms can be correctly evaluated for spherical images. With the unbiased representation and calculation, we also present Spherical CenterNet, an anchor free object detection algorithm for spherical images. The experiments show that our unbiased IoU gives accurate results and the proposed Spherical CenterNet achieves better performance on one real-world and two synthetic spherical object detection datasets than existing methods.

AAAI Conference 2021 Conference Paper

Automated Model Design and Benchmarking of Deep Learning Models for COVID-19 Detection with Chest CT Scans

  • Xin He
  • Shihao Wang
  • Xiaowen Chu
  • Shaohuai Shi
  • Jiangping Tang
  • Xin Liu
  • Chenggang Yan
  • Jiyong Zhang

The COVID-19 pandemic has spread globally for several months. Because its transmissibility and high pathogenicity seriously threaten people’s lives, it is crucial to accurately and quickly detect COVID-19 infection. Many recent studies have shown that deep learning (DL) based solutions can help detect COVID-19 based on chest CT scans. However, most existing work focuses on 2D datasets, which may result in low quality models as the real CT scans are 3D images. Besides, the reported results span a broad spectrum on different datasets with a relatively unfair comparison. In this paper, we first use three state-of-the-art 3D models (ResNet3D101, DenseNet3D121, and MC3 18) to establish the baseline performance on three publicly available chest CT scan datasets. Then we propose a differentiable neural architecture search (DNAS) framework to automatically search the 3D DL models for 3D chest CT scans classification and use the Gumbel Softmax technique to improve the search efficiency. We further exploit the Class Activation Mapping (CAM) technique on our models to provide the interpretability of the results. The experimental results show that our searched models (CovidNet3D) outperform the baseline human-designed models on three datasets with tens of times smaller model size and higher accuracy. Furthermore, the results also verify that CAM can be well applied in CovidNet3D for COVID- 19 datasets to provide interpretability for medical diagnosis. Code: https: //github. com/HKBU-HPML/CovidNet3D.

YNIMG Journal 2021 Journal Article

Learning dynamic graph embeddings for accurate detection of cognitive state changes in functional brain networks

  • Yi Lin
  • Defu Yang
  • Jia Hou
  • Chenggang Yan
  • Minjeong Kim
  • Paul J Laurienti
  • Guorong Wu

Mounting evidence shows that brain functions and cognitive states are dynamically changing even in the resting state rather than remaining at a single constant state. Due to the relatively small changes in BOLD (blood-oxygen-level-dependent) signals across tasks, it is difficult to detect the change of cognitive status without requiring prior knowledge of the experimental design. To address this challenge, we present a dynamic graph learning approach to generate an ensemble of subject-specific dynamic graph embeddings, which allows us to use brain networks to disentangle cognitive events more accurately than using raw BOLD signals. The backbone of our method is essentially a representation learning process for projecting BOLD signals into a latent vertex-temporal domain with the greater biological underpinning of brain activities. Specifically, the learned representation domain is jointly formed by (1) a set of harmonic waves that govern the topology of whole-brain functional connectivities and (2) a set of Fourier bases that characterize the temporal dynamics of functional changes. In this regard, our dynamic graph embeddings provide a new methodology to investigate how these self-organized functional fluctuation patterns oscillate along with the evolving cognitive status. We have evaluated our proposed method on both simulated data and working memory task-based fMRI datasets, where our dynamic graph embeddings achieve higher accuracy in detecting multiple cognitive states than other state-of-the-art methods.

IJCAI Conference 2020 Conference Paper

Real-World Automatic Makeup via Identity Preservation Makeup Net

  • Zhikun Huang
  • Zhedong Zheng
  • Chenggang Yan
  • Hongtao Xie
  • Yaoqi Sun
  • Jianzhong Wang
  • Jiyong Zhang

This paper focuses on the real-world automatic makeup problem. Given one non-makeup target image and one reference image, the automatic makeup is to generate one face image, which maintains the original identity with the makeup style in the reference image. In the real-world scenario, face makeup task demands a robust system against the environmental variants. The two main challenges in real-world face makeup could be summarized as follow: first, the background in real-world images is complicated. The previous methods are prone to change the style of background as well; second, the foreground faces are also easy to be affected. For instance, the ``heavy'' makeup may lose the discriminative information of the original identity. To address these two challenges, we introduce a new makeup model, called Identity Preservation Makeup Net (IPM-Net), which preserves not only the background but the critical patterns of the original identity. Specifically, we disentangle the face images to two different information codes, i. e. , identity content code and makeup style code. When inference, we only need to change the makeup style code to generate various makeup images of the target person. In the experiment, we show the proposed method achieves not only better accuracy in both realism (FID) and diversity (LPIPS) in the test set, but also works well on the real-world images collected from the Internet.

AAAI Conference 2019 Conference Paper

Dual-View Ranking with Hardness Assessment for Zero-Shot Learning

  • Yuchen Guo
  • Guiguang Ding
  • Jungong Han
  • Xiaohan Ding
  • Sicheng Zhao
  • Zheng Wang
  • Chenggang Yan
  • Qionghai Dai

Zero-shot learning (ZSL) is to build recognition models for previously unseen target classes which have no labeled data for training by transferring knowledge from some other related auxiliary source classes with abundant labeled samples to the target ones with class attributes as the bridge. The key is to learn a similarity based ranking function between samples and class labels using the labeled source classes so that the proper (unseen) class label for a test sample can be identified by the function. In order to learn the function, single-view ranking based loss is widely used which aims to rank the true label prior to the other labels for a training sample. However, we argue that the ranking can be performed from the other view, which aims to place the images belonging to a label before the images from the other classes. Motivated by it, we propose a novel DuAl-view RanKing (DARK) loss for zeroshot learning simultaneously ranking labels for an image by point-to-point metric and ranking images for a label by pointto-set metric, which is capable of better modeling the relationship between images and classes. In addition, we also notice that previous ZSL approaches mostly fail to well exploit the hardness of training samples, either using only very hard ones or using all samples indiscriminately. In this work, we also introduce a sample hardness assessment method to ZSL which assigns different weights to training samples based on their hardness, which leads to a more accurate and robust ZSL model. Experiments on benchmarks demonstrate that DARK outperforms the state-of-the-arts for (generalized) ZSL.

IJCAI Conference 2019 Conference Paper

Landmark Selection for Zero-shot Learning

  • Yuchen Guo
  • Guiguang Ding
  • Jungong Han
  • Chenggang Yan
  • Jiyong Zhang
  • Qionghai Dai

Zero-shot learning (ZSL) is an emerging research topic whose goal is to build recognition models for previously unseen classes. The basic idea of ZSL is based on heterogeneous feature matching which learns a compatibility function between image and class features using seen classes. The function is constructed based on one-vs-all training in which each class has only one class feature and many image features. Existing ZSL works mostly treat all image features equivalently. However, in this paper we argue that it is more reasonable to use some representative cross-domain data instead of all. Motivated by this idea, we propose a novel approach, termed as Landmark Selection(LAST) for ZSL. LAST is able to identify representative cross-domain features which further lead to better image-class compatibility function. Experiments on several ZSL datasets including ImageNet demonstrate the superiority of LAST to the state-of-the-arts.

AAAI Conference 2019 Conference Paper

Recurrent Attention Model for Pedestrian Attribute Recognition

  • Xin Zhao
  • Liufang Sang
  • Guiguang Ding
  • Jungong Han
  • Na Di
  • Chenggang Yan

Pedestrian attribute recognition is to predict attribute labels of pedestrian from surveillance images, which is a very challenging task for computer vision due to poor imaging quality and small training dataset. It is observed that many semantic pedestrian attributes to be recognised tend to show spatial locality and semantic correlations by which they can be grouped while previous works mostly ignore this phenomenon. Inspired by Recurrent Neural Network (RNN)’s super capability of learning context correlations and Attention Model’s capability of highlighting the region of interest on feature map, this paper proposes end-to-end Recurrent Convolutional (RC) and Recurrent Attention (RA) models, which are complementary to each other. RC model mines the correlations among different attribute groups with convolutional LSTM unit, while RA model takes advantage of the intra-group spatial locality and inter-group attention correlation to improve the performance of pedestrian attribute recognition. Our RA method combines the Recurrent Learning and Attention Model to highlight the spatial position on feature map and mine the attention correlations among different attribute groups to obtain more precise attention. Extensive empirical evidence shows that our recurrent model frameworks achieve state-of-the-art results, based on pedestrian attribute datasets, i. e. standard PETA and RAP datasets.

IJCAI Conference 2018 Conference Paper

Watching a Small Portion could be as Good as Watching All: Towards Efficient Video Classification

  • Hehe Fan
  • Zhongwen Xu
  • Linchao Zhu
  • Chenggang Yan
  • Jianjun Ge
  • Yi Yang

We aim to significantly reduce the computational cost for classification of temporally untrimmed videos while retaining similar accuracy. Existing video classification methods sample frames with a predefined frequency over entire video. Differently, we propose an end-to-end deep reinforcement approach which enables an agent to classify videos by watching a very small portion of frames like what we do. We make two main contributions. First, information is not equally distributed in video frames along time. An agent needs to watch more carefully when a clip is informative and skip the frames if they are redundant or irrelevant. The proposed approach enables the agent to adapt sampling rate to video content and skip most of the frames without the loss of information. Second, in order to have a confident decision, the number of frames that should be watched by an agent varies greatly from one video to another. We incorporate an adaptive stop network to measure confidence score and generate timely trigger to stop the agent watching videos, which improves efficiency without loss of accuracy. Our approach reduces the computational cost significantly for the large-scale YouTube-8M dataset, while the accuracy remains the same.

v2026.09.13