Arrow Research search

Author name cluster

Kunpeng Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

7 papers
2 author rows

Possible papers

7

AAAI Conference 2026 Conference Paper

Taming Cascaded Mixture-of-Experts for Modality-missing Multi-modal Salient Object Detection

  • Kunpeng Wang
  • Feifan Sun
  • Keke Chen

Multi-modal Salient Object Detection (SOD) shows an improvement over its uni-modal counterpart by exploiting the complementary benefits between modalities. However, this improvement relies on complete multi-modal information, which is difficult to be guaranteed in practice due to sensor failures and transmission errors. To address this issue, we propose a robust multi-modal SOD framework that enhances the adaptability to modality-missing conditions, while maintaining comparable performance in the modality-complete condition. Nevertheless, flexibly handling modality-missing and modality-complete cases and integrating their corresponding multi-modal features in a unified framework is non-trivial. To this end, we achieve this framework by designing a Cascaded Mixture-of-Experts (CMoE) network that sequentially incorporates missing-aware and multi-modal MoE. Specifically, the missing-aware MoE employs three modality-reconstruction experts with a soft router to adaptively reconstruct feature representations for both missing and available modalities, assisted by an expert modulation loss that guides the router to assign expert weights according to missing conditions. The multi-modal MoE adopts two homogeneous uni-modal experts with learned modality-specific knowledge tailored for integrating modality features, which are dynamically combined via the soft router. The cascaded architecture fully empowers CMoE with the flexibility across varying input cases. Extensive experiments on modality-missing and modality-complete conditions demonstrate the effectiveness of the proposed method.

AAAI Conference 2025 Conference Paper

Alignment-Free RGB-T Salient Object Detection: A Large-Scale Dataset and Progressive Correlation Network

  • Kunpeng Wang
  • Keke Chen
  • Chenglong Li
  • Zhengzheng Tu
  • Bin Luo

Alignment-free RGB-Thermal (RGB-T) salient object detection (SOD) aims to achieve robust performance in complex scenes by directly leveraging the complementary information from unaligned visible-thermal image pairs, without requiring manual alignment. However, the labor-intensive process of collecting and annotating image pairs limits the scale of existing benchmarks, hindering the advancement of alignment-free RGB-T SOD. In this paper, we construct a large-scale and high-diversity unaligned RGB-T SOD dataset named UVT20K, comprising 20,000 image pairs, 407 scenes, and 1256 object categories. All samples are collected from real-world scenarios with various challenges, such as low illumination, image clutter, complex salient objects, and so on. To support the exploration for further research, each sample in UVT20K is annotated with a comprehensive set of ground truths, including saliency masks, scribbles, boundaries, and challenge attributes. In addition, we propose a Progressive Correlation Network (PCNet), which models inter- and intra-modal correlations on the basis of explicit alignment to achieve accurate predictions in unaligned image pairs. Extensive experiments conducted on two unaligned three weakly aligned three aligned datasets demonstrate the effectiveness of our method.

ICML Conference 2025 Conference Paper

EmoGrowth: Incremental Multi-label Emotion Decoding with Augmented Emotional Relation Graph

  • Kaicheng Fu
  • Changde Du
  • Jie Peng
  • Kunpeng Wang
  • Shuangchen Zhao
  • Xiaoyu Chen
  • Huiguang He

Emotion recognition systems face significant challenges in real-world applications, where novel emotion categories continually emerge and multiple emotions often co-occur. This paper introduces multi-label fine-grained class incremental emotion decoding, which aims to develop models capable of incrementally learning new emotion categories while maintaining the ability to recognize multiple concurrent emotions. We propose an Augmented Emotional Semantics Learning (AESL) framework to address two critical challenges: past- and future-missing partial label problems. AESL incorporates an augmented Emotional Relation Graph (ERG) for reliable soft label generation and affective dimension-based knowledge distillation for future-aware feature learning. We evaluate our approach on three datasets spanning brain activity and multimedia domains, demonstrating its effectiveness in decoding up to 28 fine-grained emotion categories. Results show that AESL significantly outperforms existing methods while effectively mitigating catastrophic forgetting. Our code is available at https: //github. com/ChangdeDu/EmoGrowth.

EAAI Journal 2025 Journal Article

Erasure-based interaction network for red-green-blue and thermal object detection and a unified benchmark

  • Qishun Wang
  • Zhengzheng Tu
  • Chenglong Li
  • Hongshun Wang
  • Kunpeng Wang

Recently, many breakthroughs have been made in the field of video object detection, but the performance is still limited due to the imaging limitations of RGB (red-green-blue) sensors in adverse illumination conditions. To alleviate this issue, this work introduces a new computer vision task called RGBT (red-green-blue and thermal) video object detection by introducing the thermal modality that is insensitive to adverse illumination conditions. To promote the research and development of RGBT video object detection, we design a novel Erasure-based Interaction Network (EINet) and establish a comprehensive benchmark dataset for this task. Traditional methods often leverage temporal information by using many auxiliary frames, and thus have a large computational burden. Considering thermal images exhibit less noise than RGB ones, we develop a negative activation function that is used to erase the noise of RGB features with the help of thermal image features. Furthermore, with the benefits from thermal images, we rely only on a small temporal window to model the spatio temporal information to greatly improve efficiency while maintaining detection accuracy. Our dataset consists of 50 pairs of RGBT video sequences with complex backgrounds, various objects and different illuminations, which are collected in real traffic scenarios. Extensive experiments on the proposed dataset demonstrate the effectiveness and efficiency of EINet. Compared with existing detectors, EINet achieves a relatively balanced performance with a detection accuracy of 46. 3% and a speed of 92. 6 frames per second. This project will be released to the public for free academic usage at https: //github. com/tzz-ahu.

TCS Journal 2025 Journal Article

Full domain functional bootstrapping using the prime cyclotomic ring

  • Ruida Wang
  • Xianhui Lu
  • Yundi Wen
  • Zhihao Li
  • Benqiang Wei
  • Kunpeng Wang
  • Lixia Luo

Functional bootstrapping can evaluate a look up table (LUT) almost for free during refreshing ciphertexts, making fully homomorphic encryption (FHE) more efficient. But the original functional bootstrapping requires the LUT to be negacyclic, or the plaintext should be distributed in the first half torus. This limitation lies in the selection of the negacyclic cyclotomic ring in the blind rotation procedure, which greatly restricts the power of functional bootstrapping. Full domain functional bootstrapping (FDFB) is more attractive. While the plaintext can be distributed throughout the whole torus, it lends us an opportunity to use the natural homomorphism of Regev encryption to evaluate affine functions for free, while maintaining the efficiency of the LUT evaluation. In this paper, we analyze in depth the manifestation of blind rotation on the prime cyclotomic ring, and use this structure to turn the traditional functional bootstrapping into the full domain. Compared with existing full domain solutions, our approach achieves the lowest noise growth and the fewest polynomial multiplications, since the expensive blind rotation procedure only needs to be called once while other solutions require at least twice. Furthermore, our scheme presents good scalability, which can be extended to more functional variants, such as private FDFB, multi-value FDFB and high precision FDFB.

EAAI Journal 2025 Journal Article

Hierarchical semantics guided multi-scale correlation network for alignment-free red-green-blue and thermal salient object detection

  • Chengmei Han
  • Lei Liu
  • Kunpeng Wang
  • Fei Xie
  • Bing Wei

RGBT (red-green-blue and thermal) salient object detection (SOD) aims to identify and highlight the most visually salient objects in an image by leveraging the complementary information from both RGB and thermal (TIR) modalities. It is particularly effective for 24/7 intelligent surveillance and autonomous perception in smart city security and traffic monitoring, especially under low light and adverse weather. However, existing methods primarily rely on manually aligned datasets, which are limited in handling the challenges posed by unaligned multi-modal data in real-world applications. Furthermore, these methods usually extract complementary information from both modalities using fixed-size windows (Liuet al. , 2022, Wanget al. , 2024b). However, such fixed-size windows are not effective in dealing with unaligned multi-modal images due to spatial inconsistencies. Additionally, existing methods often use single-layer high-level feature to represent semantic information, which fails to fully exploit the complementary benefits of multi-level features, thereby reducing the effectiveness of semantic guidance. To address these challenges, we propose a Hierarchical Semantics guided Multi-scale correlation Network (HSMNet) for alignment-free RGBT SOD. A Hierarchical Semantic Fusion Module (HSFM) dynamically assigns weights to features from multiple levels, enabling adaptive fusion of multi-level semantic information. A Multi-scale Asymmetric Correlation Module (MACM) employs windows of various sizes to capture asymmetric correlations between unaligned multi-modal data, enhancing cross-modal complementary information extraction even when data are not perfectly aligned. We conduct extensive experiments on unaligned, weakly aligned and aligned RGBT SOD datasets, with results demonstrating that our method outperforms state-of-the-art algorithms, achieving superior accuracy and robustness in both unaligned and weakly aligned RGBT SOD scenarios.

EAAI Journal 2023 Journal Article

Multimodal salient object detection via adversarial learning with collaborative generator

  • Zhengzheng Tu
  • Wenfang Yang
  • Kunpeng Wang
  • Amir Hussain
  • Bin Luo
  • Chenglong Li

Multimodal salient object detection(MSOD), which utilizes multimodal information (e. g. , RGB image and thermal infrared or depth image) to detect common salient objects, has received much attention recently. Different modalities reflect different appearance properties of salient objects, some of which could contribute to improving the precision and/or recall of MSOD. To greatly improve both Precision and Recall by fully exploring multimodal data, in this work, we propose an effective adversarial learning framework based on a novel collaborative generator for accurate multimodal salient object detection. In particular, the collaborative generator consists of three generators (generator1, generator2 and generator3), which aim at decreasing the false positive and false negative of the generated saliency maps and improving F-measure of the final saliency maps respectively. Generator1 and generator2 contain two encoder–decoder networks for multimodal inputs, and we propose a new co-attention model to perform adaptive interactions between different modalities. Furthermore, we apply generator3 to integrate feature maps from generator1 and generator2 in a complementary way. Through adversarially learning the collaborative generator and discriminator, both Precision and Recall of the predicted maps are boosted with the complementary benefits of multimodal data. Extensive experiments on three RGBT datasets and six RGBD datasets show that our method performs quite well against state-of-the-art MSOD methods.

v2026.09.13