Arrow Research search

Author name cluster

Lihua Zhou

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
2 author rows

Possible papers

6

EAAI Journal 2025 Journal Article

A visual data unsupervised disentangled representation learning framework: Contrast disentanglement based on variational auto-encoder

  • Chengquan Huang
  • Jianghai Cai
  • Senyan Luo
  • Shunxia Wang
  • Guiyan Yang
  • Huan Lei
  • Lihua Zhou

To discover and learn interpretable factors behind the visual data, many approaches use extra regularization terms in learning disentangled representations, which lead to poor results between disentanglement and generative quality. The variational auto-encoder has the ability to learn more semantic information, and the traversal of the generated images along different factor directions shows meaningful and interpretable variations in the latent space. Therefore, we exploit the scalability and training stability of the variational auto-encoder, and focus on meaningful traversal directions in the latent space. We propose contrast disentanglement based on variational auto-encoder, a visual data unsupervised disentangled representation learning framework. Specifically, we explore meaningful interpretable directions in the latent space by constructing the encoder of the typical network module, and then obtain interpretable directions for candidate traversals of target variation images. In this way, the interpretable directions with rich semantic information and disentanglement properties can be obtained. Furthermore, unlike the existing methods, to further improve the ability of interaction and collaborative learning between latent factors and learn more robust and generalized representations, we design the variation space based on disentangled encoders from the contrastive learning perspective to simulate the various variation of the image data, then extract disentangled representations and generate images. Extensive experiments on six disentanglement datasets have demonstrated that the proposed method achieves competing performance on both quantitative metrics and visual quality. Among them, the proposed method achieves better factor variational auto-encoder score and β-variational auto-encoder score (0. 97 ± 0. 04 and 0. 99 ± 0. 01, respectively) on the 3Dshapes dataset compared to existing methods.

NeurIPS Conference 2025 Conference Paper

Multimodal Causal Reasoning for UAV Object Detection

  • Nianxin Li
  • Mao Ye
  • Lihua Zhou
  • Shuaifeng Li
  • Song Tang
  • Luping Ji
  • Ce Zhu

Unmanned Aerial Vehicle (UAV) object detection faces significant challenges due to complex environmental conditions and different imaging conditions. These factors introduce significant changes in scale and appearance, particularly for small objects that occupy limited pixels and exhibit limited information, complicating detection tasks. To address these challenges, we propose a Multimodel Causal Reasoning framework based on YOLO backbone for UAV Object Detection (MCR-UOD). The key idea is to use the backdoor adjustment to discover the condition-invariant object representation for easy detection. Specifically, the YOLO backbone is first adjusted to incorporate the pre-trained vision-language model. The original category labels are replaced with semantic text prompts, and the detection head is replaced with text-image contrastive learning. Based on this backbone, our method consists of two parts. The first part, named language guided region exploration, discovers the regions with high probability of object existence using text embeddings based on vision-language model such as CLIP. Another part is the backdoor adjustment casual reasoning module, which constructs a confounder dictionary tailored to different imaging conditions to capture global image semantics and derives a prior probability distribution of shooting conditions. During causal inference, we use the confounder dictionary and the prior to intervene on local instance features, disentangling condition variations, and obtaining condition-invariant representations. Experimental results on several public datasets confirm the state-of-the-art performance of our approach. The code, data and models will be released upon publication of this paper.

AAAI Conference 2025 Conference Paper

Self-Prompting Analogical Reasoning for UAV Object Detection

  • Nianxin Li
  • Mao Ye
  • Lihua Zhou
  • Song Tang
  • Yan Gan
  • Zizhuo Liang
  • Xiatian Zhu

Unmanned Aerial Vehicle Object Detection (UAVOD) presents unique challenges due to varying altitudes, dynamic backgrounds, and the small size of objects. Traditional detection methods often struggle with these challenges, as they typically rely on visual feature only and fail to extract the semantic relations between the objects. To address these limitations, we propose a novel approach named Self-Prompting Analogical Reasoning (SPAR). Our method utilizes the vision-language model (CLIP) to generate context-aware prompts based on image feature, providing rich semantic information that guides analogical reasoning. SPAR includes two main modules: self-prompting and analogical reasoning. Self-prompting module based on learnable description and CLIP-text encoder generates context-aware prompt by combining specific image feature; then an objectness prompt score map is produced by computing the similarity between pixel-level features and context-aware prompt. With this score map, multi-scale image features are enhanced and pixel-level features are chosen for graph construction. While for analogical reasoning module, graph nodes consists of category-level prompt nodes and pixel-level image feature nodes. Analogical inference is based graph convolution. Under the guidance of category-level nodes, different-scale object features have been enhanced, which helps achieve more accurate detection of challenging objects. Extensive experiments illustrate that SPAR outperforms traditional methods, offering a more robust and accurate solution for UAVOD.

NeurIPS Conference 2024 Conference Paper

Cloud Object Detector Adaptation by Integrating Different Source Knowledge

  • Shuaifeng Li
  • Mao Ye
  • Lihua Zhou
  • Nianxin Li
  • Siying Xiao
  • Song Tang
  • Xiatian Zhu

We propose to explore an interesting and promising problem, Cloud Object Detector Adaptation (CODA), where the target domain leverages detections provided by a large cloud model to build a target detector. Despite with powerful generalization capability, the cloud model still cannot achieve error-free detection in a specific target domain. In this work, we present a novel Cloud Object detector adaptation method by Integrating different source kNowledge (COIN). The key idea is to incorporate a public vision-language model (CLIP) to distill positive knowledge while refining negative knowledge for adaptation by self-promotion gradient direction alignment. To that end, knowledge dissemination, separation, and distillation are carried out successively. Knowledge dissemination combines knowledge from cloud detector and CLIP model to initialize a target detector and a CLIP detector in target domain. By matching CLIP detector with the cloud detector, knowledge separation categorizes detections into three parts: consistent, inconsistent and private detections such that divide-and-conquer strategy can be used for knowledge distillation. Consistent and private detections are directly used to train target detector; while inconsistent detections are fused based on a consistent knowledge generation network, which is trained by aligning the gradient direction of inconsistent detections to that of consistent detections, because it provides a direction toward an optimal target detector. Experiment results demonstrate that the proposed COIN method achieves the state-of-the-art performance.

EAAI Journal 2024 Journal Article

Interval Type-2 enhanced possibilistic fuzzy C-means noisy image segmentation algorithm amalgamating weighted local information

  • Chengquan Huang
  • Huan Lei
  • Yang Chen
  • Jianghai Cai
  • Xiaosu Qin
  • Jialei Peng
  • Lihua Zhou
  • Lan Zheng

The fuzzy clustering algorithms based on interval type-2 are effective methods for data clustering and image segmentation with some potential advantages in dealing with higher-order uncertainty. However, the fuzzy clustering algorithms based on interval type-2 have limited accuracy in noisy image segmentation. Therefore, we propose a new noisy image segmentation algorithm based on weighted local information for interval type-2 enhanced possibilistic fuzzy C-means clustering. Firstly, a new possibilistic fuzzy weighted local information factor, which amalgamates the mutually guided image filtering, is devised to control the relationship between mutually guided image filtering and the original image utilizing the absolute difference image between the original image and mutually guided image filtering. Additionally, two weighted sum functions are also introduced into the new possibilistic fuzzy weighted local information factor to calculate the grayscale differences between the current pixel and its neighborhood pixels. Secondly, the distance of the objective function is corrected using mutually guided image filtering to obtain the local features and global information of the image. Finally, this paper introduces the new possibilistic fuzzy weighted local information factor into the interval type-2 possibilistic fuzzy C-means and proposes a new interval type-2 enhanced possibilistic fuzzy C-means clustering algorithm with local information and mutually guided image filtering constraints to improve noisy image segmentation results. Through extensive experiments on noisy synthetic images and noisy real images, the proposed algorithm can achieve higher performance and more accurate segmentation results compared with several existing algorithms.

ICRA Conference 2001 Conference Paper

Stiffness Estimation of a Tripod-based Parallel Kinematic Machine

  • Tian Huang
  • Jiangping Mei
  • Xingyu Zhao 0004
  • Lihua Zhou

This paper presents a simple yet comprehensive approach that enables the stiffness of a tripod-based parallel kinematic machine to be quickly estimated. The approach can be implemented by two steps. In the first step, the machine structure is decomposed into two substructures associated with the machine frame and the parallel mechanism. The stiffness model of each substructure is formulated by assuming that the components in the other substructure are rigid. This is followed by the second step that enables the stiffness model of the machine structure to be achieved by linear superposition. A 3D representation of the stiffness distributions within the usable workspace are depicted with comparison to those obtained through the finite element analysis.

v2026.09.13