Arrow Research search

Author name cluster

Ping Hu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

14 papers
1 author row

Possible papers

14

AAAI Conference 2026 Conference Paper

Cross-Domain Few-Shot Learning via Multi-View Collaborative Optimization with Vision-Language Models

  • Dexia Chen
  • Wentao Zhang
  • Qianjie Zhu
  • Ping Hu
  • Weibing Li
  • Tong Zhang
  • Ruixuan Wang

Vision-language models (VLMs) pre-trained on natural image and language data, such as CLIP, have exhibited significant potential in few-shot image recognition tasks, leading to development of various efficient transfer learning methods. These methods exploit inherent pre-learned knowledge in VLMs and have achieved strong performance on standard image datasets. However, their effectiveness is often limited when confronted with cross-domain tasks where imaging domains differ from natural images. To address this limitation, we propose Consistency-guided Multi-view Collaborative Optimization (CoMuCo), a novel fine-tuning strategy for VLMs. This strategy employs two functionally complementary expert modules to extract multi-view features, while incorporating prior knowledge-based consistency constraints and information geometry-based consensus mechanisms to enhance the robustness of feature learning. Additionally, a new cross-domain few-shot benchmark is established to help comprehensively evaluate methods on imaging domains distinct from natural images. Extensive empirical evaluations on both existing and newly proposed benchmarks suggest CoMuCo consistently outperforms current methods.

AAAI Conference 2026 Conference Paper

Graph Smoothing for Enhanced Local Geometry Learning in Point Cloud Analysis

  • Shangbo Yuan
  • Jie Xu
  • Ping Hu
  • Xiaofeng Zhu
  • Na Zhao

Graph-based methods have proven to be effective in capturing relationships among points for 3D point cloud analysis. However, these methods often suffer from suboptimal graph structures, particularly due to sparse connections at boundary points and noisy connections in junction areas. To address these challenges, we propose a novel method that integrates a graph smoothing module with an enhanced local geometry learning module. Specifically, we identify the limitations of conventional graph structures, particularly in handling boundary points and junction areas. In response, we introduce a graph smoothing module designed to optimize the graph structure and minimize the negative impact of unreliable sparse and noisy connections. Based on the optimized graph structure, we improve the feature extract function with local geometry information. These include shape features derived from adaptive geometric descriptors based on eigenvectors and distribution features obtained through cylindrical coordinate transformation. Experimental results on real-world datasets validate the effectiveness of our method in various point cloud learning tasks, i.e., classification, part segmentation, and semantic segmentation.

EAAI Journal 2025 Journal Article

An attention-guided multi-scale feature cascade network for underwater fish counting

  • Hanyu Zhang
  • Mengping Dong
  • Fei Li
  • Zhenbo Li
  • Ping Hu

Visual counting is essential for advancing fisheries intelligence, but fish scale variation in open underwater environments has made underwater fish counting a constant challenge. Therefore, we propose an Attention-guided Multi-scale Feature Cascade Network, named AMFCNet, which resolves scale variation and improves the accuracy of fish counting in complex underwater environments. AMFCNet utilizes a multi-scale attention gate for multi-scale feature fusion, and integrates a multi-scale convolution module to capture complex spatial relationships. It also employs a multi-head supervision fusion strategy to mask irrelevant regions, ensuring targeted learning for each scale and generating high-quality multi-scale density maps. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on the proposed dataset with the lowest computational cost, significantly outperforming 11 mainstream counting methods. It also achieves excellent results on other publicly available underwater datasets, with Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Normalized Absolute Error (NAE) values of 1. 26, 1. 71, and 0. 08, respectively. This method shows significant potential for practical applications in aquaculture, such as in marine ranching and pond farming, to assess fish growth conditions and adjust feeding strategies accordingly.

AAAI Conference 2025 Conference Paper

Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding

  • Wenbo Zhang
  • Lu Zhang
  • Ping Hu
  • Liqian Ma
  • Yunzhi Zhuge
  • Huchuan Lu

Injecting semantics into 3D Gaussian Splatting (3DGS) has recently garnered significant attention. While current approaches typically distill 3D semantic features from 2D foundational models (e.g., CLIP and SAM) to facilitate novel view segmentation and semantic understanding, their heavy reliance on 2D supervision can undermine cross-view semantic consistency and necessitate complex data preparation processes, therefore hindering view-consistent scene understanding. In this work, we present FreeGS, an unsupervised semantic-embedded 3DGS framework that achieves view-consistent 3D scene understanding without the need for 2D labels. Instead of directly learning semantic features, we introduce the IDentity-coupled Semantic Field (IDSF) into 3DGS, which captures both semantic representations and view-consistent instance indices for each Gaussian. We optimize IDSF with a two-step alternating strategy: semantics help to extract coherent instances in 3D space, while the resulting instances regularize the injection of stable semantics from 2D space. Additionally, we adopt a 2D-3D joint contrastive loss to enhance the complementarity between view-consistent 3D geometry and rich semantics during the bootstrapping process, enabling FreeGS to uniformly perform tasks such as novel-view semantic segmentation, object selection, and 3D object detection. Extensive experiments on LERF-Mask, 3D-OVS, and ScanNet datasets demonstrate that FreeGS performs comparably to state-of-the-art methods while avoiding the complex data preprocessing workload.

NeurIPS Conference 2025 Conference Paper

FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning

  • Lu Zhang
  • Jiazuo Yu
  • Haomiao Xiong
  • Ping Hu
  • Yunzhi Zhuge
  • Huchuan Lu
  • You He

Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and localizing visual details in high-resolution images---particularly when dealing with extra-small objects embedded in cluttered contexts. To address this issue, we propose FineRS, a two-stage MLLM-based reinforcement learning framework for jointly reasoning and segmenting extremely small objects within high-resolution scenes. FineRS adopts a coarse-to-fine pipeline comprising Global Semantic Exploration (GSE) and Localized Perceptual Refinement (LPR). Specifically, GSE performs instruction-guided reasoning to generate a textural response and a coarse target region, while LPR refines this region to produce an accurate bounding box and segmentation mask. To couple the two stages, we introduce a locate-informed retrospective reward, where LPR's outputs are used to optimize GSE for more robust coarse region exploration. Additionally, we present FineRS-4k, a new dataset for evaluating MLLMs on attribute-level reasoning and pixel-level segmentation on subtle, small-scale targets in complex high-resolution scenes. Experimental results on FineRS-4k and public datasets demonstrate that our method consistently outperforms state-of-the-art MLLM-based approaches on both instruction-guided segmentation and visual reasoning tasks.

IJCAI Conference 2025 Conference Paper

Seeking Proxy Point via Stable Feature Space for Noisy Correspondence Learning

  • Yucheng Xie
  • Songyue Cai
  • Tao Tong
  • Ping Hu
  • Xiaofeng Zhu

To meet the growing demand for cross-modal training data, directly collecting multimodal data from the Internet has become prevalent. However, such data inevitably suffer from Noisy Correspondence. Previous works focused on recasting soft labels to mitigate noise's negative impact. We explore a novel perspective to solve this problem: pursuing proxy representation for noisy data to enable reliable feature learning. To this end, we propose a novel framework: Seeking Proxy Point via Stable Feature Space (SPS). This framework employs a fine-grained partitioning strategy to obtain a high-confidence reliable set. By imposing intermodal cross-transformation consistency constraints and intramodal metric consistency constraints, a stable feature space is constructed. Building on this foundation, SPS seeks proxy points for noisy data, enabling even noisy data to be accurately embedded into appropriate positions within the feature space. Combined with partial alignment for partially matched data pairs, SPS ultimately achieves robust learning under Noisy Correspondence. Experiments on three widely used cross-modal datasets demonstrate that SPS significantly outperforms previous methods. Our code is available at https: //github. com/C-TeaRanger/SPS.

IJCAI Conference 2024 Conference Paper

Exploring the Role of Node Diversity in Directed Graph Representation Learning

  • Jincheng Huang
  • Yujie Mo
  • Ping Hu
  • Xiaoshuang Shi
  • Shangbo Yuan
  • Zeyu Zhang
  • Xiaofeng Zhu

Many methods of Directed Graph Neural Networks (DGNNs) are designed to equally treat nodes in the same neighbor set (i. e. , out-neighbor set and in-neighbor set) for every node, without considering the node diversity in directed graphs, so they are often unavailable to adaptively acquire suitable information from neighbors of different directions. To alleviate this issue, in this paper, we investigate a new way to first consider node diversity for representation learning on directed graphs, i. e. , neighbor diversity and degree diversity, and then propose a new NDDGNN framework to adaptively assign weights to both outgoing information and incoming information at the node level. Extensive experiments on seven real-world datasets validate the superior performance of our method compared to state-of-the-art methods in terms of both node classification and link prediction tasks.

IJCAI Conference 2024 Conference Paper

Towards Dynamic-Prompting Collaboration for Source-Free Domain Adaptation

  • Mengmeng Zhan
  • Zongqian Wu
  • Rongyao Hu
  • Ping Hu
  • Heng Tao Shen
  • Xiaofeng Zhu

In domain adaptation, challenges such as data privacy constraints can impede access to source data, catalyzing the development of source-free domain adaptation (SFDA) methods. However, current approaches heavily rely on models trained on source data, posing the risk of overfitting and suboptimal generalization. This paper introduces a dynamic prompt learning paradigm that harnesses the power of large-scale vision-language models to enhance the semantic transfer of source models. Specifically, our approach fosters robust and adaptive collaboration between the source-trained model and the vision-language model, facilitating the reliable extraction of domain-specific information from unlabeled target data, while consolidating domain-invariant knowledge. Without the need for accessing source data, our method amalgamates the strengths inherent in both traditional SFDA approaches and vision-language models, formulating a collaborative framework for addressing SFDA challenges. Extensive experiments conducted on three benchmark datasets showcase the superiority of our framework over previous SOTA methods.

YNIMG Journal 2023 Journal Article

Deep learning-assisted identification and quantification of aneurysmal subarachnoid hemorrhage in non-contrast CT scans: Development and external validation of Hybrid 2D/3D UNet

  • Ping Hu
  • Haizhu Zhou
  • Tengfeng Yan
  • Hongping Miu
  • Feng Xiao
  • Xinyi Zhu
  • Lei Shu
  • Shuang Yang

Accurate stroke assessment and consequent favorable clinical outcomes rely on the early identification and quantification of aneurysmal subarachnoid hemorrhage (aSAH) in non-contrast computed tomography (NCCT) images. However, hemorrhagic lesions can be complex and difficult to distinguish manually. To solve these problems, here we propose a novel Hybrid 2D/3D UNet deep-learning framework for automatic aSAH identification and quantification in NCCT images. We evaluated 1824 consecutive patients admitted with aSAH to four hospitals in China between June 2018 and May 2022. Accuracy and precision, Dice scores and intersection over union (IoU), and interclass correlation coefficients (ICC) were calculated to assess model performance, segmentation performance, and correlations between automatic and manual segmentation, respectively. A total of 1355 patients with aSAH were enrolled: 931, 101, 179, and 144 in four datasets, of whom 326 were scanned with Siemens, 640 with Philips, and 389 with GE Medical Systems scanners. Our proposed deep-learning method accurately identified (accuracies 0.993-0.999) and segmented (Dice scores 0.550-0.897) hemorrhage in both the internal and external datasets, even combinations of hemorrhage subtypes. We further developed a convenient AI-assisted platform based on our algorithm to assist clinical workflows, whose performance was comparable to manual measurements by experienced neurosurgeons (ICCs 0.815-0.957) but with greater efficiency and reduced cost. While this tool has not yet been prospectively tested in clinical practice, our innovative hybrid network algorithm and platform can accurately identify and quantify aSAH, paving the way for fast and cheap NCCT interpretation and a reliable AI-based approach to expedite clinical decision-making for aSAH patients.

NeurIPS Conference 2023 Conference Paper

Joint Attribute and Model Generalization Learning for Privacy-Preserving Action Recognition

  • Duo Peng
  • Li Xu
  • Qiuhong Ke
  • Ping Hu
  • Jun Liu

Privacy-Preserving Action Recognition (PPAR) aims to transform raw videos into anonymous ones to prevent privacy leakage while maintaining action clues, which is an increasingly important problem in intelligent vision applications. Despite recent efforts in this task, it is still challenging to deal with novel privacy attributes and novel privacy attack models that are unavailable during the training phase. In this paper, from the perspective of meta-learning (learning to learn), we propose a novel Meta Privacy-Preserving Action Recognition (MPPAR) framework to improve both generalization abilities above (i. e. , generalize to novel privacy attributes and novel privacy attack models ) in a unified manner. Concretely, we simulate train/test task shifts by constructing disjoint support/query sets w. r. t. privacy attributes or attack models. Then, a virtual training and testing scheme is applied based on support/query sets to provide feedback to optimize the model's learning toward better generalization. Extensive experiments demonstrate the effectiveness and generalization of the proposed framework compared to state-of-the-arts.

NeurIPS Conference 2022 Conference Paper

DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited Annotations

  • Ximeng Sun
  • Ping Hu
  • Kate Saenko

Solving multi-label recognition (MLR) for images in the low-label regime is a challenging task with many real-world applications. Recent work learns an alignment between textual and visual spaces to compensate for insufficient image labels, but loses accuracy because of the limited amount of available MLR annotations. In this work, we utilize the strong alignment of textual and visual features pretrained with millions of auxiliary image-text pairs and propose \textit{Dual Context Optimization} (DualCoOp) as a unified framework for partial-label MLR and zero-shot MLR. \ours encodes positive and negative contexts with class names as part of the linguistic input (i. e. prompts). Since \ours only introduces a very light learnable overhead upon the pretrained vision-language framework, it can quickly adapt to multi-label recognition tasks that have limited annotations and even unseen classes. Experiments on standard multi-label recognition benchmarks across two challenging low-label settings demonstrate the advantages of our approach over state-of-the-art methods. Our code will be publicly available. Project page: https: //cs-people. bu. edu/sunxm/DualCoOp/project. html

AAAI Conference 2021 Conference Paper

Static-Dynamic Interaction Networks for Offline Signature Verification

  • Huan Li
  • Ping Wei
  • Ping Hu

Offline signature verification is a challenging issue that is widely used in various fields. Previous approaches model this task as a static feature matching or distance metric problem of two images. In this paper, we propose a novel Static- Dynamic Interaction Network (SDINet) model which introduces sequential representation into static signature images. A static signature image is converted to sequences by assuming pseudo dynamic processes in the static image. A static representation extracting deep features from signature images describes the global information of signatures. A dynamic representation extracting sequential features with LSTM networks characterizes the local information of signatures. A dynamic-to-static attention is learned from the sequences to refine the static features. Through the static-to-dynamic conversion and the dynamic-to-static attention, the static representation and dynamic representation are unified into a compact framework. The proposed method was evaluated on four popular datasets of different languages. The extensive experimental results manifest the strength of our model.

YNIMG Journal 2021 Journal Article

The effect of eye gaze direction on emotional mimicry: A multimodal study with electromyography and electroencephalography

  • Beibei Kuang
  • Xueting Li
  • Xintong Li
  • Mingxiao Lin
  • Shanrou Liu
  • Ping Hu

Emotional mimicry plays an important role in social interaction and is influenced by social context, especially eye gaze direction. However, the neural mechanism underlying the effect of eye gaze direction on emotional mimicry is unclear. Here, we explored how eye gaze direction influenced emotional mimicry with a combination of electromyography (EMG) and electroencephalography (EEG) techniques, which may provide a more comprehensive measure. To do this, we recorded facial EMG and scalp EEG signals simultaneously while participants observed emotional faces (happy vs. angry) with direct or averted gaze. Then, we split the EEG trials into two mimicry intensity categories (high mimicry intensity, HMI vs. low mimicry intensity, LMI) according to EMG activity. The ERP difference between HMI and LMI EEG trials revealed four ERP components (P50, P150, N200 and P300), and the effect of eye gaze direction on emotional mimicry was prominent on P300 at P7 and P8. Moreover, we also observed differences in the effect of eye gaze direction on mimicry of happy faces and angry faces, which were found on P300 at P7, as well as P150 at P7 and N200 at P7 and Pz. In short, the present study isolated the neural signals of emotional mimicry with a new multimodal method, and provided empirical neural evidence that eye gaze direction affected emotional mimicry.

NeurIPS Conference 2020 Conference Paper

Uncertainty-Aware Learning for Zero-Shot Semantic Segmentation

  • Ping Hu
  • Stan Sclaroff
  • Kate Saenko

Zero-shot semantic segmentation (ZSS) aims to classify pixels of novel classes without training examples available. Recently, most ZSS methods focus on learning the visual-semantic correspondence to transfer knowledge from seen classes to unseen classes at the pixel level. Yet, few works study the adverse effects caused by the noisy and outlying training samples in the seen classes. In this paper, we identify this challenge and address it with a novel framework that learns to discriminate noisy samples based on Bayesian uncertainty estimation. Specifically, we model the network outputs with Gaussian and Laplacian distributions, with the variances accounting for the observation noise as well as the uncertainty of input samples. Learning objectives are then derived with the estimated variances playing as adaptive attenuation for individual samples in training. Consequently, our model learns more attentively from representative samples of seen classes while suffering less from noisy and outlying ones, thus providing better reliability and generalization toward unseen categories. We demonstrate the effectiveness of our framework through comprehensive experiments on multiple challenging benchmarks, and show that our method achieves significant accuracy improvement over previous approaches for large open-set segmentation.

v2026.09.13