Arrow Research search

Author name cluster

Aimin Hao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

10 papers
1 author row

Possible papers

10

AAAI Conference 2026 Conference Paper

DECON: Reconstruction of Clothed-Geometric Multiple Humans from a Single Image via Geometry-Guided Decoupling

  • Yiming Jiang
  • Wenfeng Song
  • Shuai Li
  • Aimin Hao

3D multi-human reconstruction from single images holds significant potential for advancing AR/VR applications. While remarkable progress has been made in single-human reconstruction, existing methods face challenges when reconstructing multiple humans. These challenges include: (1) severe inter-occlusion that disrupts individual body structures, and (2) the absence of physically plausible relative positioning among subjects. We present DECON, a novel DEcouple-and-reCONstruct framework that systematically addresses these limitations through two technical innovations: (1) a decouple-and-reconstruct framework with multi-view synthesis. It separates individuals and reconstructs detailed 3D bodies from a single image. (2) a Perspective-Aware Position Optimization (PAPO) approach. It ensures realistic positioning by fixing overlaps and gaps between subjects. Extensive experiments demonstrate our method's capability to reconstruct fully separated, anatomically complete 3D humans with clothed-geometric details and plausible interactions. Quantitative evaluations show a 54% reduction in Chamfer Distance and 35% in Point-to-Surface Distance compared to state-of-the-art methods.

AAAI Conference 2026 Conference Paper

IntentMotion: Learning Intent-Aware Human Motion from Language in 3D Scenes

  • Wenfeng Song
  • Shi Zheng
  • Xinyu Zhang
  • Xingliang Jin
  • Aimin Hao
  • Fei Hou
  • Xia Hou
  • Shuai Li

Generating human motion in complex 3D scenes from text is a challenging task with broad applications. However, existing methods often overlook realistic physical contact, resulting in visually plausible but physically unrealistic motion, e.g., penetration. To alleviate this, we propose IntentMotion, a novel framework that generates human motion in 3D scenes from natural language instructions by explicitly modeling intent. We first introduce the Intention-Guided Contact Field (IGCF). This differentiable voxel-based contact region representation explicitly aligns parsed language roles with spatial contact regions through a hierarchical attention mechanism. IGCF is jointly trained with a diffusion-based motion generator, allowing contact predictions to adapt dynamically through gradient feedback. To improve the controllability and physics-aware motion, we further propose an Intention-Aware Diffusion Model (IADM), which decouples the high-level semantic planning from the low-level contact refinement in a coarse-to-fine process. The optimized contact cues are utilized to guide the synthesis of a coarse trajectory, followed by refining detailed pose sequences under IGCF supervision. Experiments on the HUMANISE and LINGO datasets demonstrate that our IntentMotion outperforms recent baselines in contact accuracy, semantic alignment, and generalization to unseen scenes.

AAAI Conference 2025 Conference Paper

CtrlAvatar: Controllable Avatars Generation via Disentangled Invertible Networks

  • Wenfeng Song
  • Yang Ding
  • Fei Hou
  • Shuai Li
  • Aimin Hao
  • Xia Hou

As virtual experiences grow in popularity, the demand for realistic, personalized, and animatable human avatars increases. Traditional methods, relying on fixed templates, often produce costly avatars that lack expressiveness and realism. To overcome these challenges, we introduce Controllable Avatars generation via disentangled invertible networks (CtrlAvatar), a real-time framework for generating lifelike and customizable avatars. CtrlAvatar uses disentangled invertible networks to separate the deformation process into implicit body geometry and explicit texture components. This approach eliminates the need for repeated occupancy reconstruction, enabling detailed and coherent animations. The body geometry component ensures anatomical accuracy, while the texture component allows for complex, artifact-free clothing customization. This architecture ensures smooth integration between body movements and surface details. By optimizing transformations with position-varying offsets from the avatar’s initial Linear Blend Skinning vertices, CtrlAvatar achieves flexible, natural deformations that adapt to various scenarios. Extensive experiments show that CtrlAvatar outperforms other methods in quality, diversity, controllability, and cost-efficiency, marking a significant advancement in avatar generation.

AAAI Conference 2024 Conference Paper

Weakly Supervised Multimodal Affordance Grounding for Egocentric Images

  • Lingjing Xu
  • Yang Gao
  • Wenfeng Song
  • Aimin Hao

To enhance the interaction between intelligent systems and the environment, locating the affordance regions of objects is crucial. These regions correspond to specific areas that provide distinct functionalities. Humans often acquire the ability to identify these regions through action demonstrations and verbal instructions. In this paper, we present a novel multimodal framework that extracts affordance knowledge from exocentric images, which depict human-object interactions, as well as from accompanying textual descriptions that describe the performed actions. The extracted knowledge is then transferred to egocentric images. To achieve this goal, we propose the HOI-Transfer Module, which utilizes local perception to disentangle individual actions within exocentric images. This module effectively captures localized features and correlations between actions, leading to valuable affordance knowledge. Additionally, we introduce the Pixel-Text Fusion Module, which fuses affordance knowledge by identifying regions in egocentric images that bear resemblances to the textual features defining affordances. We employ a Weakly Supervised Multimodal Affordance (WSMA) learning approach, utilizing image-level labels for training. Through extensive experiments, we demonstrate the superiority of our proposed method in terms of evaluation metrics and visual results when compared to existing affordance grounding models. Furthermore, ablation experiments confirm the effectiveness of our approach. Code:https://github.com/xulingjing88/WSMA.

AAAI Conference 2023 Conference Paper

Pixel Is All You Need: Adversarial Trajectory-Ensemble Active Learning for Salient Object Detection

  • Zhenyu Wu
  • Lin Wang
  • Wei Wang
  • Qing Xia
  • Chenglizhao Chen
  • Aimin Hao
  • Shuo Li

Although weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by proving a hypothesis: there is a point-labeled dataset where saliency models trained on it can achieve equivalent performance when trained on the densely annotated dataset. To prove this conjecture, we proposed a novel yet effective adversarial trajectory-ensemble active learning (ATAL). Our contributions are three-fold: 1) Our proposed adversarial attack triggering uncertainty can conquer the overconfidence of existing active learning methods and accurately locate these uncertain pixels. 2) Our proposed trajectory-ensemble uncertainty estimation method maintains the advantages of the ensemble networks while significantly reducing the computational cost. 3) Our proposed relationship-aware diversity sampling algorithm can conquer oversampling while boosting performance. Experimental results show that our ATAL can find such a point-labeled dataset, where a saliency model trained on it obtained 97%-99% performance of its fully-supervised version with only 10 annotated points per image.

JBHI Journal 2022 Journal Article

Automatic Dental Plaque Segmentation Based on Local-to-Global Features Fused Self-Attention Network

  • Shuai Li
  • Yuting Guo
  • Zhennan Pang
  • Wenfeng Song
  • Aimin Hao
  • Bin Xia
  • Hong Qin

The accurate detection of dental plaque at an early stage will definitely prevent periodontal diseases and dental caries. However, it remains difficult for the current dental examination to accurately recognize dental plaque without using medical dyeing reagent due to the low contrast between dental plaque and healthy teeth. To combat this problem, this paper proposes a novel network enhanced by a self-attention module for intelligent dental plaque segmentation. The key motivation is to directly utilize oral endoscope images (bypassing the need for dyeing reagent) and get accurate pixel-level dental plaque segmentation results. The algorithm needs to conduct self-attention at the super-pixel level and fuse the super-pixels’ local-to-global features. Our newly-designed network architecture will afford the simultaneous fusion of multiple-scale complementary information guided by the powerful deep learning paradigm. The critical fused information includes the statistical distribution of the plaques color, the heat kernel signature (HKS) based local-to-global structure relationship, and the circle-LBP based local texture pattern in the nearby regions centering around the plaque area. To further refine the fuzed multiple-scale features, we devise an attention module based on CNN, which could focalize the regions of interest in plaque more easily, especially for many challenging cases. Extensive experiments and comprehensive evaluations confirm that, for a small-scale training dataset, our method could outperform the state-of-the-art methods. Meanwhile, the user studies verify the claim that our method is more accurate than conventional dental practice conducted by experienced dentists.

NeurIPS Conference 2021 Conference Paper

Knowledge-inspired 3D Scene Graph Prediction in Point Cloud

  • Shoulong Zhang
  • Shuai Li
  • Aimin Hao
  • Hong Qin

Prior knowledge integration helps identify semantic entities and their relationships in a graphical representation, however, its meaningful abstraction and intervention remain elusive. This paper advocates a knowledge-inspired 3D scene graph prediction method solely based on point clouds. At the mathematical modeling level, we formulate the task as two sub-problems: knowledge learning and scene graph prediction with learned prior knowledge. Unlike conventional methods that learn knowledge embedding and regular patterns from encoded visual information, we propose to suppress the misunderstandings caused by appearance similarities and other perceptual confusion. At the network design level, we devise a graph auto-encoder to automatically extract class-dependent representations and topological patterns from the one-hot class labels and their intrinsic graphical structures, so that the prior knowledge can avoid perceptual errors and noises. We further devise a scene graph prediction model to predict credible relationship triplets by incorporating the related prototype knowledge with perceptual information. Comprehensive experiments confirm that, our method can successfully learn representative knowledge embedding, and the obtained prior knowledge can effectively enhance the accuracy of relationship predictions. Our thorough evaluations indicate the new method can achieve the state-of-the-art performance compared with other scene graph prediction methods.

AAAI Conference 2021 Conference Paper

Point Cloud Semantic Scene Completion from RGB-D Images

  • Shoulong Zhang
  • Shuai Li
  • Aimin Hao
  • Hong Qin

In this paper, we devise a novel semantic completion network, called point cloud semantic scene completion network (PCSSC-Net), for indoor scenes solely based on point clouds. Existing point cloud completion networks still suffer from their inability of fully recovering complex structures and contents from global geometric descriptions neglecting semantic hints. To extract and infer comprehensive information from partial input, we design a patch-based contextual encoder to hierarchically learn point-level, patch-level, and scene-level geometric and contextual semantic information with a divideand-conquer strategy. Consider that the scene semantics afford a high-level clue of constituting geometry for an indoor scene environment, we articulate a semantics-guided completion decoder where semantics could help cluster isolated points in the latent space and infer complicated scene geometry. Given the fact that real-world scans tend to be incomplete as ground truth, we choose to synthesize scene dataset with RGB-D images and annotate complete point clouds as ground truth for the supervised training purpose. Extensive experiments validate that our new method achieves the stateof-the-art performance, in contrast with the current methods applied to our dataset.

JBHI Journal 2019 Journal Article

Multitask Cascade Convolution Neural Networks for Automatic Thyroid Nodule Detection and Recognition

  • Wenfeng Song
  • Shuai Li
  • Ji Liu
  • Hong Qin
  • Bo Zhang
  • Shuyang Zhang
  • Aimin Hao

Thyroid ultrasonography is a widely used clinical technique for nodule diagnosis in thyroid regions. However, it remains difficult to detect and recognize the nodules due to low contrast, high noise, and diverse appearance of nodules. In today's clinical practice, senior doctors could pinpoint nodules by analyzing global context features, local geometry structure, and intensity changes, which would require rich clinical experience accumulated from hundreds and thousands of nodule case studies. To alleviate doctors’ tremendous labor in the diagnosis procedure, we advocate a machine learning approach to the detection and recognition tasks in this paper. In particular, we develop a multitask cascade convolution neural network (MC-CNN) framework to exploit the context information of thyroid nodules. It may be noted that our framework is built upon a large number of clinically confirmed thyroid ultrasound images with accurate and detailed ground truth labels. Other key advantages of our framework result from a multitask cascade architecture, two stages of carefully designed deep convolution networks in order to detect and recognize thyroid nodules in a pyramidal fashion, and capturing various intrinsic features in a global-to-local way. Within our framework, the potential regions of interest after initial detection are further fed to the spatial pyramid augmented CNNs to embed multiscale discriminative information for fine-grained thyroid recognition. Experimental results on 4309 clinical ultrasound images have indicated that our MC-CNN is accurate and effective for both thyroid nodules detection and recognition. For the correct diagnosis rate of malignant and benign thyroid nodules, its mean Average Precision (mAP) performance can achieve up to $\text{98. 2}\%$ accuracy, which outperforms the common CNNs by $\text{5}\%$ on average. In addition, we conduct rigorous user studies to confirm that our MC-CNN outperforms experienced doctors, yet only consuming roughly $\text{2}\%$ ( $1/48$ ) of doctors’ examination time on average. Therefore, the accuracy and efficiency of our new method exhibit its great potential in clinical applications.

JBHI Journal 2016 Journal Article

Robust Optimization-Based Coronary Artery Labeling From X-Ray Angiograms

  • Xinglong Liu
  • Fei Hou
  • Hong Qin
  • Aimin Hao

In this paper, we present an efficient robust labeling method for coronary arteries from X-ray angiograms based on energy optimization. The fundamental goal of this research is to facilitate the analysis and diagnosis of interventional surgery in the most efficient way, and such effort could also improve the performance during doctor training, and surgery simulation and planning. Compared to the prior state-of-the-art, our method is much more robust to resist noises and is tolerant to even incomplete data because of the “ built-in ” nature of global optimization. We start with a fully parallelized algorithm based on Hessian matrix to extract the tubular structure from the X-ray angiograms as vessel candidates. Then, instead of using the candidates directly, we use the grow cut (Vezhnevets and V. Konouchine, Growcut: Interactive multi-label N-D image segmentation by cellular automata, in Proc. of Graphicon, 2005, pp. 150–156.) method, which is similar to graph cut (Boykov et al. , Fast approximate energy minimization via graph cuts, IEEE Trans. Pattern Anal. Mach. Intell. , vol. 23, no. 11, pp. 1222–1239, Nov. 2001.)but with better performance to extract the precise vessel structure from the images. Next, we use the fast marching method with second derivatives and cross neighbors to extract the accurate skeleton segments. After that, we propose an efficient method based on iterative closest point (Z. Zhang, Iterative point matching for registration of free-form curves and surfaces, Int J. Comput. Vis. , vol. 13, no. 2, pp. 119–152, 1994.) to organize the skeleton segments by treating the continuity and similarity as extra constraints. Finally, we formulate the vessel labeling problem as an energy optimization problem and solve it using belief propagation. We also demonstrate several typical applications including flow velocity estimation, heart beat estimation, and vessel diameter estimation to show its practical uses in clinical diagnosis and treatment. Our experiments exhibit the correctness and robustness, as well as the high performance of our algorithm. We envision that our system would be of high utility for diagnosis and therapy to treat vessel-related diseases in a clinical setting in the near future.

v2026.09.13