Arrow Research search

Author name cluster

Junhao Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

10 papers
2 author rows

Possible papers

10

JBHI Journal 2026 Journal Article

DiffSpkSync: A Muscle Synergy-Guided Spiking Diffusion Model for EMG Signal Generation to Improve Gesture Recognition Performance

  • Kejia Su
  • Bo Wan
  • Jiayang Huang
  • Zhi-Qiang Zhang
  • Junhao Zhang
  • Pengfei Yang
  • Quan Wang

High-density surface electromyography (HD-sEMG) based hand gesture recognition (HGR) has shown great promise for intuitive human-machine interaction. However, the performance of HGR model is often hindered by a scarcity of available training data, especially in the fields of gesture recognition, rehabilitation, and medicine. To address these issues, we propose DiffSpkSync, a novel generative framework that integrates (1) muscle synergy-guided diffusion modeling for physiologically plausible signal reconstruction, (2) spiking neuron-based sparsification to reduce energy cost, and (3) a time-series mixup strategy to preserve local dynamics during augmentation. Experiments on a public Hyser dataset and a self-collected XDHDEMG dataset demonstrate that training gesture classifiers with data augmented by DiffSpkSync consistently improves classification accuracy in both intrasession and intersession scenarios. Comparative results further demonstrate superior performance over representative generative baselines, including VAE, DCGAN, DANN-CRC, and PatchEMG. Furthermore, real-time validation demonstrates that the proposed method achieves an average of 130. 22 ms end-to-end latency and an average of 95. 87% accuracy predictions, supporting their applicability in real-world applications.

AAAI Conference 2026 System Paper

Placing Any Object at Any 3D Position

  • Junhao Zhang
  • Ming Kong
  • Zhanbin Hu
  • Hao Qin
  • Zhijie Xu
  • Xiaojun Zhu
  • Qiang Zhu

In this work, we propose a diffusion-based method for 3D-aware image composition. Previous approaches have focused on 2D-view image composition, which limits their handling of complex 3D spatial relationships. Consequently, they are not well-suited for applications requiring precise 3D object control and iterative refinement, including interior design visualization, visual effects prototyping, and virtual reality scene construction. In contrast, our method extracts 3D bounding boxes for all objects in the scene image. Users can then specify a new 3D bounding box based on existing spatial context and provide an image of the target object. Leveraging a fine-tuned diffusion model, our approach enables high-fidelity image composition while preserving the underlying 3D structure of the scene.

YNIMG Journal 2025 Journal Article

ACTION: Augmentation and computation toolbox for brain network analysis with functional MRI

  • Yuqi Fang
  • Junhao Zhang
  • Linmin Wang
  • Qianqian Wang
  • Mingxia Liu

Functional magnetic resonance imaging (fMRI) has been increasingly employed to investigate functional brain activity. Many fMRI-related software/toolboxes have been developed, providing specialized algorithms for fMRI analysis. However, existing toolboxes seldom consider fMRI data augmentation, which is quite useful, especially in studies with limited or imbalanced data. Moreover, current studies usually focus on analyzing fMRI using conventional machine learning models that rely on human-engineered fMRI features, without investigating deep learning models that can automatically learn data-driven fMRI representations. In this work, we develop an open-source toolbox, called Augmentation and Computation Toolbox for braIn netwOrk aNalysis (ACTION), offering comprehensive functions to streamline fMRI analysis. The ACTION is a Python-based and cross-platform toolbox with graphical user-friendly interfaces. It enables automatic fMRI augmentation, covering blood-oxygen-level-dependent (BOLD) signal augmentation and brain network augmentation. Many popular methods for brain network construction and network feature extraction are included. In particular, it supports constructing deep learning models, which leverage large-scale auxiliary unlabeled data (3,800+ resting-state fMRI scans) for model pretraining to enhance model performance for downstream tasks. To facilitate multi-site fMRI studies, it is also equipped with several popular federated learning strategies. Furthermore, it enables users to design and test custom algorithms through scripting, greatly improving its utility and extensibility. We demonstrate the effectiveness and user-friendliness of ACTION on real fMRI data and present the experimental results. The software, along with its source code and manual, can be accessed online.

JBHI Journal 2025 Journal Article

Adaptive High-Order Fusion Learning for Brain Disorder Detection

  • Hengsheng Tang
  • Junji Jiang
  • Junhao Zhang
  • Yining Zhang
  • Shufeng Zhou
  • Yueying Zhou
  • Lishan Qiao

The functional brain network (FBN) serves as an important tool for investigating neurological and mental disorders. Unlike many traditional networks whose structures are often known in advance, FBNs need to be estimated from neuroimaging or electrophysiological data, and their quality generally determines the performance of downstream tasks, particularly in disorder detection. Recent studies have shown that high-order FBNs tend to achieve better discriminative performance, while some other work indicates that increasing the order of FBN does not necessarily bring additional gains in discriminative performance and may even lead to a rapid decline in discriminability. To fully leverage information across different-order FBNs and identify optimal order for downstream tasks, we design an adaptive high-order FBN fusion learning framework (AHFL) with attention mechanism for brain disorder detection. Specifically, we first construct a series of FBNs with continuously increasing orders and propose a data-driven approach to evaluate each order's contribution to the classification performance. The self-attention mechanism is employed to capture contextual dependencies during the sequential generation of multi-order FBNs, thereby offering a natural fusion approach. Experimental evidence shows that the proposed method achieves superior performance compared to the baseline. In particular, we find that third-order FBN is achieved the highest weights, playing a crucial role for the detection of autism spectrum disorder (ASD), whereas second-order FBN is the most effective for identifying patients with major depressive disorder (MDD). Our findings advance the identification of discriminative high-order FBNs and establish a generalizable diagnostic framework for brain disorders.

NeurIPS Conference 2024 Conference Paper

EvolveDirector: Approaching Advanced Text-to-Image Generation with Large Vision-Language Models

  • Rui Zhao
  • Hangjie Yuan
  • Yujie Wei
  • Shiwei Zhang
  • Yuchao Gu
  • Lingmin Ran
  • Xiang Wang
  • Zhangjie Wu

Recent advancements in generation models have showcased remarkable capabilities in generating fantastic content. However, most of them are trained on proprietary high-quality data, and some models withhold their parameters and only provide accessible application programming interfaces (APIs), limiting their benefits for downstream tasks. To explore the feasibility of training a text-to-image generation model comparable to advanced models using publicly available resources, we introduce EvolveDirector. This framework interacts with advanced models through their public APIs to obtain text-image data pairs to train a base model. Our experiments with extensive data indicate that the model trained on generated data of the advanced model can approximate its generation capability. However, it requires large-scale samples of 10 million or more. This incurs significant expenses in time, computational resources, and especially the costs associated with calling fee-based APIs. To address this problem, we leverage pre-trained large vision-language models (VLMs) to guide the evolution of the base model. VLM continuously evaluates the base model during training and dynamically updates and refines the training dataset by the discrimination, expansion, deletion, and mutation operations. Experimental results show that this paradigm significantly reduces the required data volume. Furthermore, when approaching multiple advanced models, EvolveDirector can select the best samples generated by them to learn powerful and balanced abilities. The final trained model Edgen is demonstrated to outperform these advanced models. The code and model weights are available at https: //github. com/showlab/EvolveDirector.

AAAI Conference 2021 Conference Paper

Investigate Indistinguishable Points in Semantic Segmentation of 3D Point Cloud

  • Mingye Xu
  • Zhipeng Zhou
  • Junhao Zhang
  • Yu Qiao

This paper investigates the indistinguishable points (difficult to predict label) in semantic segmentation for large-scale 3D point clouds. The indistinguishable points consist of those located in complex boundary, points with similar local textures but different categories, and points in isolate small hard areas, which largely harm the performance of 3D semantic segmentation. To address this challenge, we propose a novel Indistinguishable Area Focalization Network (IAF-Net), which select indistinguishable points adaptively by utilizing the hierarchical semantic features and enhance fine-grained features for points especially those indistinguishable points. We also introduce multi-stage loss to improve the feature representation in a progressive way. Moreover, in order to analyze the segmentation performances of indistinguishable areas, we propose a new evaluation metric called Indistinguishable Points Based Metric (IPBM). Our IAF-Net achieves the comparable results with state-of-the-art performance on several popular 3D point cloud datasets e. g. S3DIS and ScanNet, and clearly outperform other methods on IPBM. Our code will be available at https: //github. com/MingyeXu/IAF-Net

AAAI Conference 2021 Conference Paper

Learning Geometry-Disentangled Representation for Complementary Understanding of 3D Object Point Cloud

  • Mutian Xu
  • Junhao Zhang
  • Zhipeng Zhou
  • Mingye Xu
  • Xiaojuan Qi
  • Yu Qiao

In 2D image processing, some attempts decompose images into high and low frequency components for describing edge and smooth parts respectively. Similarly, the contour and flat area of 3D objects, such as the boundary and seat area of a chair, describe different but also complementary geometries. However, such investigation is lost in previous deep networks that understand point clouds by directly treating all points or local patches equally. To solve this problem, we propose Geometry-Disentangled Attention Network (GDANet). GDANet introduces Geometry-Disentangle Module to dynamically disentangle point clouds into the contour and flat part of 3D objects, respectively denoted by sharp and gentle variation components. Then GDANet exploits Sharp-Gentle Complementary Attention Module that regards the features from sharp and gentle variation components as two holistic representations, and pays different attentions to them while fusing them respectively with original point cloud features. In this way, our method captures and refines the holistic and complementary 3D geometric semantics from two distinct disentangled components to supplement the local information. Extensive experiments on 3D object classification and segmentation benchmarks demonstrate that GDANet achieves the state-of-the-arts with fewer parameters. Code is released on https: //github. com/mutianxu/GDANet.

AAAI Conference 2021 Conference Paper

PC-HMR: Pose Calibration for 3D Human Mesh Recovery from 2D Images/Videos

  • Tianyu Luan
  • Yali Wang
  • Junhao Zhang
  • Zhe Wang
  • Zhipeng Zhou
  • Yu Qiao

The end-to-end Human Mesh Recovery (HMR) approach (Kanazawa et al. 2018) has been successfully used for 3D body reconstruction. However, most HMR-based frameworks reconstruct human body by directly learning mesh parameters from images or videos, while lacking explicit guidance of 3D human pose in visual data. As a result, the generated mesh often exhibits incorrect pose for complex activities. To tackle this problem, we propose to exploit 3D pose to calibrate human mesh. Specifically, we develop two novel Pose Calibration frameworks, i. e. , Serial PC-HMR and Parallel PC-HMR. By coupling advanced 3D pose estimators and HMR in a serial or parallel manner, these two frameworks can effectively correct human mesh with guidance of a concise pose calibration module. Furthermore, since the calibration module is designed via non-rigid pose transformation, our PC- HMR frameworks can flexibly tackle bone length variations to alleviate misplacement in the calibrated mesh. Finally, our frameworks are based on generic and complementary integration of data-driven learning and geometrical modeling. Via plug-and-play modules, they can be efficiently adapted for both image/video-based human mesh recovery. Additionally, they have no requirement of extra 3D pose annotations in the testing phase, which releases inference difficulties in practice. We perform extensive experiments on the popular benchmarks, i. e. , Human3. 6M, 3DPW and SURREAL, where our PC-HMR frameworks achieve the SOTA results.

IROS Conference 2021 Conference Paper

Sensor Fusion-based Anthropomorphic Control of Under-Actuated Bionic Hand in Dynamic Environment

  • Hang Su 0001
  • Junhao Zhang
  • Junling Fu
  • Salih Ertug Ovur
  • Wen Qi 0005
  • Guoxin Li 0001
  • Yingbai Hu
  • Zhijun Li 0001

Under-actuated bionic hands have achieved tremendous popularity in many fields because of their advantages of lightweight, budget-friendly, satisfactory flexibility, and adaptability. Except for the bionic mechanical design, various anthropomorphic control strategies have been proposed and investigated in the last decades. However, due to its under-actuated characteristic, there are still many challenges for anthropomorphic control of all the degrees of freedom (DOFs) using less input. It is challenging to map the human hand kinematic synergies on robotic hands, particularly for a dynamic environment. Therefore, it is worth studying how to control the under-actuated bionic hand effectively in a dynamic environment. In this paper, an anthropomorphic control method is proposed using sensor fusion of hand kinematic inputs to control the under-actuated bionic hand. In order to map the kinematics of human fingers to the bionic hand, a novel finger bending angle is defined to represent the posture of human fingers. Multiple Leap Motion Controllers (LMC) are fused to estimate the stable and accurate finger bending angles to avoid the occlusion problem. Finally, experiments with real-time control of the under-actuated bionic hand are implemented to demonstrate the proposed approach’s effectiveness.

ICRA Conference 2020 Conference Paper

Grasp for Stacking via Deep Reinforcement Learning

  • Junhao Zhang
  • Wei Zhang 0021
  • Ran Song 0001
  • Lin Ma 0002
  • Yibin Li 0001

Integrated robotic arm system should contain both grasp and place actions. However, most grasping methods focus more on how to grasp objects, while ignoring the placement of the grasped objects, which limits their applications in various industrial environments. In this research, we propose a model-free deep Q-learning method to learn the grasping-stacking strategy end-to-end from scratch. Our method maps the images to the actions of the robotic arm through two deep networks: the grasping network (GNet) using the observation of the desk and the pile to infer the gripper’s position and orientation for grasping, and the stacking network (SNet) using the observation of the platform to infer the optimal location when placing the grasped object. To make a long-range planning, the two observations are integrated in the grasping for stacking network (GSN). We evaluate the proposed GSN on a grasping-stacking task in both simulated and real-world scenarios.

v2026.09.13