Arrow Research search

Author name cluster

Lei Hu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
2 author rows

Possible papers

11

NeurIPS Conference 2025 Conference Paper

HMVLM:Human Motion-Vision-Language Model via MoE LoRA

  • Lei Hu
  • Yongjing Ye
  • Shihong Xia

The expansion of instruction-tuning data has enabled foundation language models to exhibit improved instruction adherence and superior performance across diverse downstream tasks. Semantically-rich 3D human motion is being progressively integrated with these foundation models to enhance multimodal understanding and cross-modal generation capabilities. However, the modality gap between human motion and text raises unresolved concerns about catastrophic forgetting during this integration. In addition, developing autoregressive-compatible pose representations that preserve generalizability across heterogeneous downstream tasks remains a critical technical barrier. To address these issues, we propose the Human Motion-Vision-Language Model (HMVLM), a unified framework based on the Mixture of Expert Low-Rank Adaption(MoE LoRA) strategy. The framework leverages the gating network to dynamically allocate LoRA expert weights based on the input prompt, enabling synchronized fine-tuning of multiple tasks. To mitigate catastrophic forgetting during instruction-tuning, we introduce a novel zero expert that preserves the pre-trained parameters for general linguistic tasks. For pose representation, we implement body-part-specific tokenization by partitioning the human body into different joint groups, enhancing the spatial resolution of the representation. Experiments show that our method effectively alleviates knowledge forgetting during instruction-tuning and achieves remarkable performance across diverse human motion downstream tasks.

AAAI Conference 2025 Conference Paper

Motion Artifact Removal in Pixel-Frequency Domain via Alternate Masks and Diffusion Model

  • Jiahua Xu
  • Dawei Zhou
  • Lei Hu
  • Jianfeng Guo
  • Feng Yang
  • Zaiyi Liu
  • Nannan Wang
  • Xinbo Gao

Motion artifacts present in magnetic resonance imaging (MRI) can seriously interfere with clinical diagnosis. Removing motion artifacts is a straightforward solution and has been extensively studied. However, paired data are still heavily relied on in recent works and the perturbations in k-space (frequency domain) are not well considered, which limits their applications in the clinical field. To address these issues, we propose a novel unsupervised purification method which leverages pixel-frequency information of noisy MRI images to guide a pre-trained diffusion model to recover clean MRI images. Specifically, considering that motion artifacts are mainly concentrated in high-frequency components in k-space, we utilize the low-frequency components as the guide to ensure correct tissue textures. Additionally, given that high-frequency and pixel information are helpful for recovering shape and detail textures, we design alternate complementary masks to simultaneously destroy the artifact structure and exploit useful information. Quantitative experiments are performed on datasets from different tissues and show that our method achieves superior performance on several metrics. Qualitative evaluations with radiologists also show that our method provides better clinical feedback.

IJCAI Conference 2025 Conference Paper

ReplayCAD: Generative Diffusion Replay for Continual Anomaly Detection

  • Lei Hu
  • Zhiyong Gan
  • Ling Deng
  • Jinglin Liang
  • Lingyu Liang
  • Shuangping Huang
  • Tianshui Chen

Continual Anomaly Detection (CAD) enables anomaly detection models in learning new classes while preserving knowledge of historical classes. CAD faces two key challenges: catastrophic forgetting and segmentation of small anomalous regions. Existing CAD methods store image distributions or patch features to mitigate catastrophic forgetting, but they fail to preserve pixel-level detailed features for accurate segmentation. To overcome this limitation, we propose ReplayCAD, a novel diffusion-driven generative replay framework that replay high-quality historical data, thus effectively preserving pixel-level detailed features. Specifically, we compress historical data by searching for a class semantic embedding in the conditional space of the pre-trained diffusion model, which can guide the model to replay data with fine-grained pixel details, thus improving the segmentation performance. However, relying solely on semantic features results in limited spatial diversity. Hence, we further use spatial features to guide data compression, achieving precise control of sample space, thereby generating more diverse data. Our method achieves state-of-the-art performance in both classification and segmentation, with notable improvements in segmentation: 11. 5% on VisA and 8. 1% on MVTec. Our source code is available at https: //github. com/HULEI7/ReplayCAD.

JBHI Journal 2024 Journal Article

Protecting Prostate Cancer Classification From Rectal Artifacts via Targeted Adversarial Training

  • Lei Hu
  • Dawei Zhou
  • Jiahua Xu
  • Cheng Lu
  • Chu Han
  • Zhenwei Shi
  • Qikui Zhu
  • Xinbo Gao

Magnetic resonance imaging (MRI)-based deep neural networks (DNN) have been widely developed to perform prostate cancer (PCa) classification. However, in real-world clinical situations, prostate MRIs can be easily impacted by rectal artifacts, which have been found to lead to incorrect PCa classification. Existing DNN-based methods typically do not consider the interference of rectal artifacts on PCa classification, and do not design specific strategy to address this problem. In this study, we proposed a novel Targeted adversarial training with Proprietary Adversarial Samples (TPAS) strategy to defend the PCa classification model against the influence of rectal artifacts. Specifically, based on clinical prior knowledge, we generated proprietary adversarial samples with rectal artifact-pattern adversarial noise, which can severely mislead PCa classification models optimized by the ordinary training strategy. We then jointly exploited the generated proprietary adversarial samples and original samples to train the models. To demonstrate the effectiveness of our strategy, we conducted analytical experiments on multiple PCa classification models. Compared with ordinary training strategy, TPAS can effectively improve the single- and multi-parametric PCa classification at patient, slice and lesion level, and bring substantial gains to recent advanced models. In conclusion, TPAS strategy can be identified as a valuable way to mitigate the influence of rectal artifacts on deep learning models for PCa classification.

EAAI Journal 2023 Journal Article

MSRA-G: Combination of multi-scale residual attention network and generative adversarial networks for hyperspectral image classification

  • Jinling Zhao
  • Lei Hu
  • Linsheng Huang
  • Chuanjian Wang
  • Dong Liang

Deep learning-based technology has been introduced to increase the classification accuracy of hyperspectral imagery (HSI). Nevertheless, it is still a challenging issue to derive a satisfying classification accuracy from limited training samples. A novel method (MSRA-G) that combines multi-scale residual attention (MSRA) with Generative Adversarial Networks (GANs) was proposed. In view of the low classification accuracy with limited training samples, the is first used to generate more separable synthetic samples. A network is then proposed to extract multi-scale context information for improving HSI classification. The proposed method constructs two multi-scale feature extraction modules to identify high-level spatial–spectral features based on the 3D–2D hybrid network. In addition, the residual connection mode and the attention mechanism are combined to establish the channel and spatial residual attention modules. Different weights are assigned to different features in the channel dimension and spatial dimension, and the features are selectively learned. Furthermore, to verify the performance of MSRA-G, experiments were carried out on three publicly available HSI datasets of Indian Pines, University of Pavia and Salinas Valley. The experimental results show that our proposed MSRA-G is superior to several popular classification models. It can still achieve satisfactory classification accuracies, even in the case of insufficient training samples.

EAAI Journal 2023 Journal Article

Two-rank multi-attribute group decision-making with linguistic distribution assessments: An optimization-based integrated approach

  • Shitao Zhang
  • Lei Hu
  • Zhenzhen Ma
  • Xiaodi Liu

In real life, two-rank multi-attribute decision-making (MADM), in which all alternatives are divided into two preference-ordered categories, is common. In this paper, we investigate two-rank multi-attribute group decision-making (MAGDM) with linguistic distribution assessments (LDAs). A challenge when tackling such two-rank problems is establishing an LDAs-based two-rank model that maintains a balance between classification accuracy and computational complexity. When neither the threshold nor the number of alternatives within each category is specified in advance, determining individual two-rank results and subsequently aggregating the two-rank results for each decision-maker to resolve conflicts within the group is another challenge. Given these, we aim to propose a novel approach for two-rank MAGDM with LDAs. The main innovations and contributions of this paper are as follows. (a) From the perspective of linguistic scale function (LSF)-based cumulative expectations, we present a new LDAs-based distance measure that exhibits several desirable properties. A new score function for comparing LDAs is subsequently proposed using the new distance and the idea of the Technique for Order Preference by Similarity to Ideal Solution (TOPSIS). (b) Using the minimum intensity of reversed rankings combined with the misclassification ratio of alternatives as an integrated objective function, we construct two 0-1 integer programming models incorporating constraints associated with the centers and priorities of categories to determine the optimal individual and group two-rank results of alternatives, respectively. (c) We apply our method to two-rank MAGDM associated with short video placement platforms. Comparing the proposed approach with other two-rank MAGDM approaches further demonstrates its effectiveness and rationality.

TCS Journal 2022 Journal Article

Revisiting orthogonal lattice attacks on approximate common divisor problems

  • Jun Xu
  • Santanu Sarkar
  • Lei Hu

In this paper, we revisit three existing types of orthogonal lattice (OL) attacks and propose optimized cases to solve approximate common divisor (ACD) problems. In order to reduce both space and time costs, we also make an improved lattice using the rounding technique. Further, we present asymptotic formulas of the time complexities on our optimizations as well as three known OL attacks. Besides, we give specific conditions that the optimized OL attacks can work and show how the attack ability depends on the blocksize β in the BKZ-β algorithm.

IJCAI Conference 2021 Conference Paper

Sequential 3D Human Pose Estimation Using Adaptive Point Cloud Sampling Strategy

  • Zihao Zhang
  • Lei Hu
  • Xiaoming Deng
  • Shihong Xia

3D human pose estimation is a fundamental problem in artificial intelligence, and it has wide applications in AR/VR, HCI and robotics. However, human pose estimation from point clouds still suffers from noisy points and estimated jittery artifacts because of handcrafted-based point cloud sampling and single-frame-based estimation strategies. In this paper, we present a new perspective on the 3D human pose estimation method from point cloud sequences. To sample effective point clouds from input, we design a differentiable point cloud sampling method built on density-guided attention mechanism. To avoid the jitter caused by previous 3D human pose estimation problems, we adopt temporal information to obtain more stable results. Experiments on the ITOP dataset and the NTU-RGBD dataset demonstrate that all of our contributed components are effective, and our method can achieve state-of-the-art performance.

IROS Conference 2009 Conference Paper

A fluoroscopic-based navigation system for ACL reconstruction assisted by robot

  • Yan Hu
  • Lei Hu
  • Tianmiao Wang
  • Jun Wei 0005
  • Sun Lei
  • Wenyong Liu
  • Li Wen

Entry position of the graft is very important in anterior cruciate ligament (ACL) reconstruction. However the determination of entry position is very difficult to the surgeon. In this paper, a navigation and evaluation system assisted by the 6-DOFS robot is implemented for the simulation evaluation and planning insertion points based on quadrant method for the femur and Stäublis method for the tibia on the lateral X-ray image of knee joint. Meanwhile, the implementation of the key technologies such as image correction, image registration, C-arm calibration, video tracking, bone surface reconstruction, image fusion, 6-DOFS robot, and virtual simulation are introduced. Finally, Experiments about the tunnel planning method and real time tracking of surgical apparatus are implemented on 8 bone of plastic models (Sawbone, Swiss) and 10 bones of the goat. In the experiment, the tibia rotates around the femur under the surgeon's implementation to evaluate the planning result with the virtual simulation and evaluation module. The positioning error is 1. 59mm from analysis on 30 space targets. The virtual reconstruction ACL is satisfied with two important criteria of the best isometry and collision detection between graft and intercondylar surface of femur. The results are well accepted in operations. In order to satisfying with the request of exact operation in ACL reconstruction, we have developed 6-DOFS passive robot to assist the surgeons entry positioning and drilling of implant tunnels, implementing exact operation in knee joint.

IROS Conference 2006 Conference Paper

An Internet Robot Assistant Tele-neurosurgery System Case

  • Tianmiao Wang
  • Jun Wei 0005
  • Da Liu 0002
  • Lei Hu
  • Wenyong Liu

Robot assistant surgical system is currently a research focus. This paper introduces a tele-neurosurgery system joined developed Beihang university and General Navy Hospital of PLA, China. The content describes the system structure, important procedures, modules, safety and convenience consideration. The system has successfully performed three teleneurosurgery operation between Beijing and Yan'an, Shanxi province via public Internet. The experiment results proved the system is feasible, safe and low-cost

IROS Conference 2004 Conference Paper

BPOR: a fluoroscopy-based robot navigating system for distal locking of intramedullary nails

  • Tianmiao Wang
  • Wenyong Liu
  • Lei Hu

According to the biplanar imaging principle, this paper proposes a novel fluoroscopy-guided robot navigating system for distal locking of intramedullary nails which only needs two X-ray projection images from lateral and AP views of the nail respectively. Error distributions are analyzed based on the mathematical model and the structural model. The system performance is also tested and evaluated through preliminary experiments with plastic bones and cadaver limbs. It is shown that this system has enough positioning accuracy and stability with a compact robot structure, and it can dramatically decrease the operating time. The modular design makes it applicable for most kinds of intramedullary nails existed nowadays.

v2026.09.13