Arrow Research search

Author name cluster

Weiqing Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
2 author rows

Possible papers

6

EAAI Journal 2026 Journal Article

A general framework for interactive semantic segmentation refinement of point clouds

  • Peng Zhang
  • Ting Wu
  • Jinsheng Sun
  • Weiqing Li
  • Zhiyong Su

Although deep learning-based point cloud semantic segmentation has been extensively studied in the past decade, it is still challenging to produce high quality masks that meet high-precision downstream applications. This challenge stems from the distribution mismatch between training and testing data. The pre-trained segmentation networks optimized on the training dataset may perform sub-optimally on individual unseen testing data, resulting in performance drop. To this end, we propose a general interactive framework to enhance off-the-shelf networks. This framework integrates with off-the-shelf semantic segmentation networks in a fully test-time manner, allowing users to refine mis-segmented regions with a few corrective clicks. Specifically, we formulate a correction energy that treats user clicks as sparse training examples for test-time optimization. To mitigate catastrophic overfitting caused by sparse supervision, we formulate a stabilization energy that selectively minimizes the entropy of global points. Both the correction and stabilization energies constitute the test-time loss, promoting effective refinement of mis-segmented regions while maintaining the stability of others. Furthermore, a warm-up pre-process and an interaction simulation scheme are proposed for performance improvement and reproducible evaluation, respectively. We evaluate our framework on indoor and outdoor datasets with off-the-shelf networks, showing promising results in semantic segmentation refinement. The source code is available at https: //github. com/Pengz98/ISSR.

ICRA Conference 2025 Conference Paper

Gradient-Based Adversarial Attacks on Deep LiDAR Odometry

  • Zhenbo Song
  • Xuanzhu Chen
  • Zhenyuan Zhang 0001
  • Kaihao Zhang
  • Jianfeng Lu 0003
  • Weiqing Li

Adversarial attacks have been recently investigated in LiDAR perception problems for autonomous driving, where a small perturbation of source inputs can result in incorrect predictions. However, most previous studies focus on attacks on single-frame perception modules, lacking explorations of attacks on consecutive-frame tasks, i. e. the LiDAR odometry. In this paper, we propose a gradient optimization-based adversarial attack towards deep LiDAR odometry networks. To generate point clouds consistent with real-world scenarios, we constrain adversarial points within the range of a small object, e. g. a traffic cone, and render new points to simulate real LiDAR measurements. By incorporating such adversarial points in consecutive frames, we demonstrate a significant decrease in pose estimation accuracy of current popular LiDAR odometry networks. In addition, we also evaluate traditional geometric odometry approaches and report their robustness against adversarial points. Extensive experiments on the KITTI and Waymo datasets illustrate the effectiveness of the proposed attack method and the vulnerability of deep LiDAR odometry networks against adversarial points.

AAAI Conference 2024 Conference Paper

Enhanced Fine-Grained Motion Diffusion for Text-Driven Human Motion Synthesis

  • Dong Wei
  • Xiaoning Sun
  • Huaijiang Sun
  • Shengxiang Hu
  • Bin Li
  • Weiqing Li
  • Jianfeng Lu

The emergence of text-driven motion synthesis technique provides animators with great potential to create efficiently. However, in most cases, textual expressions only contain general and qualitative motion descriptions, while lack fine depiction and sufficient intensity, leading to the synthesized motions that either (a) semantically compliant but uncontrollable over specific pose details, or (b) even deviates from the provided descriptions, bringing animators with undesired cases. In this paper, we propose DiffKFC, a conditional diffusion model for text-driven motion synthesis with KeyFrames Collaborated, enabling realistic generation with collaborative and efficient dual-level control: coarse guidance at semantic level, with only few keyframes for direct and fine-grained depiction down to body posture level. Unlike existing inference-editing diffusion models that incorporate conditions without training, our conditional diffusion model is explicitly trained and can fully exploit correlations among texts, keyframes and the diffused target frames. To preserve the control capability of discrete and sparse keyframes, we customize dilated mask attention modules where only partial valid tokens participate in local-to-global attention, indicated by the dilated keyframe mask. Additionally, we develop a simple yet effective smoothness prior, which steers the generated frames towards seamless keyframe transitions at inference. Extensive experiments show that our model not only achieves state-of-the-art performance in terms of semantic fidelity, but more importantly, is able to satisfy animator requirements through fine-grained guidance without tedious labor.

ICLR Conference 2024 Conference Paper

NeRM: Learning Neural Representations for High-Framerate Human Motion Synthesis

  • Dong Wei 0007
  • Huaijiang Sun
  • Bin Li 0084
  • Xiaoning Sun
  • Shengxiang Hu 0001
  • Weiqing Li
  • Jianfeng Lu 0003

Generating realistic human motions with high framerate is an underexplored task, due to the varied framerates of training data, huge memory burden brought by high framerates and slow sampling speed of generative models. Recent advances make a compromise for training by downsampling high-framerate details away and discarding low-framerate samples, which suffer from severe information loss and restricted-framerate generation. In this paper, we found that the recent emerging paradigm of Implicit Neural Representations (INRs) that encode a signal into a continuous function can effectively tackle this challenging problem. To this end, we introduce NeRM, a generative model capable of taking advantage of varied-size data and capturing variational distribution of motions for high-framerate motion synthesis. By optimizing latent representation and a auto-decoder conditioned on temporal coordinates, NeRM learns continuous motion fields of sampled motion clips that ingeniously avoid explicit modeling of raw varied-size motions. This expressive latent representation is then used to learn a diffusion model that enables both unconditional and conditional generation of human motions. We demonstrate that our approach achieves competitive results with state-of-the-art methods, and can generate arbitrary framerate motions. Additionally, we show that NeRM is not only memory-friendly, but also highly efficient even when generating high-framerate motions.

AAAI Conference 2023 Conference Paper

Human Joint Kinematics Diffusion-Refinement for Stochastic Motion Prediction

  • Dong Wei
  • Huaijiang Sun
  • Bin Li
  • Jianfeng Lu
  • Weiqing Li
  • Xiaoning Sun
  • Shengxiang Hu

Stochastic human motion prediction aims to forecast multiple plausible future motions given a single pose sequence from the past. Most previous works focus on designing elaborate losses to improve the accuracy, while the diversity is typically characterized by randomly sampling a set of latent variables from the latent prior, which is then decoded into possible motions. This joint training of sampling and decoding, however, suffers from posterior collapse as the learned latent variables tend to be ignored by a strong decoder, leading to limited diversity. Alternatively, inspired by the diffusion process in nonequilibrium thermodynamics, we propose MotionDiff, a diffusion probabilistic model to treat the kinematics of human joints as heated particles, which will diffuse from original states to a noise distribution. This process not only offers a natural way to obtain the "whitened'' latents without any trainable parameters, but also introduces a new noise in each diffusion step, both of which facilitate more diverse motions. Human motion prediction is then regarded as the reverse diffusion process that converts the noise distribution into realistic future motions conditioned on the observed sequence. Specifically, MotionDiff consists of two parts: a spatial-temporal transformer-based diffusion network to generate diverse yet plausible motions, and a flexible refinement network to further enable geometric losses and align with the ground truth. Experimental results on two datasets demonstrate that our model yields the competitive performance in terms of both diversity and accuracy.

AAAI Conference 2023 Conference Paper

Meta-Auxiliary Learning for Adaptive Human Pose Prediction

  • Qiongjie Cui
  • Huaijiang Sun
  • Jianfeng Lu
  • Bin Li
  • Weiqing Li

Predicting high-fidelity future human poses, from a historically observed sequence, is crucial for intelligent robots to interact with humans. Deep end-to-end learning approaches, which typically train a generic pre-trained model on external datasets and then directly apply it to all test samples, emerge as the dominant solution to solve this issue. Despite encouraging progress, they remain non-optimal, as the unique properties (e.g., motion style, rhythm) of a specific sequence cannot be adapted. More generally, once encountering out-of-distributions, the predicted poses tend to be unreliable. Motivated by this observation, we propose a novel test-time adaptation framework that leverages two self-supervised auxiliary tasks to help the primary forecasting network adapt to the test sequence. In the testing phase, our model can adjust the model parameters by several gradient updates to improve the generation quality. However, due to catastrophic forgetting, both auxiliary tasks typically have a low ability to automatically present the desired positive incentives for the final prediction performance. For this reason, we also propose a meta-auxiliary learning scheme for better adaptation. Extensive experiments show that the proposed approach achieves higher accuracy and more realistic visualization.

v2026.09.13