Arrow Research search

Author name cluster

Zhenyu Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

22 papers
2 author rows

Possible papers

22

JBHI Journal 2026 Journal Article

A Novel Binocular-Encoded SSVEP Framework for Efficient VR-Based Brain-Computer Interface

  • Haifeng Liu
  • Zhenyu Wang
  • Ruxue Li
  • Xi Zhao
  • Tianheng Xu
  • Ting Zhou
  • Honglin Hu

This paper presents a novel binocular-encoded SSVEP (beSSVEP) method, leveraging binocular vision in virtual reality (VR) to enhance brain-computer interface (BCI) applications. We introduce the Binocular Periodically Repeated Component Analysis (bPRCA) algorithm, designed to address the unique characteristics of binocular-encoded targets, which include combinations of monocular single-frequency SSVEP units or void units, with frequency units being reused multiple times in the encoded interface. To further optimize performance, we propose the Fusion Component Analysis (FusionCA) framework, which integrates bPRCA with Task-Related Component Analysis (TRCA), effectively utilizing both steady-state periodic components and cross-trial aperiodic components. Experimental results demonstrate that ensemble-FusionCA achieves the highest information transfer rate (ITR) with an average accuracy of 71. 39% and an ITR of 138. 50 bits/min at 0. 4 seconds, among the comparison with ensemble-bPRCA and ensemble-TRCA. Compared to traditional SSVEP approaches, beSSVEP significantly enhances frequency utilization, making VR-BCI systems more efficient and practical. This study highlights the application of physiological mechanisms of binocular vision to improve BCI systems, offering a new perspective for developing fast and scalable brain-computer interactions in VR environments.

EAAI Journal 2026 Journal Article

A novel color marker-based target tracking and motion intent detection method

  • Zhenyu Wang
  • Zenan Lu
  • Simin Tang
  • Jianmin Wang

To address the limitations of conventional target tracking methods—such as reliance on external devices, poor robustness, and low accuracy in motion intent detection under complex scenarios—this study proposes a novel color marker-based target tracking and motion intent detection method (CMTT). Specifically, the target object is first annotated using color markers. Then, color segmentation and occlusion modeling techniques are applied for image preprocessing, enabling preliminary localization of the target region. Building on this, the method integrates GoogLeNet with gradient-weighted class activation mapping (Grad-CAM) for fine-grained binary classification within the identified region, enhancing the detection of key target areas. A Kalman filter is subsequently employed to perform dynamic system state estimation, achieving real-time tracking performance. To further improve robustness under varying lighting conditions, an image reconstruction strategy based on average grayscale values is introduced, effectively mitigating the impact of illumination interference on tracking accuracy. Experimental results demonstrate that the proposed CMTT method achieves substantial performance gains across several benchmark datasets. On the visual object tracking (VOT) dataset, compared to the baseline model, adaptive target classification and interaction strategy (ATCAIS), CMTT achieves a 34. 3 % increase in expected average overlap (EAO), an 8. 0 % improvement in accuracy (ACC), and a 23. 8 % enhancement in Robustness (ROB). On the depth-based tracking benchmark (DepthTrack) dataset, CMTT achieves 31. 9 %, 38. 7 %, and 25. 6 % in EAO, ACC, and ROB, respectively, highlighting its superior stability and adaptability in complex environments. An empirical study with 15 participants confirmed the consistency between the proposed method and data from inertial measurement units (IMU).

JBHI Journal 2026 Journal Article

Enhancing the Reliability of Affective Brain-Computer Interfaces by Using Specifically Designed Confidence Estimator

  • Jiaheng Wang
  • Zhenyu Wang
  • Tianheng Xu
  • Ang Li
  • Yuan Si
  • Ting Zhou
  • Xi Zhao
  • Honglin Hu

In recent years, the diverse applications of electroencephalography (EEG) - based affective brain-computer interfaces (aBCIs) are being extensively explored. However, due to adverse factors like noise and physiological variability, the recognition capability of aBCIs can unforeseeably suffer abrupt declines. Since the timing of these aBCI failures is unknown, placing trust in aBCIs without scrutiny can lead to undesirable consequences. To alleviate this issue, we propose an algorithm for estimating the reliability of aBCI (primarily Graph Convolutional Network), synchronously delivering a probabilistic confidence score upon aBCI decision completion, thereby reflecting the aBCI’s real-time recognition capabilities. Methodologically, we use the Maximum Softmax Probability (MSP) from EEG recognition networks as confidence scores and leverage the Scaling Operator to calibrate them. Then, the Projection Operator is employed to address confidence estimation biases caused by noise and subject variability. For the numerical concentration of MSP, we provide fresh insights into its causes and propose corresponding solutions. The derivation of the estimator from the Maximum Entropy Principle is also substantiated for robust theoretical underpinnings. Finally, we confirm theoretically that the estimator does not compromise BCI performance. In experiments conducted on public datasets SEED and SEED-IV, the proposed algorithm demonstrates superior performance in estimating aBCIs reliability compared to other benchmarks, and commendable adaptability to new subjects. This research has the potential to lead to more trustworthy aBCIs and advance their broader application in complex real-world scenarios.

AAAI Conference 2026 Conference Paper

MUSE: Multimodal Uncertainty-Based Self-Driven Evolution for Robust Physiological-Signal–Based Driver Fatigue Detection

  • Jiaheng Wang
  • Yuan Si
  • Ang Li
  • Zhenyu Wang
  • Tianheng Xu
  • Honglin Hu

Precise detection of driver mental fatigue is critical for reducing traffic accidents and enhancing road safety. Compared with vision-based detection—which is susceptible to illumination and occlusion—multimodal physiological‑signal-based approaches integrate complementary information from diverse biosignals, delivering more faithful and objective fatigue assessments. However, adverse factors such as motion artifacts and environmental noise induce ceaseless deterioration to physiological signals, which markedly degrade the performance of existing multimodal fusion methods. To address this challenge, we propose Multimodal Uncertainty-based Self-driven Evolution, MUSE, reallocating modality contributions in real time via overall uncertainty minimization, thereby enabling efficient collaborative fusion of multi‐source predictions. Theoretically, MUSE guarantees a provably bounded cumulative error, and its generalization error approaches the Bayesian‑optimal fusion as iterations progress. Operating in a closed loop without labels or manual recalibration, MUSE presents superior suitability for real‑world driving scenarios compared to supervised algorithms. On the large‑scale driving fatigue dataset SEED‑VIG, MUSE outperforms existing models in both classification and regression tasks, substantiating its robustness and practicality as a promising driving fatigue detection solution.

EAAI Journal 2025 Journal Article

A novel semi-local centrality to identify influential nodes in complex networks by integrating multidimensional factors

  • Kun Zhang
  • Zaiyi Pu
  • Chuan Jin
  • Yu Zhou
  • Zhenyu Wang

This study addresses the critical problem of identifying influential nodes in complex networks, a task that plays a pivotal role in understanding network dynamics, optimizing information spread, and controlling epidemic outbreaks. Although semi-local centrality metrics are a valid approach for identifying influential nodes, they face challenges such as inefficiency when dealing with large-scale networks and neglecting semantic relationships, often relying on unidimensional criteria that limit their effectiveness. To tackle this challenge, this study presents a novel Semi-Local Centrality metric designed to identify influential nodes in complex networks by incorporating Multidimensional Factors (SLCMF). SLCMF combines structural, social, and semantic factors to find seed nodes in complex networks. To improve scalability, SLCMF utilizes distributed local subgraphs and redefines semi-local centrality by employing the average shortest path theory. Additionally, SLCMF incorporates a semantic graph embedding model by an augmented graph to capture distant and latent relationships among nodes. Extensive experiments on real-world networks demonstrate the effectiveness and efficiency of the proposed centrality metric, showcasing its superior performance in ranking influential nodes. Specifically, SLCMF outperforms the best traditional and advanced centrality metrics, improving Kendall's correlation coefficient by 8. 94% and 1. 61%, respectively. Additionally, the proposed metric demonstrates enhanced efficiency, reducing runtime by 4. 7% and 0. 21% compared to the top-performing traditional and advanced metrics, respectively.

JBHI Journal 2025 Journal Article

BSAN: A Self-Adapted Motor Imagery Decoding Framework Based on Contextual Information

  • Zikai Wang
  • Ang Li
  • Zhenyu Wang
  • Ting Zhou
  • Tianheng Xu
  • Honglin Hu

In motor imagery (MI) decoding, it still remains challenging to excavate enough contextual information of MI in different brain regions and to bridge the cross-session variance in feature distributions. In light of these issues, our study presents an innovative Bi-Stream Adaptation Network (BSAN) to bolster network efficacy, aiming to improve MI-based brain-computer interface (BCI) robustness across sessions. Our framework consists of the Bi-attention module, feature extractor, classifier, and Bi-discriminator. Precisely, we devise the Bi-attention module to reveal granular context information of MI with performing multi-scale convolutions asymptotically. Then, after features extraction, Bi-discriminator is involved to align the features from different MI sessions such that a uniform and accurate representation of neural patterns is achieved. By such a workflow, the proposed BSAN allows for the effective fusion of context coherence and session-invariance within the network architecture, therefore diminishing the reliance of redundant MI trials for MI-BCI re-calibration. To empirically substantiate BSAN, comprehensive experiments are conducted based on two public MI datasets. With average accuracies of 78. 97% and 83. 79% on two public datasets, and an inference time of 2. 99 ms on CPU-only devices, it is believed that our approach has the potential to accelerate the practical deployment of MI-BCI.

ICRA Conference 2025 Conference Paper

Geometry and Force-Informed Robotic Assembly with Small Relative Initial Deviations for Circular Electrical Connectors

  • Zhenyu Wang
  • Xiangfei Li
  • Huan Zhao 0001
  • Lingjun Shao
  • Hao Zhang
  • Han Ding 0001

Circular electrical connectors (CECs) have a wide range of applications in scenarios that require reliable connections. However, sockets are often located in narrow scenes with random spatial orientations, complex lighting conditions, and obstructions from cables, making it difficult to accurately locate them through cameras. Besides, due to the complex geometric structure of CECs and the presence of electrode protection slots, the existing research on the assembly of cylindrical or polygonal pegs and holes may not be applicable to the assembly of such components. To this end, this article proposes a novel robotic assembly strategy for CECs with small relative initial deviations, whose core is to design a search trajectory and heuristic force strategy to perceive force/pose (F/P) discontinuity characteristics under different geometric constraints. This assembly strategy is independent of the CEC's size and is not affected by the socket's spatial orientation. The experiments with two different sizes of CECs on a robot equipped with a 6-dimensional force/torque ( $\mathbf{F} / \mathbf{T}$ ) sensor are conducted, and the effectiveness and robustness of the proposed assembly strategy for CECs are demonstrated.

IJCAI Conference 2025 Conference Paper

PatternCIR Benchmark and TisCIR: Advancing Zero-Shot Composed Image Retrieval in Remote Sensing

  • Zhechun Liang
  • Tao Huang
  • Fangfang Wu
  • Shiwen Xue
  • Zhenyu Wang
  • Weisheng Dong
  • Xin Li
  • Guangming Shi

Remote sensing composed image retrieval (RSCIR) is a new vision-language task that takes a composed query of an image and text, aiming to search for a target remote sensing image satisfying two conditions from intricate remote sensing imagery. However, the existing attribute-based benchmark Patterncom in RSCIR has significant flaws, including the lack of query text sentences and paired triplets, thus making it unable to evaluate the latest methods. To address this, we propose the Zero-Shot Query Text Generator (ZS-QTG) that can generate full query text sentences based on attributes, and then, by capitalizing on ZS-QTG, we develop the PatternCIR benchmark. PatternCIR rectifies Patterncom’s deficiencies and enables the evaluation of existing methods. Additionally, we explore zero-shot composed image retrieval methods that do not rely on massive pre-collected triplets for training. Existing methods use only the text during retrieval, performing poorly in RSCIR. To improve this, we propose Text-image Sequential Training of Composed Image Retrieval (TisCIR). TisCIR undergoes sequential training of multiple self-masking projection and fine-grained image attention modules, which endows it with the capacity to filter out conflicting information between the image and text, enhancing the retrieval by utilizing both modalities in harmony. TisCIR outperforms existing methods by 12. 40% to 62. 03% on PatternCIR, achieving state-of-the-art performance in RSCIR. The data and code are available here.

ICLR Conference 2025 Conference Paper

PortLLM: Personalizing Evolving Large Language Models with Training-Free and Portable Model Patches

  • Rana Muhammad Shahroz
  • Pingzhi Li
  • Sukwon Yun
  • Zhenyu Wang
  • Shahriar Nirjon
  • Chau-Wai Wong
  • Tianlong Chen 0001

As large language models (LLMs) increasingly shape the AI landscape, fine-tuning pretrained models has become more popular than in the pre-LLM era for achieving optimal performance in domain-specific tasks. However, pretrained LLMs such as ChatGPT are periodically evolved (i.e., model parameters are frequently updated), making it challenging for downstream users with limited resources to keep up with fine-tuning the newest LLMs for their domain application. Even though fine-tuning costs have nowadays been reduced thanks to the innovations of parameter-efficient fine-tuning such as LoRA, not all downstream users have adequate computing for frequent personalization. Moreover, access to fine-tuning datasets, particularly in sensitive domains such as healthcare, could be time-restrictive, making it crucial to retain the knowledge encoded in earlier fine-tuned rounds for future adaptation. In this paper, we present PORTLLM, a training-free framework that (i) creates an initial lightweight model update patch to capture domain-specific knowledge, and (ii) allows a subsequent seamless plugging for the continual personalization of evolved LLM at minimal cost. Our extensive experiments cover seven representative datasets, from easier question-answering tasks {BoolQ, SST2} to harder reasoning tasks {WinoGrande, GSM8K}, and models including {Mistral-7B,Llama2, Llama3.1, and Gemma2}, validating the portability of our designed model patches and showcasing the effectiveness of our proposed framework. For instance, PORTLLM achieves comparable performance to LoRA fine-tuning with reductions of up to 12.2× in GPU memory usage. Finally, we provide theoretical justifications to understand the portability of our model update patches, which offers new insights into the theoretical dimension of LLMs’ personalization.

NeurIPS Conference 2025 Conference Paper

Towards 3D Objectness Learning in an Open World

  • Taichi Liu
  • Zhenyu Wang
  • Ruofeng Liu
  • Guang Wang
  • Desheng Zhang

Recent advancements in 3D object detection and novel category detection have made significant progress, yet research on learning generalized 3D objectness remains insufficient. In this paper, we delve into learning open-world 3D objectness, which focuses on detecting all objects in a 3D scene, including novel objects unseen during training. Traditional closed-set 3D detectors struggle to generalize to open-world scenarios, while directly incorporating 3D open-vocabulary models for open-world ability struggles with vocabulary expansion and semantic overlap. To achieve generalized 3D object discovery, We propose OP3Det, a class-agnostic Open-World Prompt-free 3D Detector to detect any objects within 3D scenes without relying on hand-crafted text prompts. We introduce the strong generalization and zero-shot capabilities of 2D foundation models, utilizing both 2D semantic priors and 3D geometric priors for class-agnostic proposals to broaden 3D object discovery. Then, by integrating complementary information from point cloud and RGB image in the cross-modal mixture of experts, OP3Det dynamically routes uni-modal and multi-modal features to learn generalized 3D objectness. Extensive experiments demonstrate the extraordinary performance of OP3Det, which significantly surpasses existing open-world 3D detectors by up to 16. 0% in AR and achieves a 13. 5% improvement compared to closed-world 3D detectors.

AAAI Conference 2024 Conference Paper

BVT-IMA: Binary Vision Transformer with Information-Modified Attention

  • Zhenyu Wang
  • Hao Luo
  • Xuemei Xie
  • Fan Wang
  • Guangming Shi

As a compression method that can significantly reduce the cost of calculations and memories, model binarization has been extensively studied in convolutional neural networks. However, the recently popular vision transformer models pose new challenges to such a technique, in which the binarized models suffer from serious performance drops. In this paper, an attention shifting is observed in the binary multi-head self-attention module, which can influence the information fusion between tokens and thus hurts the model performance. From the perspective of information theory, we find a correlation between attention scores and the information quantity, further indicating that a reason for such a phenomenon may be the loss of the information quantity induced by constant moduli of binarized tokens. Finally, we reveal the information quantity hidden in the attention maps of binary vision transformers and propose a simple approach to modify the attention values with look-up information tables so that improve the model performance. Extensive experiments on CIFAR-100/TinyImageNet/ImageNet-1k demonstrate the effectiveness of the proposed information-modified attention on binary vision transformers.

NeurIPS Conference 2024 Conference Paper

GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing

  • Zhenyu Wang
  • Aoxue Li
  • Zhenguo Li
  • Xihui Liu

Despite the success achieved by existing image generation and editing methods, current models still struggle with complex problems including intricate text prompts, and the absence of verification and self-correction mechanisms makes the generated images unreliable. Meanwhile, a single model tends to specialize in particular tasks and possess the corresponding capabilities, making it inadequate for fulfilling all user requirements. We propose GenArtist, a unified image generation and editing system, coordinated by a multimodal large language model (MLLM) agent. We integrate a comprehensive range of existing models into the tool library and utilize the agent for tool selection and execution. For a complex problem, the MLLM agent decomposes it into simpler sub-problems and constructs a tree structure to systematically plan the procedure of generation, editing, and self-correction with step-by-step verification. By automatically generating missing position-related inputs and incorporating position information, the appropriate tool can be effectively employed to address each sub-problem. Experiments demonstrate that GenArtist can perform various generation and editing tasks, achieving state-of-the-art performance and surpassing existing models such as SDXL and DALL-E 3, as can be seen in Fig. 1. We will open-source the code for future research and applications.

NeurIPS Conference 2024 Conference Paper

One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object Detection

  • Zhenyu Wang
  • Yali Li
  • Hengshuang Zhao
  • Shengjin Wang

The current trend in computer vision is to utilize one universal model to address all various tasks. Achieving such a universal model inevitably requires incorporating multi-domain data for joint training to learn across multiple problem scenarios. In point cloud based 3D object detection, however, such multi-domain joint training is highly challenging, because large domain gaps among point clouds from different datasets lead to the severe domain-interference problem. In this paper, we propose OneDet3D, a universal one-for-all model that addresses 3D detection across different domains, including diverse indoor and outdoor scenes, within the same framework and only one set of parameters. We propose the domain-aware partitioning in scatter and context, guided by a routing mechanism, to address the data interference issue, and further incorporate the text modality for a language-guided classification to unify the multi-dataset label spaces and mitigate the category interference issue. The fully sparse structure and anchor-free head further accommodate point clouds with significant scale disparities. Extensive experiments demonstrate the strong universal ability of OneDet3D to utilize only one trained model for addressing almost all 3D object detection tasks (Fig. 1). We will open-source the code for future research and applications.

JBHI Journal 2024 Journal Article

Riemannian Locality Preserving Method for Transfer Learning With Applications on Brain-Computer Interface

  • Guiying Xu
  • Zhenyu Wang
  • Honglin Hu
  • Xi Zhao
  • Ruxue Li
  • Ting Zhou
  • Tianheng Xu

Brain-computer interfaces (BCIs) have been widely focused and extensively studied in recent years for their huge prospect of medical rehabilitation and commercial applications. Transfer learning exploits the information in the source domain and applies in another different but related domain (target domain), and is therefore introduced into the BCIs to figure out the inter-subject variances of electroencephalography (EEG) signals. In this article, a novel transfer learning method is proposed to preserve the Riemannian locality of data structure in both the source and target domains and simultaneously realize the joint distribution adaptation of both domains to enhance the effectiveness of transfer learning. Specifically, a Riemannian graph is first defined and constructed based on the Riemannian distance to represent the Riemannian geometry information. To simultaneously align the marginal and conditional distribution of source and target domains and preserve the Riemannian locality of data structure in both domains, the Riemannian graph is embedded in the joint distribution adaptation (JDA) framework and forms the proposed Riemannian locality preserving-based transfer learning (RLPTL). To validate the effect of the proposed method, it is compared with several existing methods on two open motor imagery datasets, and both multi-source domains (MSD) and single-source domains (SSD) experiments are considered. Experimental results show that the proposed method achieves the highest accuracies in MSD and SSD experiments on three datasets and outperforms eight baseline methods, which demonstrates that the proposed method creates a feasible and efficient way to realize transfer learning.

NeurIPS Conference 2024 Conference Paper

RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation

  • Jiaming Liu
  • Mengzhen Liu
  • Zhenyu Wang
  • Pengju An
  • Xiaoqi Li
  • Kaichen Zhou
  • Senqiao Yang
  • Renrui Zhang

A fundamental objective in robot manipulation is to enable models to comprehend visual scenes and execute actions. Although existing Vision-Language-Action (VLA) models for robots can handle a range of basic tasks, they still face challenges in two areas: (1) insufficient reasoning ability to tackle complex tasks, and (2) high computational costs for VLA model fine-tuning and inference. The recently proposed state space model (SSM) known as Mamba demonstrates promising capabilities in non-trivial sequence modeling with linear inference complexity. Inspired by this, we introduce RoboMamba, an end-to-end robotic VLA model that leverages Mamba to deliver both robotic reasoning and action capabilities, while maintaining efficient fine-tuning and inference. Specifically, we first integrate the vision encoder with Mamba, aligning visual tokens with language embedding through co-training, empowering our model with visual common sense and robotic-related reasoning. To further equip RoboMamba with SE(3) pose prediction abilities, we explore an efficient fine-tuning strategy with a simple policy head. We find that once RoboMamba possesses sufficient reasoning capability, it can acquire manipulation skills with minimal fine-tuning parameters (0. 1\% of the model) and time. In experiments, RoboMamba demonstrates outstanding reasoning capabilities on general and robotic evaluation benchmarks. Meanwhile, our model showcases impressive pose prediction results in both simulation and real-world experiments, achieving inference speeds 3 times faster than existing VLA models.

EAAI Journal 2023 Journal Article

A novel seminar learning framework for weakly supervised salient object detection

  • Yan Liu
  • Yunzhou Zhang
  • Zhenyu Wang
  • Fei Yang
  • Feng Qiu
  • Sonya Coleman
  • Dermot Kerr

Weakly supervised salient object detection (SOD) is a challenging task and has drawn much attention from several research perspectives, it has revealed two problems while driving the rapid development of saliency detection. (1) Large divergence in the characteristics of saliency regions in terms of location, shape and size makes them difficult to recognize. (2) The properties of convolutional neural networks dictate that it is insensitive to various transformations, which will lead to hardly balance the application of various disturbances. To tackle these limitations, this paper proposes a novel seminar learning framework with consistent transformation ensembling (SLF-CT) for scribble supervised SOD. The framework consists of the teacher–student model and the student–student model for segmenting the salient objects. Specifically, we first design a cross attention guided network (CAGNet) as a baseline model for saliency prediction. Then we assign CAGNet to the teacher–student model, where the teacher network is based on the exponential moving average and guides the training of the student network. Moreover, we adopt multiple pseudo labels to transfer the information among students from different conditions. To further enhance the regularization of the network, a consistency transformation mechanism is also incorporated, which encourages the saliency prediction and input image of the network to be consistent. The experimental results demonstrate that the proposed approach performs favorably comparable with the state-of-the-art weakly supervised methods. As far as we know, the proposed approach is the first application of seminar learning in the SOD area.

NeurIPS Conference 2023 Conference Paper

Uni3DETR: Unified 3D Detection Transformer

  • Zhenyu Wang
  • Ya-Li Li
  • Xi Chen
  • Hengshuang Zhao
  • Shengjin Wang

Existing point cloud based 3D detectors are designed for the particular scene, either indoor or outdoor ones. Because of the substantial differences in object distribution and point density within point clouds collected from various environments, coupled with the intricate nature of 3D metrics, there is still a lack of a unified network architecture that can accommodate diverse scenes. In this paper, we propose Uni3DETR, a unified 3D detector that addresses indoor and outdoor 3D detection within the same framework. Specifically, we employ the detection transformer with point-voxel interaction for object prediction, which leverages voxel features and points for cross-attention and behaves resistant to the discrepancies from data. We then propose the mixture of query points, which sufficiently exploits global information for dense small-range indoor scenes and local information for large-range sparse outdoor ones. Furthermore, our proposed decoupled IoU provides an easy-to-optimize training target for localization by disentangling the $xy$ and $z$ space. Extensive experiments validate that Uni3DETR exhibits excellent performance consistently on both indoor and outdoor 3D detection. In contrast to previous specialized detectors, which may perform well on some particular datasets but suffer a substantial degradation on different scenes, Uni3DETR demonstrates the strong generalization ability under heterogeneous conditions (Fig. 1).

NeurIPS Conference 2022 Conference Paper

VTC-LFC: Vision Transformer Compression with Low-Frequency Components

  • Zhenyu Wang
  • Hao Luo
  • Pichao Wang
  • Feng Ding
  • Fan Wang
  • Hao Li

Although Vision transformers (ViTs) have recently dominated many vision tasks, deploying ViT models on resource-limited devices remains a challenging problem. To address such a challenge, several methods have been proposed to compress ViTs. Most of them borrow experience in convolutional neural networks (CNNs) and mainly focus on the spatial domain. However, the compression only in the spatial domain suffers from a dramatic performance drop without fine-tuning and is not robust to noise, as the noise in the spatial domain can easily confuse the pruning criteria, leading to some parameters/channels being pruned incorrectly. Inspired by recent findings that self-attention is a low-pass filter and low-frequency signals/components are more informative to ViTs, this paper proposes compressing ViTs with low-frequency components. Two metrics named low-frequency sensitivity (LFS) and low-frequency energy (LFE) are proposed for better channel pruning and token pruning. Additionally, a bottom-up cascade pruning scheme is applied to compress different dimensions jointly. Extensive experiments demonstrate that the proposed method could save 40% ~ 60% of the FLOPs in ViTs, thus significantly increasing the throughput on practical devices with less than 1% performance drop on ImageNet-1K.

AAAI Conference 2021 Conference Paper

Automatic Curriculum Learning With Over-repetition Penalty for Dialogue Policy Learning

  • Yangyang Zhao
  • Zhenyu Wang
  • Zhenhua Huang

Dialogue policy learning based on reinforcement learning is difficult to be applied to real users to train dialogue agents from scratch because of the high cost. User simulators, which choose random user goals for the dialogue agent to train on, have been considered as an affordable substitute for real users. However, this random sampling method ignores the law of human learning, making the learned dialogue policy inefficient and unstable. We propose a novel framework, Automatic Curriculum Learning-based Deep Q- Network (ACL-DQN), which replaces the traditional random sampling method with a teacher policy model to realize the dialogue policy for automatic curriculum learning. The teacher model arranges a meaningful ordered curriculum and automatically adjusts it by monitoring the learning progress of the dialogue agent and the over-repetition penalty without any requirement of prior knowledge. The learning progress of the dialogue agent reflects the relationship between the dialogue agent’s ability and the sampled goals’ difficulty for sample efficiency. The over-repetition penalty guarantees the sampled diversity. Experiments show that the ACL-DQN significantly improves the effectiveness and stability of dialogue tasks with a statistically significant margin. Furthermore, the framework can be further improved by equipping with different curriculum schedules, which demonstrates that the framework has strong generalizability.

NeurIPS Conference 2021 Conference Paper

Combating Noise: Semi-supervised Learning by Region Uncertainty Quantification

  • Zhenyu Wang
  • Ya-Li Li
  • Ye Guo
  • Shengjin Wang

Semi-supervised learning aims to leverage a large amount of unlabeled data for performance boosting. Existing works primarily focus on image classification. In this paper, we delve into semi-supervised learning for object detection, where labeled data are more labor-intensive to collect. Current methods are easily distracted by noisy regions generated by pseudo labels. To combat the noisy labeling, we propose noise-resistant semi-supervised learning by quantifying the region uncertainty. We first investigate the adverse effects brought by different forms of noise associated with pseudo labels. Then we propose to quantify the uncertainty of regions by identifying the noise-resistant properties of regions over different strengths. By importing the region uncertainty quantification and promoting multi-peak probability distribution output, we introduce uncertainty into training and further achieve noise-resistant learning. Experiments on both PASCAL VOC and MS COCO demonstrate the extraordinary performance of our method.

AAAI Conference 2021 Short Paper

Melodic Phrase Attention Network for Symbolic Data-based Music Genre Classification (Student Abstract)

  • Li Li
  • Rui Zhang
  • Zhenyu Wang

Compared with audio data-based music genre classification, researches on symbolic data-based music are scarce. Existing methods generally utilize manually extracted features, which is very time-consuming and laborious, and use traditional classifiers for label prediction without considering specific music features. To tackle this issue, we propose the Melodic Phrase Attention Network (MPAN) for symbolic data-based music genre classification. Our model is trained in three steps: First, we adopt representation learning, instead of the traditional musical feature extraction method, to obtain a vectorized representation of the music pieces. Second, the music pieces are divided into several melodic phrases through melody segmentation. Finally, the Melodic Phrase Attention Network is designed according to music characteristics, to identify the reflection of each melodic phrase on the music genre, thereby generating more accurate predictions. Experimental results show that our proposed method is superior to baseline symbolic data-based music genre classification approaches, and has achieved significant performance improvements on two large datasets.

AAAI Conference 2020 Conference Paper

Dynamic Reward-Based Dueling Deep Dyna-Q: Robust Policy Learning in Noisy Environments

  • Yangyang Zhao
  • Zhenyu Wang
  • Kai Yin
  • Rui Zhang
  • Zhenhua Huang
  • Pei Wang

Task-oriented dialogue systems provide a convenient interface to help users complete tasks. An important consideration for task-oriented dialogue systems is the ability to against the noise commonly existed in the real-world conversation. Both rule-based strategies and statistical modeling techniques can solve noise problems, but they are costly. In this paper, we propose a new approach, called Dynamic Reward-based Dueling Deep Dyna-Q (DR-D3Q). The DR-D3Q can learn policies in noise robustly, and it is easy to implement by combining dynamic reward and the Dueling Deep Q-Network (Dueling DQN) into Deep Dyna-Q (DDQ) framework. The Dueling DQN can mitigate the negative impact of noise on learning policies, but it is inapplicable to dialogue domain due to different reward mechanisms. Unlike typical dialogue reward function, we integrate dynamic reward that provides reward in real-time for agent to make Dueling DQN adapt to dialogue domain. For the purpose of supplementing the limited amount of real user experiences, we take the DDQ framework as the basic framework. Experiments using simulation and human evaluation show that the DR-D3Q significantly improve the performance of policy learning tasks in noisy environments. 1

v2026.09.13