Arrow Research search

Author name cluster

Wei Yu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

30 papers
2 author rows

Possible papers

30

EAAI Journal 2026 Journal Article

A novel model-data-driven sample generation approach for bearing fault diagnosis under imbalanced data conditions

  • Zhimin Wang
  • Wei Yu
  • Jianwu Wang
  • Lizhang Cheng
  • Chuanchuan Shao

Bearing fault diagnosis is vital for equipment reliability, but the scarcity of fault data in industrial settings hinders the performance and generalization of diagnostic models. To address this issue, a model-data-driven sample generation (MDDSG) framework is proposed that integrates physics-based modeling with data-driven augmentation. First, a nonlinear dynamic model is constructed to simulate vibration signals under various fault types, thereby enriching the sample space. Then, a feature spectrum-based generation criterion is introduced, which utilizes healthy signals as prior knowledge to extract device-specific features. These features are integrated with model-generated fault characteristics to guide realistic sample generation and reduce distribution discrepancies. To further enhance data diversity and authenticity, a threshold-guided Generative Adversarial Network is employed, incorporating an adaptive similarity-based training strategy. Finally, training on the generated dataset yields robust diagnostic performance, with average accuracies of 98. 35% on the synthetic dataset and 96. 95% on the real-world dataset. These results validate the ability of MDDSG to overcome data imbalance and support robust industrial fault diagnosis.

AAAI Conference 2026 Conference Paper

Communication-efficient Multi-Agent Reinforcement Learning with Spatiotemporal Information Hub

  • Ling Ding
  • Tianbai Lyu
  • Zhiliang Bi
  • Hao Wang
  • Shanshan Feng
  • Wei Yu

Centralized training with decentralized execution (CTDE) is a framework for MARL with wide applications. In the CTDE paradigm, agents leverage global state information during training to mitigate the non-stationarity of the MARL environment, but must rely solely on partial observations during execution. Recent work has highlighted the growing importance of inter-agent communication for more effective learning and coordination. However, most existing methods overlook the fact that real-world communication channels are often bandwidth-constrained and imperfectly reliable. Toward more communication-efficient and robust MARL, we extend the conventional CTDE framework with an information hub. The hub collects local observations from the agents to restore the global state, which is then delivered to the agents on demand. To this end, technical mechanisms are designed to enable effective global reconstruction with incomplete observations, as well as agent-specific attention to the reconstructed global information. Experiments on multiple cooperative MARL benchmarks demonstrate that our method achieves state-of-the-art performance compared to popular MARL algorithms while substantially reducing communication overhead and exhibiting strong robustness under imperfect communication channels.

AAAI Conference 2026 Conference Paper

DialoGen: Towards Dialog Gesture Generation via Identity-Decoupled Style Guidance in Interactive Diffusion Model

  • Weiyu Zhao
  • Chenyang Wang
  • Liangxiao Hu
  • Zonglin Li
  • Wei Yu
  • Shengping Zhang

We propose DialoGen, a novel framework for generating realistic gestures for both interlocutors in dialog scenarios, conditioned on conversational audios. Unlike most existing methods that focus solely on a single speaker, DialoGen simultaneously generates synchronized gestures for both participants while also embedding identity-decoupled style into generated gestures that enhance realism and expressiveness. To ensure precise synchronization between interlocutors, DialoGen adopts an interactive dual-diffusion model with mutual interaction estimation, which integrates interaction correlation into the diffusion process. More importantly, by leveraging supervised contrastive learning, we develop the identity-decoupled style guidance to adaptively decompose the identity-specific style of interlocutors into latent space, enabling multi-style dialog gesture generation. Extensive experimental results demonstrate that our model significantly outperforms existing methods in generating realistic, speech-aligned, identity-specific gestures, offering a high-quality solution for various dialog scenarios.

IROS Conference 2025 Conference Paper

Design and Performance Analysis of a Pipeline Crawling Robot Based on Spring-Roll Dielectric Elastomer Actuators

  • Qinghai Zhang
  • Wei Yu
  • Ziqi Zhang
  • Jianghua Zhao
  • Shijie Guo

With the increasing complexity of pipeline systems in various industrial and environmental applications, there is a critical need for flexible and efficient robotic solutions that can navigate and inspect confined spaces. This paper introduces a lightweight pipeline crawling robot based on spring-roll dielectric elastomer actuators (DEAs). Inspired by the adaptability of caterpillars, the robot combines anisotropic friction feet with a spring-roll DEA structure to achieve high-speed movement. It operates effectively in pipes with diameters ranging from 16 mm to 20 mm, reaching a maximum speed of 357 mm/s (5. 95 BL/s) under a 3. 5 kV driving voltage. The optimized design enhances actuator performance and friction distribution, significantly outperforming existing soft crawling robots. This innovation demonstrates great potential for high-speed, lightweight pipeline inspection applications and advances the field of soft robotics for diverse industrial tasks.

AIIM Journal 2025 Journal Article

Difficulty-aware coupled contour regression network with IoU loss for efficient IVUS delineation

  • Yuan Yang
  • Xu Yu
  • Wei Yu
  • Shengxian Tu
  • Su Zhang
  • Wei Yang

The lumen and external elastic lamina contour delineation is crucial for quantitative analyses of intravascular ultrasound (IVUS) images. However, the various artifacts in IVUS images pose substantial challenges for accurate delineation. Existing mask-based methods often produce anatomically implausible contours in artifact-affected images, while contour-based methods suffer from the over-smooth problem within the artifact regions. In this paper, we directly regress the contour pairs instead of mask-based segmentation. A coupled contour representation is adopted to learn a low-dimensional contour signature space, where the embedded anatomical prior enables the model to avoid producing unreasonable results. Further, a PIoU loss is proposed to capture the overall shape of the contour points and maximize the similarity between the regressed contours and manually delineated contours with various irregular shapes, alleviating the over-smooth problem. For the images with severe artifacts, a difficulty-aware training strategy is designed for contour regression, which gradually guides the model focus on hard samples and improves contour localization accuracy. We evaluate the proposed framework on a large IVUS dataset, consisting of 7204 frames from 185 pullbacks. The mean Dice similarity coefficients of the method for the lumen and external elastic lamina are 0. 951 and 0. 967, which significantly outperforms other state-of-the-art (SOTA) models. All regressed contours in the test images are anatomically plausible. On the public IVUS-2011 dataset, the proposed method attains comparable performance to the SOTA models with the highest processing speed at 100 fps. The code is available at https: //github. com/SMU-MedicalVision/ContourRegression.

AAAI Conference 2025 Conference Paper

Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

  • Wenbin Wang
  • Liang Ding
  • Minyan Zeng
  • Xiabin Zhou
  • Li Shen
  • Yong Luo
  • Wei Yu
  • Dacheng Tao

Multimodal large language models (MLLMs) have experienced significant advancements recently, but still struggle to recognize and interpret intricate details in high-resolution (HR) images effectively. While state-of-the-art (SOTA) MLLMs claim to process images at 4K resolution, existing MLLM benchmarks only support up to 2K, leaving the capabilities of SOTA models on true HR images largely untested. Furthermore, existing methods for enhancing HR image perception in MLLMs rely on computationally expensive visual instruction tuning. To address these limitations, we introduce HR-Bench, the first deliberately designed benchmark to rigorously evaluate MLLM performance on 4K & 8K images. Through extensive experiments, we demonstrate that while downsampling HR images leads to vision information loss, leveraging complementary modalities, e.g., text, can effectively compensate for this loss. Building upon this insight, we propose Divide, Conquer and Combine, a novel training-free framework for enhancing MLLM perception of HR images. Our method follows a three-staged approach: 1) Divide: recursively partitioning the HR image into patches and merging similar patches to minimize computational overhead, 2) Conquer: leveraging the MLLM to generate accurate textual descriptions for each image patch, and 3) Combine: utilizing the generated text descriptions to enhance the MLLM's understanding of the overall HR image. Extensive experiments show that: 1) the SOTA MLLM achieves 63% accuracy, which is markedly lower than the 87% accuracy achieved by humans on HR-Bench; 2) our method brings consistent and significant improvements (a relative increase of +6% on HR-Bench and +8% on general multimodal benchmarks).

JBHI Journal 2025 Journal Article

DPPAT: Dual-Level Periodic Pattern-Aware Transformer for Heart Sound Murmur Identification

  • Zilan Hong
  • Wei Yu
  • Chunming Li
  • Botao Yang
  • Zehao Fan
  • Runguo Wei
  • Shengxian Tu

Developing heart sound classification algorithms for murmur identification is critical for early screening of heart diseases. However, identifying murmurs in long-duration heart sound signals can be challenging due to their weak features and interference from noise. Considering the periodic patterns of heart sounds and murmurs, periodic priors can be introduced to enhance murmur identification, an approach that remains underutilized in current methods. In this study, we propose a novel Dual-level Periodic Pattern-Aware Transformer (DPPAT) to implicitly leverage the periodic priors of heart sound signals without requiring cycle segmentation. In the regional-level, an Adaptive Period-Aligned Window Selection algorithm is designed for the model to extract periodic components while suppressing random noise using a Periodic Pattern Attention module. In the global-level, the model further integrates these periodic features in global-modeling to enhance the identification of murmur-discriminative features. Validated on the dataset from 2022 George B. Moody PhysioNet Challenge, our proposed method achieves a weighted accuracy of 84. 27% and an F1-score of 70. 38% through 10-fold cross-validation. The generalizability of DPPAT is further verified on two additional public datasets, including both heart sound and respiratory sound signals. Furthermore, attention visualizations provide a clear understanding of the focus of the model, highlighting the decision-making basis for murmur identification.

AAAI Conference 2025 Conference Paper

Federated Graph Anomaly Detection Through Contrastive Learning with Global Negative Pairs

  • Nannan Wu
  • Yazheng Zhao
  • Hongdou Dong
  • Keao Xi
  • Wei Yu
  • Wenjun Wang

Anomaly detection on attributed graphs has applications in various domains such as finance and email spam detection, thus gaining substantial attention. Distributed scenarios can also involve issues related to anomaly detection in attribute graphs, such as in medical scenarios. However, most of the existing anomaly detection methods are designed for centralized scenarios, and directly applying them to distributed settings may lead to reduced performance. One possible reason for this issue is that, when graph data are distributed across multiple clients, federated graph learning may struggle to fully exploit the potential of the dispersed data, leading to suboptimal performance. Building on this insight, we propose FedCLGN, a federated graph anomaly detection framework that leverages contrastive self-supervised learning. First, we put forward an augmentation method to maintain global negative pairs on the server. This involves identifying anomalous nodes using pseudo-labels, extracting embedding representations of the negative pairs corresponding to these anomalous nodes from clients, and uploading them to the server. Then, we adopt graph diffusion to enhance the feature representation of nodes, capturing the global structure and local connection patterns. This strategy can strengthen the differentiation between positive and negative instance pairs. Finally, the effectiveness of our approach is verified by experimental results on four real graph datasets.

TAAS Journal 2025 Journal Article

HAG-MTF: Higher-Order Adaptive Generative Graph for Massive Traffic Forecasting in Industry 5.0

  • Lei Wang
  • Huaming Wu
  • Fan Zhang
  • Keqiu Li
  • Wei Yu
  • Shuo Chen

With the evolution of urban smart transportation, the complexity of urban traffic networks escalates, emphasizing the importance of large-scale traffic data prediction in traffic management and urban planning. Traditional spatiotemporal graph models, such as Graph-WaveNet and MTGCN, face exponentially increasing computational complexity as the spatial dimensions expand. To address this challenge, we propose a novel Higher-order Adaptive Generative graph for Massive Traffic Forecasting (HAG-MTF) approach, which utilizes generative AI and high-order graph structures to model the intricate spatial dependencies in large-scale traffic data. The HAG-MTF incorporates a high-order dimensionality reduction module to optimize traffic node processing, utilizing prior graph relationships to generate a fusion graph that dynamically incorporates neighborhood information for efficient, localized graph convolution. The model further incorporates the high-order spatiotemporal relationship extraction module (H-net), enhancing the capacity and speed of traffic data processing while boosting prediction accuracy for complex spatial structures. Furthermore, HAG-MTF introduces a fusion loss function that hierarchically balances multiple objectives, ensuring both precision and computational efficiency. HAG-MTF adaptively handles large-scale real-world traffic data, meeting the needs of traffic controllers and urban planners for predicting massive datasets in practical settings. It supports efficient, flexible interactions via parameter tuning and model outputs, ultimately integrating human insights into traffic analysis and decision-making. This dynamic human-machine collaboration differs from non-Industry 5.0 approaches, which rely on purely automated systems without human input. Those lead to inflexible, brittle conclusions and recommendations, neglecting shifts in traffic patterns driven by human behavior. Extensive experiments on real-world traffic datasets demonstrate that HAG-MTF significantly improves processing efficiency for high-complexity spatial data while delivering precise, human-informed predictions through generative AI-driven operations.

AAAI Conference 2025 Conference Paper

MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-Reading

  • Wenhao Zhang
  • Jun Wang
  • Yong Luo
  • Lei Yu
  • Wei Yu
  • Zheng He
  • Jialie Shen

Lip-reading is to utilize the visual information of the speaker’s lip movements to recognize words and sentences. Existing event-based lip-reading solutions integrate different frame rate branches to learn spatio-temporal features of varying granularities. However, aggregating events into event frames inevitably leads to the loss of fine-grained temporal information within frames. To remedy this drawback, we propose a novel framework termed Multi-view Temporal Granularity aligned Aggregation (MTGA). Specifically, we first present a novel event representation method, namely time-segmented voxel graph list, where the most significant local voxels are temporally connected into a graph list. Then we design a spatio-temporal fusion module based on temporal granularity alignment, where the global spatial features extracted from event frames, together with the local relative spatial and temporal features contained in voxel graph list are effectively aligned and integrated. Finally, we design a temporal aggregation module that incorporates positional encoding, which enables the capture of local absolute spatial and global temporal information. Experiments demonstrate that our method outperforms both the event-based and video-based lip-reading counterparts.

AAAI Conference 2025 Conference Paper

OTPNet: ODE-inspired Tuning-free Proximal Network for Remote Sensing Image Fusion

  • Wei Yu
  • Zonglin Li
  • Qinglin Liu
  • Xin Sun

Remote sensing image fusion aims to reconstruct a high spatial and spectral resolution image by integrating the spatial and spectral information from multiple remote sensing sensor data. Despite the remarkable progress of deep learning-based fusion methods, most existing methods rely on manual network architecture design and hyperparameter tuning, lacking sufficient interpretability and adaptability. To address this limitation, we propose a novel neural Ordinary Differential Equation (ODE)-inspired tuning-free proximal splitting algorithm, which splits remote sensing image fusion as two optimization problems regularized by deep priors to model the fusion of spatial and spectral. Firstly, based on the physical properties of spatial and spectral information, the two problems are optimized by two proximal splitting operators to iteratively integrate spatial-spectral complementary information, eliminating or suppressing redundant information to reduce fusion errors. Secondly, considering the efficiency of neural ODE in reducing optimization error, we utilize a high-order numerical scheme to customize the proximal operator theoretically without additional handcrafted design and parameter tuning. Finally, by incorporating the numerical scheme as a solver into the proximal optimization algorithm, we derive an ODE-inspired Tuning-free Proximal Network, dubbed OTPNet, which achieves efficient and robust fusion reconstruction. Extensive experiments on nine datasets across three different remote sensing image fusion tasks show that our OTPNet outperforms existing state-of-the-art approaches, which validates the effectiveness of our method.

AAAI Conference 2025 Conference Paper

What Kind of Visual Tokens Do We Need? Training-Free Visual Token Pruning for Multi-Modal Large Language Models from the Perspective of Graph

  • Yutao Jiang
  • Qiong Wu
  • Wenhao Lin
  • Wei Yu
  • Yiyi Zhou

Recent Multimodal Large Language Models(MLLMs) often use a large number of visual tokens to compensate their visual shortcoming, leading to excessive computation and obvious visual redundancy. In this paper, we investigate what kind of visual tokens are needed for MLLMs, and reveal that both foreground and background tokens are critical for MLLMs given the varying difficulties of examples. Based on this observation, we propose a graph-based method towards training-free visual token pruning, termed G-Prune. In particular, G-Prune regards visual tokens as nodes, and construct their connections based on their semantic similarities. Afterwards, the information flow is propagated via weighted links, and the most important tokens after iterations are kept for MLLMs, which can be front or background. To validate G-Prune, we apply it to a recent MLLM called LLaVA-NeXT, and conduct extensive experiments on a set of benchmarks. The experiment results show that G-Prune can greatly reduce computation overhead while retaining high performance on both coarse- and fine-grained tasks. For instance, G-Prune can reduce 63.57% FLOPs of LLaVA-NeXT on VQA2.0 and TextVQA with only 0.95% and 2.34% accuracy drops, respectively.

IJCAI Conference 2024 Conference Paper

Joint Input and Output Coordination for Class-Incremental Learning

  • Shuai Wang
  • Yibing Zhan
  • Yong Luo
  • Han Hu
  • Wei Yu
  • Yonggang Wen
  • Dacheng Tao

Incremental learning is nontrivial due to severe catastrophic forgetting. Although storing a small amount of data on old tasks during incremental learning is a feasible solution, current strategies still do not 1) adequately address the class bias problem, and 2) alleviate the mutual interference between new and old tasks, and 3) consider the problem of class bias within tasks. In light of the above issues, we analyze the cause of class bias in incremental learning, as well as the drawbacks of existing approaches, and propose a joint input and output coordination (JIOC) mechanism to address these issues. This mechanism assigns different weights to different categories of data according to the gradient of the output score, and uses knowledge distillation (KD) to reduce the mutual interference between the outputs of old and new tasks. The proposed mechanism is general and flexible, and can be incorporated into different incremental learning approaches that use memory storage. Extensive experiments show that our mechanism can significantly improve their performance.

AAAI Conference 2024 Conference Paper

Levenshtein Distance Embedding with Poisson Regression for DNA Storage

  • Xiang Wei
  • Alan J.X. Guo
  • Sihan Sun
  • Mengyi Wei
  • Wei Yu

Efficient computation or approximation of Levenshtein distance, a widely-used metric for evaluating sequence similarity, has attracted significant attention with the emergence of DNA storage and other biological applications. Sequence embedding, which maps Levenshtein distance to a conventional distance between embedding vectors, has emerged as a promising solution. In this paper, a novel neural network-based sequence embedding technique using Poisson regression is proposed. We first provide a theoretical analysis of the impact of embedding dimension on model performance and present a criterion for selecting an appropriate embedding dimension. Under this embedding dimension, the Poisson regression is introduced by assuming the Levenshtein distance between sequences of fixed length following a Poisson distribution, which naturally aligns with the definition of Levenshtein distance. Moreover, from the perspective of the distribution of embedding distances, Poisson regression approximates the negative log likelihood of the chi-squared distribution and offers advancements in removing the skewness. Through comprehensive experiments on real DNA storage data, we demonstrate the superior performance of the proposed method compared to state-of-the-art approaches.

NeurIPS Conference 2024 Conference Paper

Rethinking Imbalance in Image Super-Resolution for Efficient Inference

  • Wei Yu
  • Bowen Yang
  • Qinglin Liu
  • Jianing Li
  • Shengping Zhang
  • Xiangyang Ji

Existing super-resolution (SR) methods optimize all model weights equally using $\mathcal{L}_1$ or $\mathcal{L}_2$ losses by uniformly sampling image patches without considering dataset imbalances or parameter redundancy, which limits their performance. To address this, we formulate the image SR task as an imbalanced distribution transfer learning problem from a statistical probability perspective, proposing a plug-and-play Weight-Balancing framework (WBSR) to achieve balanced model learning without changing the original model structure and training data. Specifically, we develop a Hierarchical Equalization Sampling (HES) strategy to address data distribution imbalances, enabling better feature representation from texture-rich samples. To tackle model optimization imbalances, we propose a Balanced Diversity Loss (BDLoss) function, focusing on learning texture regions while disregarding redundant computations in smooth regions. After joint training of HES and BDLoss to rectify these imbalances, we present a gradient projection dynamic inference strategy to facilitate accurate and efficient inference. Extensive experiments across various models, datasets, and scale factors demonstrate that our method achieves comparable or superior performance to existing approaches with about 34\% reduction in computational cost.

IROS Conference 2024 Conference Paper

Theoretical Modeling and Bio-inspired Trajectory Optimization of A Multiple-locomotion Origami Robot

  • Keqi Zhu
  • Haotian Guo
  • Wei Yu
  • Hassen Nigatu
  • Tong Li
  • Ruihong Dong
  • Huixu Dong

Recent research on mobile robots has focused on increasing their adaptability to unpredictable and unstructured environments using soft materials and structures. However, the determination of key design parameters and control over these compliant robots are predominantly iterated through experiments, lacking a solid theoretical foundation. To improve their efficiency, this paper aims to provide mathematics modeling over two locomotion, crawling and swimming. Specifically, a dynamic model is first devised to reveal the influence of the contact surfaces’ frictional coefficients on displacements in different motion phases. Besides, a swimming kinematics model is provided using coordinate transformation, based on which, we further develop an algorithm that systematically plans human-like swimming gaits, with maximum thrust obtained. The proposed algorithm is highly generalizable and has the potential to be applied in other soft robots with similar multiple joints. Simulation experiments have been conducted to illustrate the effectiveness of the proposed modeling.

JBHI Journal 2023 Journal Article

Coupled Contour Regression for Efficient Delineation of Lumen and External Elastic Lamina in Intravascular Ultrasound Images

  • Yuan Yang
  • Wei Yu
  • Haiyan Du
  • Li Ling
  • Qianjin Feng
  • Shengxian Tu
  • Wei Yang

Automatic delineation of the lumen and vessel contours in intravascular ultrasound (IVUS) images is crucial for the subsequent IVUS-based analysis. Existing methods usually address this task through mask-based segmentation, which cannot effectively handle the anatomical plausibility of the lumen and external elastic lamina (EEL) contours and thus limits their performance. In this article, we propose a contour encoding based method called coupled contour regression network (CCRNet) to directly predict the lumen and EEL contour pairs. The lumen and EEL contours are resampled, coupled, and embedded into a low-dimensional space to learn a compact contour representation. Then, we employ a convolutional network backbone to predict the coupled contour signatures and reconstruct the signatures to the object contours by a linear decoder. Assisted by the implicit anatomical prior of the paired lumen and EEL contours in the signature space and contour decoder, CCRNet has the potential to avoid producing unreasonable results. We evaluated our proposed method on a large IVUS dataset consisting of 7204 cross-sectional frames from 185 pullbacks. The CCRNet can rapidly extract the contours at 100 fps. Without any post-processing, all produced contours are anatomically reasonable in the test 19 pullbacks. The mean Dice similarity coefficients of our CCRNet for the lumen and EEL are 0. 940 and 0. 958, which are comparable to the mask-based models. In terms of the contour metric Hausdorff distance, our CCRNet achieves 0. 258 mm for lumen and 0. 268 mm for EEL, which outperforms the mask-based models.

NeurIPS Conference 2023 Conference Paper

On the choice of Perception Loss Function for Learned Video Compression

  • Sadaf Salehkalaibar
  • Truong Buu Phan
  • Jun Chen
  • Wei Yu
  • Ashish Khisti

We study causal, low-latency, sequential video compression when the output is subjected to both a mean squared-error (MSE) distortion loss as well as a perception loss to target realism. Motivated by prior approaches, we consider two different perception loss functions (PLFs). The first, PLF-JD, considers the joint distribution (JD) of all the video frames up to the current one, while the second metric, PLF-FMD, considers the framewise marginal distributions (FMD) between the source and reconstruction. Using information theoretic analysis and deep-learning based experiments, we demonstrate that the choice of PLF can have a significant effect on the reconstruction, especially at low-bit rates. In particular, while the reconstruction based on PLF-JD can better preserve the temporal correlation across frames, it also imposes a significant penalty in distortion compared to PLF-FMD and further makes it more difficult to recover from errors made in the earlier output frames. Although the choice of PLF decisively affects reconstruction quality, we also demonstrate that it may not be essential to commit to a particular PLF during encoding and the choice of PLF can be delegated to the decoder. In particular, encoded representations generated by training a system to minimize the MSE (without requiring either PLF) can be {\em near universal} and can generate close to optimal reconstructions for either choice of PLF at the decoder. We validate our results using (one-shot) information-theoretic analysis, detailed study of the rate-distortion-perception tradeoff of the Gauss-Markov source model as well as deep-learning based experiments on moving MNIST and KTH datasets.

NeurIPS Conference 2023 Conference Paper

Parameter and Computation Efficient Transfer Learning for Vision-Language Pre-trained Models

  • Qiong Wu
  • Wei Yu
  • Yiyi Zhou
  • Shubin Huang
  • Xiaoshuai Sun
  • Rongrong Ji

With ever increasing parameters and computation, vision-language pre-trained (VLP) models exhibit prohibitive expenditure in downstream task adaption. Recent endeavors mainly focus on parameter efficient transfer learning (PETL) for VLP models by only updating a small number of parameters. However, excessive computational overhead still plagues the application of VLPs. In this paper, we aim at parameter and computation efficient transfer learning (PCETL) for VLP models. In particular, PCETL not only needs to limit the number of trainable parameters in VLP models, but also to reduce the computational redundancy during inference, thus enabling a more efficient transfer. To approach this target, we propose a novel dynamic architecture skipping (DAS) approach towards effective PCETL. Instead of directly optimizing the intrinsic architectures of VLP models, DAS first observes the significances of their modules to downstream tasks via a reinforcement learning (RL) based process, and then skips the redundant ones with lightweight networks, i. e. adapters, according to the obtained rewards. In this case, the VLP model can well maintain the scale of trainable parameters while speeding up its inference on downstream tasks. To validate DAS, we apply it to two representative VLP models, namely ViLT and METER, and conduct extensive experiments on a bunch of VL tasks. The experimental results not only show the great advantages of DAS in reducing computational complexity, e. g. -11. 97% FLOPs of METER on VQA2. 0, but also confirm its competitiveness against existing PETL methods in terms of parameter scale and performance. Our source code is given in our appendix.

TCS Journal 2022 Journal Article

Approximation algorithms for the min-max clustered k-traveling salesmen problems

  • Xiaoguang Bao
  • Lei Xu
  • Wei Yu
  • Wei Song

Given a complete undirected graph G = ( V, E ), where V is the vertex set partitioned into K clusters V 1, V 2, …, V K and E is the edge set with edge weights satisfying triangle inequality, and a positive integer k, the min-max clustered k-traveling salesmen problem (min-max Ck-TSP) asks to find a set of k tours to visit all vertices, such that each cluster is visited by exactly one tour and the vertices of each cluster are visited consecutively. The objective is to minimize the weight of the maximum weight tour. The problem is known to be NP-hard even when k = 1 and K = 1. In this paper, we consider two variants of the problem. The first one is all the k tours have a common predefined starting vertex, and the other one is no starting vertex of any tour is specified. For both the variants we propose the first constant-factor approximation algorithms with ratios 5. 5 and 16, respectively.

JBHI Journal 2022 Journal Article

Time-Frequency Analysis of Scalp EEG With Hilbert-Huang Transform and Deep Learning

  • Jingyi Zheng
  • Mingli Liang
  • Sujata Sinha
  • Linqiang Ge
  • Wei Yu
  • Arne Ekstrom
  • Fushing Hsieh

Electroencephalography (EEG) is a brain imaging approach that has been widely used in neuroscience and clinical settings. The conventional EEG analyses usually require pre-defined frequency bands when characterizing neural oscillations and extracting features for classifying EEG signals. However, neural responses are naturally heterogeneous by showing variations in frequency bands of brainwaves and peak frequencies of oscillatory modes across individuals. Fail to account for such variations might result in information loss and classifiers with low accuracy but high variation across individuals. To address these issues, we present a systematic time-frequency analysis approach for analyzing scalp EEG signals. In particular, we propose a data-driven method to compute the subject-specific frequency bands for brain oscillations via Hilbert-Huang Transform, lifting the restriction of using fixed frequency bands for all subjects. Then, we propose two novel metrics to quantify the power and frequency aspects of brainwaves represented by sub-signals decomposed from the EEG signals. The effectiveness of the proposed metrics are tested on two scalp EEG datasets and compared with four commonly used features sets extracted from wavelet and Hilbert-Huang Transform. The validation results show that the proposed metrics are more discriminatory than other features leading to accuracies in the range of 94. 93% to 99. 84%. Besides classification, the proposed metrics show great potential in quantification of neural oscillations and serving as biomarkers in the neuroscience research.

TCS Journal 2020 Journal Article

New LP relaxations for Minimum Cycle/Path/Tree Cover Problems

  • Wei Yu
  • Zhaohui Liu
  • Xiaoguang Bao

Given an undirected complete graph G = ( V, E ) with nonnegative edge weight function obeying the triangle inequality, a set { C 1, C 2, …, C k } of cycles is called a cycle cover if V ⊆ ⋃ i = 1 k V ( C i ), where V ( C i ) represents the set of vertices in C i, and its cost is given by the maximum weight of the cycles. The Minimum Cycle Cover Problem (MCCP) aims to find a cycle cover of cost at most λ with the minimum number of cycles. We propose new LP relaxations for MCCP as well as its variants, called the Minimum Path Cover Problem (MPCP) and the Minimum Tree Cover Problem, where the cycles are replaced by paths or trees. Moreover, we give new LP relaxations for a special case of the rooted version of MCCP/MPCP. We show that these LP relaxations have significantly better integrality gaps than the previous relaxations.

IJCAI Conference 2019 Conference Paper

Learning a Generative Model for Fusing Infrared and Visible Images via Conditional Generative Adversarial Network with Dual Discriminators

  • Han Xu
  • Pengwei Liang
  • Wei Yu
  • Junjun Jiang
  • Jiayi Ma

In this paper, we propose a new end-to-end model, called dual-discriminator conditional generative adversarial network (DDcGAN), for fusing infrared and visible images of different resolutions. Unlike the pixel-level methods and existing deep learning-based methods, the fusion task is accomplished through the adversarial process between a generator and two discriminators, in addition to the specially designed content loss. The generator is trained to generate real-like fused images to fool discriminators. The two discriminators are trained to calculate the JS divergence between the probability distribution of downsampled fused images and infrared images, and the JS divergence between the probability distribution of gradients of fused images and gradients of visible images, respectively. Thus, the fused images can compensate for the features that are not constrained by the single content loss. Consequently, the prominence of thermal targets in the infrared image and the texture details in the visible image can be preserved or even enhanced in the fused image simultaneously. Moreover, by constraining and distinguishing between the downsampled fused image and the low-resolution infrared image, DDcGAN can be preferably applied to the fusion of different resolution images. Qualitative and quantitative experiments on publicly available datasets demonstrate the superiority of our method over the state-of-the-art.

TCS Journal 2019 Journal Article

New approximation algorithms for the minimum cycle cover problem

  • Wei Yu
  • Zhaohui Liu
  • Xiaoguang Bao

Given an undirected weighted graph G = ( V, E ) with nonnegative weight function obeying the triangle inequality, a set { C 1, C 2, …, C k } of cycles is called a cycle cover if V ⊆ ⋃ i = 1 k V ( C i ) and its cost is given by the maximum weight of the cycles. The Minimum Cycle Cover Problem aims to find a cycle cover of cost at most λ with the minimum number of cycles. An O ( n 2 ) 24/5-approximation algorithm and an O ( n 5 ) 14/3-approximation algorithm are given by Yu and Liu (Improved approximation algorithms for some min-max cycle cover problems. Theoretical Computer Science 654 (2016) 45–58). However, the original proofs for approximation ratios are incomplete. In this paper we first present a corrected simplified analysis on the 24/5-approximation algorithm. Based on the simplified approach of analysis and some new observations, we present a new 14/3-approximation algorithm that runs in O ( n 3 ) and give an improved 32/7-approximation algorithm that runs in O ( n 5 ).

TCS Journal 2016 Journal Article

Improved approximation algorithms for some min-max and minimum cycle cover problems

  • Wei Yu
  • Zhaohui Liu

Given an undirected weighted graph G = ( V, E ), a set { C 1, C 2, …, C k } of cycles is called a cycle cover of the vertex subset V ′ if V ′ ⊆ ∪ i = 1 k V ( C i ) and its cost is given by the maximum weight of the cycles. The Min-Max Cycle Cover Problem (MMCCP) is to find a minimum cost cycle cover of V with at most k cycles. The Rooted Min-Max Cycle Cover Problem (RMMCCP) is to find a minimum cost cycle cover of V ∖ D with at most k cycles, each of which contains one vertex in D. The Minimum Cycle Cover Problem (MCCP) aims to find a cycle cover of V of cost at most λ with the minimum number of cycles. We propose approximation algorithms for MMCCP and RMMCCP with performance ratios 5 and 6, respectively. These results improve the previous algorithms in term of both approximation ratios and running times. For MCCP we obtain a 14 3 -approximation algorithm that has the same time complexity as the previous best 5-approximation algorithm. Moreover, we transform a ρ-approximation algorithm for TSP into approximation algorithms for MMCCP, RMMCCP and MCCP with ratios 4ρ, 4 ρ + 1 and 4ρ, respectively.

TCS Journal 2013 Journal Article

Strategy-proof approximation mechanisms for an obnoxious facility game on networks

  • Yukun Cheng
  • Wei Yu
  • Guochuan Zhang

We study a new facility game, namely, an obnoxious facility game, on a network where the facility is undesirable and all agents try to be as far away from the facility as possible. The following process is considered: at first the agents declare their locations, then, given these bids, a mechanism selects a place on the network to locate the facility. The aim of the mechanism is to maximize the obnoxious social welfare, i. e. , the total distance between the agents and the facility. The objective of each agent is to maximize his/her utility, i. e. , the distance from the facility. Thus an agent may lie if, by doing so, he/she can get strictly more benefit. We are interested in mechanisms without money to decide the facility location so that the obnoxious social welfare is maximized and all agents are enforced to report their true locations. In this paper we give a first attempt at this game on different networks. Our main results are the following. When the network is a path, we show a 3-approximation group strategy-proof deterministic mechanism which is best possible if the facility can only take one of the endpoints on the path, and a group strategy-proof randomized mechanism with tight approximation ratio of 3 2. When the networks are a circle (known as a ring in the case of computer networks) and a tree, we propose two group strategy-proof deterministic mechanisms that each provides the approximation ratio of 3. Furthermore, when all agents are on a general network, we propose a 4-approximation group strategy-proof deterministic mechanism and a 2-approximation group strategy-proof randomized mechanism.

ICRA Conference 2011 Conference Paper

Motion planning for steep hill climbing

  • Damion D. Dunlap
  • Wei Yu
  • Emmanuel G. Collins Jr.
  • Charmane V. Caldwell

The motors or engines of an autonomous ground vehicles (AGV) have torque and power limitations, which limit their abilities to climb steep hills, which are defined to be hills that have high grade sections in which the vehicle is forced to decelerate. Traversal of a steep hill requires the vehicle to have sufficient momentum before entering the hill. This problem is part of a larger class of momentum-based motion planning problems such as the problem of lifting heavy objects with manipulators. Hence, solutions to the steep hill climbing problem have much wider applicability. The motion planning here is accomplished using a dynamic model of the skid-steered AGV used in the experiments along with Sampling Based Model Predictive Control (SBMPC), a recently developed input sampling planning algorithm that may be viewed as a generalization of LPA* to the direct use of kinodynamic models. The motion planning is demonstrated experimentally using two scenarios, one in which the robot starts at rest at the bottom of a hill and one in which the robot starts at rest a distance from the hill. The first scenario requires the AGV to first reverse direction so that the vehicle can gather enough momentum before reaching the hill. This corresponds to having the vehicle begin at a local minimum, which results in a problem that many traditional model predictive control methods cannot solve. It is seen that, whereas open loop trajectories can lead to vehicle immobilization, SBMPC successfully uses the information provided by the dynamic model to ensure that the AGV has the requisite momentum.

IROS Conference 2009 Conference Paper

Dynamic modeling of a skid-steered wheeled vehicle with experimental verification

  • Wei Yu
  • Oscar Chuy
  • Emmanuel G. Collins Jr.
  • Patrick Hollis

Skid-steered vehicles are often used as outdoor mobile robots due to their robust mechanical structure and high maneuverability. Sliding along with rolling is inherent to general curvilinear motion, which makes both kinematic and dynamic modeling difficult. For the purpose of motion planning this paper develops and experimentally verifies dynamic models of a skid-steered wheeled vehicle for general planar (2D) motion and for linear 3D motion. These models are characterized by the coefficient of rolling resistance, the coefficient of friction, and the shear deformation modulus, which have terrain-dependent values. The dynamic models also include motor saturation and motor power limitations, which enable correct prediction of vehicle velocities when traversing hills. It is shown that the closed-loop system that results from inclusion of the dynamics of the (PID) speed controllers for each set of wheels does a much better job than the open loop model of predicting the vehicle linear and angular velocities. Hence, the closed-loop model is recommended for motion planning.

ICRA Conference 2009 Conference Paper

Power modeling of a skid steered wheeled robotic ground vehicle

  • Oscar Chuy
  • Emmanuel G. Collins Jr.
  • Wei Yu
  • Camilo Ordonez

Analysis of the power consumption of a robotic ground vehicle (RGV) is important for planning since it enables motion plans that do not violate the power limitations of the motors, energy efficient path planning, prediction of the ability to complete a task based upon the vehicle's current energy supply, and estimation of when the vehicle will need to refuel or recharge. Power modeling is particularly difficult for skid steered vehicles because of the complexities of properly taking into account the skidding that is used for vehicle turning. This paper begins with a 2-dimensional, second order differential equation of a skid steered, wheeled RGV and shows that the power model is terrain dependent and is a function of both the turning radius and linear velocity of the vehicle. This model was verified experimentally, and a comprehensive set of experiments was performed to describe the power consumption of a skid steered RGV on asphalt.

YNIMG Journal 2004 Journal Article

Semantic processing of Chinese in left inferior prefrontal cortex studied with reversible words

  • John X. Zhang
  • Jie Zhuang
  • Lifei Ma
  • Wei Yu
  • Danling Peng
  • Guosheng Ding
  • Zhaoqi Zhang
  • Xuchu Weng

This study utilized fast event-related fMRI with reversible words to examine the role of left inferior prefrontal cortex (PFC) in semantic processing of Chinese. As a special linguistic phenomenon in Chinese, a reversible word is a two-character word (AB) that, when read from right to left (BA), opposite to the normal left to right reading direction, is also a real word. The two words, AB and BA, can have very different meanings. Fourteen native Chinese saw a reversible word (BA) and were asked to read it backward silently to obtain the meaning of AB, defined as the target meaning. They then saw two test words and decided which of the two was semantically related to the target meaning. Activity in a subregion of BA47 was found to be modulated by the extent to which irrelevant semantic activation of the distractor word BA interfered with semantic retrieval of the target word AB. This finding demonstrated the involvement of the left inferior PFC in the control processes of semantic retrieval in Chinese. In addition, comparing conditions using reversible with that using nonreversible words, we found evidence suggesting a semantic/phonological functional subdivision in left inferior PFC, consistent with that in English.

v2026.09.13