Arrow Research search

Author name cluster

Wensheng Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
2 author rows

Possible papers

12

AAAI Conference 2026 Conference Paper

Is Your (Reasoning) Multimodal Language Model Vulnerable Toward Distractions?

  • Ming Liu
  • Hao Chen
  • Jindong Wang
  • Liwen Wang
  • Jingchen Sun
  • Wensheng Zhang

Vision-Language Models (VLMs) have achieved success in tasks such as visual question answering, yet their resilience to distractions remains underexplored. Understanding how distractions affect VLMs' performance is crucial for real-world applications, as input data often contains noisy or irrelevant content. This paper assesses the robustness of VLMs—including general-purpose models and those specialized for reasoning—against distractions in the context of science question answering. We introduce I-ScienceQA, a new benchmark based on the ScienceQA dataset, which systematically injects distractions into both visual and textual contexts. We evaluate how distractions perturb the underlying reasoning processes of these models by analyzing changes in textual explanations leading to answers. Our findings show that most VLMs are vulnerable to distractions, with a noticeable degradation in reasoning when extraneous content is present. In particular, some models (including GPT-o4 mini) exhibit a higher degree of robustness. We also observe that textual distractions generally cause greater performance declines than visual distractions. Finally, we explore mitigation strategies such as prompt engineering. Although these strategies improve resilience modestly, our analysis highlights considerable room for further improvement in the robustness of VLMs.

JBHI Journal 2026 Journal Article

MuGEP: Multiplex Graph-Based Brain Network Modeling for Epileptic Seizure Prediction Using Intracranial EEG

  • Minyu Zhou
  • Yajing Wu
  • Yongqiang Tang
  • Xiaohu Zhou
  • Ying Zhang
  • Runshi Gao
  • Wang Jia
  • Wensheng Zhang

Accurate seizure prediction in advance is crucial for patients with epilepsy, as it helps prevent harm and improve life quality. Intracranial electroencephalogram (iEEG), enabling precise characterization of epileptogenic and propagation networks from deep brain tissue, is a reliable foundation for epilepsy research. While deep learning models have shown success in automating seizure prediction, existing methods often overlook diverse brain network relationships and the rich information within each channel, thereby failing to fully exploit the advantages of iEEG signals. To address these issues, we propose a Multiplex Graph-based brain network modeling framework for Epileptic seizure Prediction (MuGEP) to represent the diverse and fine-grained relationships in the brain network effectively, involving the relationships between amplitude and phase of different frequency bands. Specifically, a specialized multiplex graph called the Multiplex Brain Graph (MBG) is designed for brain network connectivity, which is inspired by Cross-Frequency Coupling (CFC) in neuroscience. In MBG, nodes represent the frequency bands in each channel, and edges are computed following the idea of three types of CFC, resulting in three distinct subgraphs. Further, a novel MBG learning network is proposed to incorporate intra- and inter-subgraph patterns to obtain the final representation leveraging graph convolution networks and a joint fusion module. A sufficient evaluation is conducted on the Kaggle and the SWEC-ETHZ datasets, and the promising results confirm the advantage of MuGEP for iEEG seizure prediction.

JBHI Journal 2025 Journal Article

Contrastive Learning With Transformer to Predict the Chronicity of Children With Immune Thrombocytopenia

  • Yuntian Wang
  • Yongqiang Tang
  • Jingyao Ma
  • Zhenping Chen
  • Chang Cui
  • Mingda Li
  • Runhui Wu
  • Wensheng Zhang

Immune thrombocytopenia (ITP) is a typically self-limiting and immune-mediated bleeding disorder in children. Approximately 20% of children with ITP experience chronicity, leading to reduced quality of life and increased treatment burden. The accurate prediction of chronicity would enable clinicians to make personalized treatment plans at an early stage. However, due to the self-limiting nature of ITP and the scarcity of available children patients, the data presents two prominent issues: small data and imbalanced class, which are unfavorable for effectively training a deep learning model. To handle these issues concurrently, we proposed a novel method that integrates contrastive learning with the Transformer. First, we adopt the FT-Transformer as our backbone, which allows our model to flexibly process heterogeneous tabular data. Second, we amplify and balance the original data via random masking and oversampling, respectively. Lastly, we build contrastive pairs according to the latent representations generated by the FT-Transformer encoder, such that the amplified and oversampled synthetic data can be utilized thoroughly. The experimental results on real-world ITP children data show that our proposal outperforms the state-of-the-art methods, and demonstrate the significant advantages of dealing with insufficient and imbalanced problems.

ICLR Conference 2025 Conference Paper

Is Your Video Language Model a Reliable Judge?

  • Ming Liu
  • Wensheng Zhang

As video language models (VLMs) gain more applications in various scenarios, the need for robust and scalable evaluation of their performance becomes increasingly critical. The traditional human expert-based evaluation of VLMs has limitations in consistency and scalability, which sparked interest in automatic methods such as employing VLMs to evaluate VLMs. However, the reliability of VLMs as judges remains underexplored. Existing methods often rely on a single VLM as the evaluator. However, this approach can be unreliable or biased because such a model may lack the ability to fully understand the content and may have inherent biases, ultimately compromising evaluation reliability. A remedy is to apply the principle of collective thoughts, aggregating evaluations from multiple VLMs to enhance reliability. This study investigates the efficacy of such approaches, particularly when the pool of judges includes both reliable and unreliable models. Our findings reveal that incorporating collective judgments from such a mixed pool does not necessarily improve the accuracy of the final evaluation. The inclusion of less reliable judges can introduce noise, undermining the overall reliability of the outcomes. To explore the factors that impact evaluation reliability, we fine-tune an underperforming VLM judge, Video-LLaVA, and observe that improved understanding ability alone is insufficient to make VLM judges more reliable. These findings stress the limitations of collective thought approaches and highlight the need for more advanced methods that can account for the reliability of individual models. Our study promotes the development of more reliable evaluation methods for VLMs

NeurIPS Conference 2025 Conference Paper

On Fairness of Unified Multimodal Large Language Model for Image Generation

  • Ming Liu
  • Hao Chen
  • Jindong Wang
  • Liwen Wang
  • Bhiksha Raj
  • Wensheng Zhang

Unified multimodal large language models (U-MLLMs) have demonstrated impressive performance in end-to-end visual understanding and generation tasks. However, compared to generation-only systems (e. g. , Stable Diffusion), the unified architecture of U-MLLMs introduces new risks of propagating demographic stereotypes. In this paper, we benchmark several state-of-the-art U-MLLMs and show that they exhibit significant gender and race biases in the generated outputs. To diagnose the source of these biases, we propose a locate-then-fix framework: we first audit the vision and language components — using techniques such as linear probing and controlled generation — and find that the language model appears to be a primary origin of the observed generative bias. Moreover, we observe a ``partial alignment'' phenomenon, where the U-MLLMs exhibit less bias in understanding tasks yet produce substantially biased images. To address this, we introduce a novel \emph{balanced preference loss} that enforces uniform generation probabilities across demographics by leveraging a synthetically balanced dataset. Extensive experiments show that our approach significantly reduces demographic bias while preserving semantic fidelity and image quality. Our findings underscore the need for targeted debiasing strategies in unified multimodal systems and introduce a practical approach to mitigate biases.

JBHI Journal 2025 Journal Article

RTGN: Robust Traditional Chinese Medicine Graph Networks for Patient Similarity Learning

  • Junjie Long
  • Jinghao Niu
  • Heping Wang
  • Jiaxi Liu
  • Jie Li
  • Wensheng Zhang

Traditional Chinese Medicine (TCM) boasts a long history and a unique diagnostic and therapeutic paradigm. Integrating TCM with Western medicine and modern medical devices has yielded numerous successful cases in recent years. TCM treatment has developed a special knowledge framework focusing on precise differentiation based on multidimensional information such as the patient's diseases, symptoms, and syndromes. This offers significant opportunities for AI research in similar patient scenarios within TCM contexts. However, traditional medicine's reliance on the physiological sensory judgment of human physicians to gather clinical information might lead to non-standardized descriptions and disturbances in patient assessments. Additionally, how to integrate TCM's fine-grained differentiation knowledge to design a patient similarity measure remains an open question. To address this, we first constructed a real-world dataset of TCM gastrointestinal malignancies (TCMGI) based on real cases in the Guang'anmen Hospital, China Academy of Chinese Medical Sciences. It contains 406 types of multidimensional information from 719 patients, organized in a graph structure. Second, we develop a novel deep learning framework, Robust Traditional Chinese Medicine Graph Networks (RTGN), which employs a Siamese network architecture with self-attention and self-supervision strategies to enhance robustness in patient retrieval. Lastly, we design a patient similarity metric integrating TCM and Western medicine approaches, demonstrating superior performance in depicting fine-grained patient similarities. Experimental results show our method outperforms existing best practices in patient retrieval accuracy. Moreover, the proposed similarity metric exhibits excellent performance in clustering tasks at various granularity levels, possibly supporting precision TCM patient retrieval and downstream tasks, such as prescription generation.

NeurIPS Conference 2024 Conference Paper

Addressing Hidden Confounding with Heterogeneous Observational Datasets for Recommendation

  • Yanghao Xiao
  • Haoxuan Li
  • Yongqiang Tang
  • Wensheng Zhang

The collected data in recommender systems generally suffers selection bias. Considerable works are proposed to address selection bias induced by observed user and item features, but they fail when hidden features (e. g. , user age or salary) that affect both user selection mechanism and feedback exist, which is called hidden confounding. To tackle this issue, methods based on sensitivity analysis and leveraging a few randomized controlled trial (RCT) data for model calibration are proposed. However, the former relies on strong assumptions of hidden confounding strength, whereas the latter relies on the expensive RCT data, thereby limiting their applicability in real-world scenarios. In this paper, we propose to employ heterogeneous observational data to address hidden confounding, wherein some data is subject to hidden confounding while the remaining is not. We argue that such setup is more aligned with practical scenarios, especially when some users do not have complete personal information (thus assumed with hidden confounding), while others do have (thus assumed without hidden confounding). To achieve unbiased learning, we propose a novel meta-learning based debiasing method called MetaDebias. This method explicitly models oracle error imputation and hidden confounding bias, and utilizes bi-level optimization for model training. Extensive experiments on three public datasets validate our method achieves state-of-the-art performance in the presence of hidden confounding, regardless of RCT data availability.

JBHI Journal 2024 Journal Article

Deep Survival Analysis With Latent Clustering and Contrastive Learning

  • Chang Cui
  • Yongqiang Tang
  • Wensheng Zhang

Survival analysis is employed to analyze the time before the event of interest occurs, which is broadly applied in many fields. The existence of censored data with incomplete supervision information about survival outcomes is one key challenge in survival analysis tasks. Although some progress has been made on this issue recently, the present methods generally treat the instances as separate ones while ignoring their potential correlations, thus rendering unsatisfactory performance. In this study, we propose a novel Deep Survival Analysis model with latent Clustering and Contrastive learning (DSACC). Specifically, we jointly optimize representation learning, latent clustering and survival prediction in a unified framework. In this way, the clusters distribution structure in latent representation space is revealed, and meanwhile the structure of the clusters is well incorporated to improve the ability of survival prediction. Besides, by virtue of the learned clusters, we further propose a contrastive loss function, where the uncensored data in each cluster are set as anchors, and the censored data are treated as positive/negative sample pairs according to whether they belong to the same cluster or not. This design enables the censored data to make full use of the supervision information of the uncensored samples. Through extensive experiments on four popular clinical datasets, we demonstrate that our proposed DSACC achieves advanced performance in terms of both C-index (0. 6722, 0. 6793, 0. 6350, and 0. 7943) and Integrated Brier Score (IBS) (0. 1616, 0. 1826, 0. 2028, and 0. 1120).

AIIM Journal 2023 Journal Article

CEHMR: Curriculum learning enhanced hierarchical multi-label classification for medication recommendation

  • Mengxuan Sun
  • Jinghao Niu
  • Xuebing Yang
  • Yifan Gu
  • Wensheng Zhang

The medication recommendation (MR) or medication combination prediction task aims to predict effective prescriptions given accurate patient representations derived from electronic health records (EHRs), which contributes to improving the quality of clinical decision-making, especially for patients with multi-morbidity. Although in recent years deep learning technology has achieved great success in MR, the performance of current multi-label based MR solutions is unsatisfactory. They mainly focus on improving the patient representation module and modeling the medication label dependencies such as drug–drug interaction (DDI) correlation and co-occurrence relationship. However, the hierarchical dependency among medication labels and diversity of difficulty among MR training examples lack sufficient consideration. In this paper, we propose a framework of Curriculum learning Enhanced Hierarchical multi-label classification for MR (CEHMR). Motivated by the category hierarchy of medications which organizes standard medication codes in a hierarchical structure, we utilize it to provide more trustworthy prior knowledge for modeling label dependency. Specifically, we design a hierarchical multi-label classifier with a learnable gate fusion layer, to simultaneously capture the level-independent (local) and level-dependent (global) hierarchical information in the medication hierarchy. In addition, to overcome the diversity of training example difficulties, and progressively achieve a smoother training process, we introduce a bootstrap-based curriculum learning strategy. Hence, the example difficulty can be measured based on the predictive performance of the MR model, and then all training examples would be retrained from easy to hard under the guidance of a predefined training scheduler. Experiments on the real-world medical MIMIC-III database demonstrate that the proposed framework can achieve state-of-the-art performance compared with seven representative baselines, and extensive ablation studies validate the effectiveness of each component of CEHMR.

JBHI Journal 2023 Journal Article

Leveraging Summary Guidance on Medical Report Summarization

  • Yunqi Zhu
  • Xuebing Yang
  • Yuanyuan Wu
  • Wensheng Zhang

This study presents three deidentified large medical text datasets, named DISCHARGE, ECHO and RADIOLOGY, which contain 50 K, 16 K and 378 K pairs of report and summary that are derived from MIMIC-III, respectively. We implement convincing baselines of automated abstractive summarization on the created datasets with pre-trained encoder-decoder language models, including BERT2BERT, BERTShare, RoBERTaShare, Pegasus, ProphetNet, T5-large, BART and GSUM. Further, based on the BART model, we leverage the sampled summaries from the training set as prior knowledge guidance, for encoding additional contextual representations of the guidance with the encoder and enhancing the decoding representations in the decoder. The experimental results confirm the improvement of ROUGE scores and BERTScore made by the proposed method.

JBHI Journal 2020 Journal Article

Inter-Patient ECG Classification With Symbolic Representations and Multi-Perspective Convolutional Neural Networks

  • Jinghao Niu
  • Yongqiang Tang
  • Zhengya Sun
  • Wensheng Zhang

This paper presents a novel deep learning framework for the inter-patient electrocardiogram (ECG) heartbeat classification. A symbolization approach especially designed for ECG is introduced, which can jointly represent the morphology and rhythm of the heartbeat and alleviate the influence of inter-patient variation through baseline correction. The symbolic representation of the heartbeat is used by a multi-perspective convolutional neural network (MPCNN) to learn features automatically and classify the heartbeat. We evaluate our method for the detection of the supraventricular ectopic beat (SVEB) and ventricular ectopic beat (VEB) on MIT-BIH arrhythmia dataset. Compared with the state-of-the-art methods based on manual features or deep learning models, our method shows superior performance: the overall accuracy of 96. 4%, F1 scores for SVEB and VEB of 76. 6% and 89. 7%, respectively. The ablation study on our method validates the effectiveness of the proposed symbolization approach and joint representation architecture, which can help the deep learning model to learn more general features and improve the ability of generalization for unseen patients. Because our method achieves a competitive inter-patient heartbeat classification performance without complex handcrafted features or the intervention of the human expert, it can also be adjusted to handle various other tasks relative to ECG classification.

IJCAI Conference 2018 Conference Paper

Finite Sample Analysis of LSTD with Random Projections and Eligibility Traces

  • Haifang Li
  • Yingce Xia
  • Wensheng Zhang

Policy evaluation with linear function approximation is an important problem in reinforcement learning. When facing high-dimensional feature spaces, such a problem becomes extremely hard considering the computation efficiency and quality of approximations. We propose a new algorithm, LSTD(lambda)-RP, which leverages random projection techniques and takes eligibility traces into consideration to tackle the above two challenges. We carry out theoretical analysis of LSTD(lambda)-RP, and provide meaningful upper bounds of the estimation error, approximation error and total generalization error. These results demonstrate that LSTD(lambda)-RP can benefit from random projection and eligibility traces strategies, and LSTD(lambda)-RP can achieve better performances than prior LSTD-RP and LSTD(lambda) algorithms.

v2026.09.13