Arrow Research search

Author name cluster

Zhiwei Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

27 papers
2 author rows

Possible papers

27

AAAI Conference 2026 Conference Paper

FIA-Edit: Frequency-Interactive Attention for Efficient and High-Fidelity Inversion-Free Text-Guided Image Editing

  • Kaixiang Yang
  • Boyang Shen
  • Xin Li
  • Yuchen Dai
  • Yuxuan Luo
  • Yueran Ma
  • Wei Fang
  • Qiang Li

Text-guided image editing has advanced rapidly with the rise of diffusion models. While flow-based inversion-free methods offer high efficiency by avoiding latent inversion, they often fail to effectively integrate source information, leading to poor background preservation, spatial inconsistencies, and over-editing due to the lack of effective integration of source information. In this paper, we present FIA-Edit, a novel inversion-free framework that achieves high-fidelity and semantically precise edits through a Frequency-Interactive Attention. Specifically, we design two key components: (1) a Frequency Representation Interaction (FRI) module that enhances cross-domain alignment by exchanging frequency components between source and target features within self-attention, and (2) a Feature Injection (FIJ) module that explicitly incorporates source-side queries, keys, values, and text embeddings into the target branch's cross-attention to preserve structure and semantics. Comprehensive and extensive experiments demonstrate that FIA-Edit supports high-fidelity editing at low computational cost (~6s per 512 * 512 image on an RTX 4090) and consistently outperforms existing methods across diverse tasks in visual quality, background fidelity, and controllability. Furthermore, we are the first to extend text-guided image editing to clinical applications. By synthesizing anatomically coherent hemorrhage variations in surgical images, FIA-Edit opens new opportunities for medical data augmentation and delivers significant gains in downstream bleeding classification.

EAAI Journal 2026 Journal Article

GRHP: Graph-Fused Hierarchical Planning for Embodied Long-Horizon Robotic Task

  • Xiaodong Li
  • Guohui Tian
  • Yongcheng Cui
  • Xuyang Shao
  • Zhiwei Wang

Embodied Long-Horizon task planning is crucial for robots to perform complex household tasks. Existing methods typically rely on Vision Language Models (VLMs) to generate high-level semantic planning, but face critical limitations. First, the plans often fail to be grounded in the physical environment due to an inability to perceive spatial layouts and object relationships. Second, the high-level plans lack direct executability, as they cannot be readily mapped to the robot’s atomic actions. To overcome these challenges, we present Graph-Fused Hierarchical Planning (GRHP), a novel framework for Long-Horizon task planning. GRHP employs a unified dual-graph perception structure, where a scene graph captures spatial context and relationships, and a task graph models action dependencies and intent. Environmental constraints are directly integrated into task planning through explicit cross-graph fusion. Furthermore, GRHP features a hierarchical planning architecture that strategically decouples the planning process. It leverages large models for efficient high-level semantic planning, while small models handle precise action generation. This decomposition ensures that high-level instruction translates directly into executable low-level actions. Extensive experiments on GATP (Graph-Aware Task Planning), an enhanced version of the challenging ALFRED benchmark, demonstrate the effectiveness of GRHP. Qualitative analysis corroborates the significant contributions of scene perception, task modeling, and the hierarchical design for agent performance. Our code is available at: https: //github. com/LL00qw/GRHP.

JBHI Journal 2026 Journal Article

Improving 3D Thin Vessel Segmentation in Brain TOF-MRA via a Dual-Space Context-Aware Network

  • Wenqi Shan
  • Xudong Li
  • Xuan Wang
  • Qiang Li
  • Zhiwei Wang

3D cerebrovascular segmentation poses a significant challenge, akin to locating a line within a vast 3D environment. This complexity can be substantially reduced by projecting the vessels onto a 2D plane, enabling easier segmentation. In this paper, we create a vessel-segmentation-friendly space using a clinical visualization technique called maximum intensity projection (MIP). Leveraging this, we propose a Dual-space Context-Aware Network (DCANet) for 3D vessel segmentation, designed to capture even the finest vessel structures accurately. DCANet begins by transforming a magnetic resonance angiography (MRA) volume into a 3D Regional-MIP volume, where each Regional-MIP slice is constructed by projecting adjacent MRA slices. This transformation highlights vessels as prominent continuous curves rather than the small circular or ellipsoidal cross-sections seen in MRA slices. DCANet encodes vessels separately in the MRA and the projected Regional-MIP spaces and introduces the Regional-MIP Image Fusion Block (MIFB) between these dual spaces to selectively integrate contextual features from Regional-MIP into MRA. Following dual-space encoding, DCANet employs a Dual-mask Spatial Guidance TransFormer (DSGFormer) decoder to focus on vessel regions while effectively excluding background areas, which reduces the learning burden and improves segmentation accuracy. We benchmark DCANet on four datasets: two public datasets, TubeTK and IXI-IOP, and two in-house datasets, Xiehe and IXI-HH. The results demonstrate that DCANet achieves superior performance, with improvements in average DSC values of at least 2. 26%, 2. 17%, 2. 62%, and 2. 58% for thin vessels, respectively.

AAAI Conference 2026 Conference Paper

Pairing-free Group-level Knowledge Distillation for Robust Gastrointestinal Lesion Classification in White-Light Endoscopy

  • Qiang Hu
  • Qimei Wang
  • Yingjie Guo
  • Qiang Li
  • Zhiwei Wang

White-Light Imaging (WLI) is the standard for endoscopic cancer screening, but Narrow-Band Imaging (NBI) offers superior diagnostic details. A key challenge is transferring knowledge from NBI to enhance WLI-only models, yet existing methods are critically hampered by their reliance on paired NBI-WLI images of the same lesion, a costly and often impractical requirement that leaves vast amounts of clinical data untapped. In this paper, we break this paradigm by introducing PaGKD, a novel Pairing-free Group-level Knowledge Distillation framework that that enables effective cross-modal learning using unpaired WLI and NBI data. Instead of forcing alignment between individual, often semantically mismatched image instances, PaGKD operates at the group level to distill more complete and compatible knowledge across modalities. Central to PaGKD are two complementary modules: (1) Group-level Prototype Distillation (GKD-Pro) distills compact group representations by extracting modality-invariant semantic prototypes via shared lesion-aware queries; (2) Group-level Dense Distillation (GKD-Den) performs dense cross-modal alignment by guiding group-aware attention with activation-derived relation maps. Together, these modules enforce global semantic consistency and local structural coherence without requiring image-level correspondence. Extensive experiments on four clinical datasets demonstrate that PaGKD consistently and significantly outperforms state-of-the-art methods, boosting AUC by 3.3%, 1.1%, 2.8%, and 3.2%, respectively, establishing a new direction for cross-modal learning from unpaired data.

AAAI Conference 2026 Conference Paper

RouterNet: Hierarchical Point Routing Network for Robust Vertebral Landmark Localization on AP X-ray Images

  • Yingjie Guo
  • Jinxin Lv
  • Wei Fang
  • Qiang Li
  • Zhiwei Wang

Locating vertebral landmarks on anteroposterior (AP) X-ray images is challenging due to the tissue overlap. Despite the great progress of heatmap-based methods, they often predict missing/false points, which are intolerable in the downstream applications like scoliosis assessment. In this paper, we instead modernize the classic point-regression scheme, and propose a novel model termed RouterNet to locate the 68 vertebral landmarks completely and accurately. RouterNet starts from an initial root point, and then gradually routes it onto more and more points with finer and finer semantics. RouterNet naturally couples such point routing process with its hierarchical and multi-scale feature learning. That is, lower-scale feature maps are utilized to regress points with coarser semantics, and the regressed points pilot a more focused local feature extraction on the next higher-scale map to route onto their subsequent positions with finer semantics. With this divide-and-conquer, RouterNet alleviates the task difficulty, and can robustly localize by routing from the whole spinal center to 17 vertebral centers, and further to their 68 corner points. Extensive and comprehensive experiments on both public and private datasets demonstrate our superior performance over other state-of-the-arts, by decreasing NMSE by 73.8% for landmark localization, and SMAPE by 14.8% for the downstream scoliosis assessment.

NeurIPS Conference 2025 Conference Paper

Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data

  • Tianyi Chen
  • Pengxiao Lin
  • Zhiwei Wang
  • Zhi-Qin Xu

State Space Models (SSMs) have emerged as promising alternatives to attention mechanisms, with the Mamba architecture demonstrating impressive performance and linear complexity for processing long sequences. However, the fundamental differences between Mamba and Transformer architectures remain incompletely understood. In this work, we use carefully designed synthetic tasks to reveal Mamba's inherent limitations. Through experiments, we identify that Mamba's nonlinear convolution introduces an asymmetry bias that significantly impairs its ability to recognize symmetrical patterns and relationships. Using composite function and inverse sequence matching tasks, we demonstrate that Mamba strongly favors compositional solutions over symmetrical ones and struggles with tasks requiring the matching of reversed sequences. We show these limitations stem not from the SSM module itself but from the nonlinear convolution preceding it, which fuses token information asymmetrically. These insights provide a new understanding of Mamba's constraints and suggest concrete architectural improvements for future sequence models.

JBHI Journal 2025 Journal Article

Boosting Few-Shot Semantic Segmentation of 3D Medical Images via Collaborative Slice Alignment

  • Ran Duan
  • Jialun Pei
  • Zhiwei Wang
  • Ruiheng Zhang
  • Qiang Li
  • Pheng-Ann Heng

Few-shot semantic segmentation (FSS) of 3D medical images requires finding a 2D slice from the labeled volume as support to ‘query’ slices of the unlabeled one. Accurately determining support slices is crucial for learning representative prototypical features, thereby enhancing segmentation accuracy. The existing methods typically resort to the true position of the query target to align the query with support slices or simply exploit one key support slice to segment all query slices, which inevitably results in poor practicality and mis-segmentation. In this regard, we seek a practical and efficient solution by proposing a novel Collaborative Slice Alignment (CSA) module, which densely assigns each query slice its own fittest support without knowing the target prior. Concretely, our CSA first estimates the confidence scores of slices from the sorting task to implicitly reflect their physical location in the human body. The estimated scores are considered as spatial references for aligning support slices and query slices so that each matching pair shares the most similar image contents. Moreover, the self-learnable ranking objective allows CSA to transfer internal knowledge into both support and query features to further boost the FSS performance. Additionally, we introduce an Information Reconciliation (InRe) module to mitigate the inconsistent feature distribution caused by the individual differences between support and query images. Experimental results demonstrate that the combination of CSA and InRe achieves an average Dice score improvement of at least 8. 61% across three datasets, consistently outperforming other state-of-the-art methods.

JBHI Journal 2025 Journal Article

Cross-Correlation Rectification for Robust Deformable Registration of Brain Tumor MRI Between Preoperative and Postoperative Phases

  • Chongwei Wu
  • Tao Gong
  • Xiaoyu Zeng
  • Shuxian Niu
  • Guangbin Wang
  • Qiang Li
  • Zhiwei Wang

A key challenge in registering pre- and post-operative brain tumor images lies in the anatomical inconsistencies caused by pathological changes and surgical resections. Recent efforts have addressed this issue by masking affected regions during optimization, but such approaches discard contextual information and rely on CNN backbones that implicitly model deformation, often overfitting to distant normal tissues and failing to capture the severe nonlinear distortions near the tumor. Correlation-based alternatives enhance generalization to diverse deformation patterns by explicitly modeling geometric correspondences, yet they frequently yield unreliable matches in and around tumor regions, disrupting the deformation field. In this paper, we propose Cross-correlation Rectification-based Registration Network (CRRNet), the first framework that introduces an active rectification mechanism specifically for robustness and structurally coherent pre- to post-operative brain tumor image registration. Specifically, CRRNet achieves this through two complementary modules: 1) a Cross-correlation Analysis-Based Inconsistency module that identifies invalid correspondences via bidirectional loop-closure evaluation on cross-correlations, and 2) a Dual-level Correspondence Rectification module that adaptively integrates contextually reliable correlations from local and long-range perspectives to restore structurally coherent matches. This synergistic design retains the strengths of cross-correlation while effectively mitigating correspondence mismatches. Extensive experiments on multiple tumor benchmarks demonstrate the superiority of CRRNet. Specifically, on the BraTS-Reg dataset, it reduces the mean registration errors by 8. 18% in near-tumor regions and 3. 95% in far-from-tumor regions, surpassing state-of-the-art methods.

JBHI Journal 2025 Journal Article

Fine-Grained Temporal Site Monitoring in EGD Streams via Visual Time-Aware Embedding and Vision-Text Asymmetric Coworking

  • Fang Peng
  • Hongkuan Shi
  • Shiquan He
  • Qiang Hu
  • Ting Li
  • Fan Huang
  • Xinxia Feng
  • Mei Liu

Esophagogastroduodenoscopy (EGD) requires inspecting plentiful upper gastrointestinal (UGI) sites completely for a precise cancer screening. Automated temporal site monitoring for EGD assistance is thus of high demand, yet often fails if directly applying the existing methods of online action detection. The key challenges are two-fold: 1) the global camera motion dominates, invalidating the temporal patterns derived from the object optical flows, and 2) the UGI sites are fine-grained, yielding highly homogenized appearances. In this paper, we propose an EGD-customized model, powered by two novel designs, i. e. , Visual Time-aware Embedding plus Vision-text Asymmetric Coworking (VTE+VAC), for real-time accurate fine-grained UGI site monitoring. Concretely, VTE learns visual embeddings by differentiating frames via classification losses, and meanwhile by reordering the sampled time-agnostic frames to be temporally coherent via a ranking loss. Such joint objective encourages VTE to capture the sequential relation without resorting to the inapplicable object optical flows, and thus to provide the time-aware frame-wise embeddings. In the subsequent analysis, VAC uses a temporal sliding window, and extracts vision-text multimodal knowledge from each frame and its corresponding textualized prediction via the learned VTE and a frozen BERT. The text embeddings help provide more representative cues, but also may cause misdirection due to prediction errors. Thus, VAC randomly drops or replaces historical predictions to increase the error tolerance to avoid collapsing onto the last few predictions. Qualitative and quantitative experiments demonstrate that the proposed method achieves superior performance compared to other state-of-the-art methods, with an average F1-score improvement of at least 7. 66%.

NeurIPS Conference 2025 Conference Paper

FSI-Edit: Frequency and Stochasticity Injection for Flexible Diffusion-Based Image Editing

  • Kaixiang Yang
  • Xin Li
  • Yuxi Li
  • Qiang Li
  • Zhiwei Wang

Latent Diffusion-based Text-to-Image (T2I) is a free image editing tool that typically reverses an image into noise, reconstructs it using its original text prompt, and then generates an edited version under a new target prompt. To preserve unaltered image content, features from the reconstruction are directly injected to replace selected features in the generation. However, this direct replacement often leads to feature incompatibility, compromising editing fidelity and limiting creative flexibility, particularly for non-rigid edits (\emph{e. g. }, structural or pose changes). In this paper, we aim to address these limitations by proposing \textbf{FSI-Edit}, a novel framework using frequency- and stochasticity-based feature injection for flexible image editing. First, FSI-Edit enhances feature consistency by injecting \emph{high-frequency} components of reconstruction features into generation features, mitigating incompatibility while preserving the editing ability for major structures encoded in low-frequency information. Second, it introduces controlled \emph{noise} into the replaced reconstruction features, expanding the generative space to enable diverse non-rigid edits beyond the original image’s constraints. Experiments on non-rigid edits, \emph{e. g. }, addition, deletion, and pose manipulation, demonstrate that FSI-Edit outperforms existing baselines in target alignment, semantic fidelity and visual quality. Our work highlights the critical roles of frequency-aware design and stochasticity in overcoming rigidity in diffusion-based editing.

AAAI Conference 2025 Conference Paper

MonoBox: Tightness-Free Box-Supervised Polyp Segmentation Using Monotonicity Constraint

  • Qiang Hu
  • Zhenyu Yi
  • Ying Zhou
  • Fan Huang
  • Mei Liu
  • Qiang Li
  • Zhiwei Wang

We propose MonoBox, an innovative box-supervised segmentation method constrained by monotonicity to liberate its training from the user-unfriendly box-tightness assumption. In contrast to conventional box-supervised segmentation, where the box edges must precisely touch the target boundaries, MonoBox leverages imprecisely-annotated boxes to achieve robust pixel-wise segmentation. The 'linchpin' is that, within the noisy zones around box edges, MonoBox discards the traditional misguiding multiple-instance learning loss, and instead optimizes a carefully-designed objective, termed monotonicity constraint. Along directions transitioning from the foreground to background, this new constraint steers responses to adhere to a trend of monotonically decreasing values. Consequently, the originally unreliable learning within the noisy zones is transformed into a correct and effective monotonicity optimization. Moreover, an adaptive label correction is introduced, enabling MonoBox to enhance the tightness of box annotations using predicted masks from the previous epoch and dynamically shrink the noisy zones as training progresses. We verify MonoBox in the box-supervised segmentation task of polyps, where satisfying box-tightness is challenging due to the vague boundaries between the polyp and normal tissues. Experiments on both public synthetic and in-house real noisy datasets demonstrate that MonoBox exceeds other anti-noise state-of-the-arts by improving Dice by at least 5.5% and 3.3%, respectively.

JBHI Journal 2025 Journal Article

Single-Slice Semi-Supervised 3D Medical Image Segmentation via Correlation Information Enhancement and Hybrid Pseudo Mask Generation

  • Quan Zhou
  • Mingwei Wen
  • Mingyue Ding
  • Yixin Su
  • Zhiwei Wang

Three-dimensional (3D) medical image segmentation typically demands extensive labeled training samples, which is prohibitively time-consuming and requires significant expertise. Although this demand can be mitigated by special learning paradigms such as semi-supervised learning, the cost is still high due to the reader-unfriendly 3D data structure. In this paper, we seek a solution of robust 3D segmentation using extremely simplified annotation that delineates only a single slice per each volume for only a subset of the 3D samples. To this end, we propose two innovative modules: a correlation-enhanced 3D segmentation model (CE-Seg) and a hybrid 3D pseudo mask generator (Hy-Gen). CE-Seg aims to comprehensively understand the 3D targets under super-sparse single-slice supervision by maximizing its ability to mine correlations across slices, spaces and scales. Specifically, CE-Seg mimics the radiologist's interpretation by ‘seeing’ a dynamically scrolling 3D image to enrich the slice-correlated context. It also introduces a drop-then-restoration self-played task to enhance the spatial correlations of features, and uses a bidirectional cascaded attention to interactively fuse features across different scales. To train CS-Seg, Hy-Gen combines learning-based and learning-free strategies to generate reliable pseudo 3D masks as supervisions. Concretely, Hy-Gen first employs a level-set evolution to ‘spread’ the single annotation to its neighboring slices as initialization. It then builds a teacher-student framework to progressively refine the initialized 3D mask by dynamically merging the predictions of the CS-Seg's teacher-copy. Extensive experiments on three public and one in-house datasets indicate that our method exceeds eight state-of-the-art semi-supervised methods by at least 3 $\%$ in dice, and is even on par with the full-supervised counterpart.

JBHI Journal 2025 Journal Article

TDSFE-Net: A Temporal Dual-Stream Feature Extraction Network for Depression Detection From EEG

  • Mingyang Li
  • Zhiwei Wang
  • Xi Yang
  • Tao Zhang

Early detection and diagnosis are critical for effective depression management. Although electroence-phalography (EEG) can provide an objective basis for the auxiliary diagnosis of depression, decoding depression-related brain activity from EEG is a highly challenging task due to the inherent complexity, dynamism, and non-linearity. Therefore, this study introduces a novel temporal dual-stream feature extraction network (TDSFE-Net) that incorporates multiple attention mechanisms. Specially, we first develop a dynamic fusion weight based local-global attention mechanism into the hierarchiclal temporal-separable convolutional network (TSCN) to automatically capture the temporal dynamic characteristics of the EEG signal. Subsequently, a channel-wise module is designed to reveal the key temporal information in spatial dimensions. Finally, a softmax with full conected layer is used as classifier. The TDSFE-Net achieved impressive classification accuracies of 98. 72%, 96. 91%, and 99. 53% on the MODMA, HUSM, and Hospital datasets, respectively. In addition, this study also reveals the pattern of correlation between the activity of specific brain regions and depression, providing a new perspective and scientific basis for discovering biomarkers and studying the neural mechanisms of depression.

IJCAI Conference 2024 Conference Paper

Improving Paratope and Epitope Prediction by Multi-Modal Contrastive Learning and Interaction Informativeness Estimation

  • Zhiwei Wang
  • Yongkang Wang
  • Wen Zhang

Accurately predicting antibody-antigen binding residues, i. e. , paratopes and epitopes, is crucial in antibody design. However, existing methods solely focus on uni-modal data (either sequence or structure), disregarding the complementary information present in multi-modal data, and most methods predict paratopes and epitopes separately, overlooking their specific spatial interactions. In this paper, we propose a novel Multi-modal contrastive learning and Interaction informativeness estimation-based method for Paratope and Epitope prediction, named MIPE, by using both sequence and structure data of antibodies and antigens. MIPE implements a multi-modal contrastive learning strategy, which maximizes representations of binding and non-binding residues within each modality and meanwhile aligns uni-modal representations towards effective modal representations. To exploit the spatial interaction information, MIPE also incorporates an interaction informativeness estimation that computes the estimated interaction matrices between antibodies and antigens, thereby approximating them to the actual ones. Extensive experiments demonstrate the superiority of our method in predicting paratopes and epitopes compared to baselines. Additionally, the ablation studies and visualizations demonstrate the superiority of MIPE owing to the better representations acquired through multi-modal contrastive learning and the interaction patterns comprehended by the interaction informativeness estimation.

NeurIPS Conference 2024 Conference Paper

Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing

  • Zhongwang Zhang
  • Pengxiao Lin
  • Zhiwei Wang
  • Yaoyu Zhang
  • Zhi-Qin J. Xu

Transformers have shown impressive capabilities across various tasks, but their performance on compositional problems remains a topic of debate. In this work, we investigate the mechanisms of how transformers behave on unseen compositional tasks. We discover that the parameter initialization scale plays a critical role in determining whether the model learns inferential (reasoning-based) solutions, which capture the underlying compositional primitives, or symmetric (memory-based) solutions, which simply memorize mappings without understanding the compositional structure. By analyzing the information flow and vector representations within the model, we reveal the distinct mechanisms underlying these solution types. We further find that inferential (reasoning-based) solutions exhibit low complexity bias, which we hypothesize is a key factor enabling them to learn individual mappings for single anchors. We validate our conclusions on various real-world datasets. Our findings provide valuable insights into the role of initialization scale in tuning the reasoning and memorizing ability and we propose the initialization rate $\gamma$ to be a convenient tunable hyper-parameter in common deep learning frameworks, where $1/d_{\mathrm{in}}^\gamma$ is the standard deviation of parameters of the layer with $d_{\mathrm{in}}$ input neurons.

JBHI Journal 2024 Journal Article

Multi-Feature Decision Fusion Network for Heart Sound Abnormality Detection and Classification

  • Haobo Zhang
  • Peng Zhang
  • Zhiwei Wang
  • Lianying Chao
  • Yuting Chen
  • Qiang Li

The heart sound reflects the movement status of the cardiovascular system and contains the early pathological information of cardiovascular diseases. Automatic heart sound diagnosis plays an essential role in the early detection of cardiovascular diseases. In this study, we aim to develop a novel end-to-end heart sound abnormality detection and classification method, which can be adapted to different heart sound diagnosis tasks. Specifically, we developed a Multi-feature Decision Fusion Network (MDFNet) composed of a Multi-dimensional Feature Extraction (MFE) module and a Multi-dimensional Decision Fusion (MDF) module. The MFE module extracted spatial features, multi-level temporal features and spatial-temporal fusion features to learn heart sound characteristics from multiple perspectives. Through deep supervision and decision fusion, the MDF module made the multi-dimensional features extracted by the MFE module more discriminative, and fused the decision results of multi-dimensional features to integrate complementary information. Furthermore, attention modules were embedded in the MDFNet to emphasize the fundamental heart sounds containing effective feature information. Finally, we proposed an efficient data augmentation method to circumvent the diagnosis performance degradation caused by the lack of cardiac cycle segmentation in other end-to-end methods. The developed method achieved an overall accuracy of 94. 44% and a F1-score of 86. 90% on the binary classification task and a F1-score of 99. 30% on the five-classification task. Our method outperformed other state-of-the-art methods and had good clinical application prospects.

AAAI Conference 2024 Conference Paper

Stealthy Adversarial Attacks on Stochastic Multi-Armed Bandits

  • Zhiwei Wang
  • Huazheng Wang
  • Hongning Wang

Adversarial attacks against stochastic multi-armed bandit (MAB) algorithms have been extensively studied in the literature. In this work, we focus on reward poisoning attacks and find most existing attacks can be easily detected by our proposed detection method based on the test of homogeneity, due to their aggressive nature in reward manipulations. This motivates us to study the notion of stealthy attack against stochastic MABs and investigate the resulting attackability. Our analysis shows that against two popularly employed MAB algorithms, UCB1 and $\epsilon$-greedy, the success of a stealthy attack depends on the environmental conditions and the realized reward of the arm pulled in the first round. We also analyze the situation for general MAB algorithms equipped with our attack detection method and find that it is possible to have a stealthy attack that almost always succeeds. This brings new insights into the security risks of MAB algorithms.

JBHI Journal 2024 Journal Article

Subgraph-Aware Graph Kernel Neural Network for Link Prediction in Biological Networks

  • Menglu Li
  • Zhiwei Wang
  • Luotao Liu
  • Xuan Liu
  • Wen Zhang

Identifying links within biological networks is important in various biomedical applications. Recent studies have revealed that each node in a network may play a unique role in different links, but most link prediction methods overlook distinctive node roles, hindering the acquisition of effective link representations. Subgraph-based methods have been introduced as solutions but often ignore shared information among subgraphs. To address these limitations, we propose a Subgraph-aware Graph Kernel Neural Network (SubKNet) for link prediction in biological networks. Specifically, SubKNet extracts a subgraph for each node pair and feeds it into a graph kernel neural network, which decomposes each subgraph into a combination of trainable graph filters with diversity regularization for subgraph-aware representation learning. Additionally, node embeddings of the network are extracted as auxiliary information, aiding in distinguishing node pairs that share the same subgraph. Extensive experiments on five biological networks demonstrate that SubKNet outperforms baselines, including methods especially designed for biological networks and methods adapted to various networks. Further investigations confirm that employing graph filters to subgraphs helps to distinguish node roles in different subgraphs, and the inclusion of diversity regularization further enhances its capacity from diverse perspectives, generating effective link representations that contribute to more accurate link prediction.

JBHI Journal 2023 Journal Article

Accurate Cobb Angle Estimation on Scoliosis X-Ray Images via Deeply-Coupled Two-Stage Network With Differentiable Cropping and Random Perturbation

  • Yuanhuai Liang
  • Jinxin Lv
  • Dun Li
  • Xin Yang
  • Zhiwei Wang
  • Qiang Li

Automated Cobb angle estimation on X-ray images is crucial to scoliosis diagnosis. The existing efforts are typically two extremes, which either laboriously detect the raw vertebral landmarks or directly regress Cobb angles from the entire image. In this paper, we propose a novel two-stage end-to-end method as a balanced solution, to avoid vulnerability to false landmarks, and to preserve flexibility in clinical usages. Concretely, we cascade two stages sequentially for detecting vertebrae and then regressing their bending directions instead of raw landmarks. In the detection stage, we combine two networks called LocNet and SegNet to robustly localize vertebrae, and meanwhile to suppress the false positives by additionally segmenting the whole spine. In the subsequent stage, we introduce a regression network named RegNet to accurately regress bending directions of localized vertebrae. Furthermore, the vertebra-aligned local regions on LocNet's intermediate features are cropped via RoIAlign-pooling, and RegNet inherits the cropped regions to learn only feature residuals. By doing so, the regression difficulty can be dramatically alleviated, and the two stages are deeply coupled and mutually guided in an end-to-end training. Moreover, a random perturbation on the inherited features further enhances RegNet's robustness. We benchmark our method on both public and private datasets, and the errors are 2. 92 $\pm$ 2. 34 $^{\circ }$ and 6. 87 $\pm$ 6. 26% in terms of CMAE and SMAPE on the widely-employed AASCE dataset, outperforming other state-of-the-arts by at least 16. 81% and 6. 15%, respectively. Also, a clinical user study verifies our promising flexibility for allowing convenient rectifications to further decrease errors by a large marge.

IJCAI Conference 2023 Conference Paper

Dual-view Correlation Hybrid Attention Network for Robust Holistic Mammogram Classification

  • Zhiwei Wang
  • Junlin Xian
  • Kangyi Liu
  • Xin Li
  • Qiang Li
  • Xin Yang

Mammogram image is important for breast cancer screening, and typically obtained in a dual-view form, i. e. , cranio-caudal (CC) and mediolateral oblique (MLO), to provide complementary information for clinical decisions. However, previous methods mostly learn features from the two views independently, which violates the clinical knowledge and ignores the importance of dual-view correlation in the feature learning. In this paper, we propose a dual-view correlation hybrid attention network (DCHA-Net) for robust holistic mammogram classification. Specifically, DCHA-Net is carefully designed to extract and reinvent deep feature maps for the two views, and meanwhile to maximize the underlying correlations between them. A hybrid attention module, consisting of local relation and non-local attention blocks, is proposed to alleviate the spatial misalignment of the paired views in the correlation maximization. A dual-view correlation loss is introduced to maximize the feature similarity between corresponding strip-like regions with equal distance to the chest wall, motivated by the fact that their features represent the same breast tissues, and thus should be highly-correlated with each other. Experimental results on the two public datasets, i. e. , INbreast and CBIS-DDSM, demonstrate that the DCHA-Net can well preserve and maximize feature correlations across views, and thus outperforms previous state-of-the-art methods for classifying a whole mammogram as malignant or not.

AAAI Conference 2023 Conference Paper

Robust One-Shot Segmentation of Brain Tissues via Image-Aligned Style Transformation

  • Jinxin Lv
  • Xiaoyu Zeng
  • Sheng Wang
  • Ran Duan
  • Zhiwei Wang
  • Qiang Li

One-shot segmentation of brain tissues is typically a dual-model iterative learning: a registration model (reg-model) warps a carefully-labeled atlas onto unlabeled images to initialize their pseudo masks for training a segmentation model (seg-model); the seg-model revises the pseudo masks to enhance the reg-model for a better warping in the next iteration. However, there is a key weakness in such dual-model iteration that the spatial misalignment inevitably caused by the reg-model could misguide the seg-model, which makes it converge on an inferior segmentation performance eventually. In this paper, we propose a novel image-aligned style transformation to reinforce the dual-model iterative learning for robust one-shot segmentation of brain tissues. Specifically, we first utilize the reg-model to warp the atlas onto an unlabeled image, and then employ the Fourier-based amplitude exchange with perturbation to transplant the style of the unlabeled image into the aligned atlas. This allows the subsequent seg-model to learn on the aligned and style-transferred copies of the atlas instead of unlabeled images, which naturally guarantees the correct spatial correspondence of an image-mask training pair, without sacrificing the diversity of intensity patterns carried by the unlabeled images. Furthermore, we introduce a feature-aware content consistency in addition to the image-level similarity to constrain the reg-model for a promising initialization, which avoids the collapse of image-aligned style transformation in the first iteration. Experimental results on two public datasets demonstrate 1) a competitive segmentation performance of our method compared to the fully-supervised method, and 2) a superior performance over other state-of-the-art with an increase of average Dice by up to 4.67%. The source code is available at: https://github.com/JinxLv/One-shot-segmentation-via-IST.

JBHI Journal 2023 Journal Article

Self-Supervised Triplet Contrastive Learning for Classifying Endometrial Histopathological Images

  • Fengjun Zhao
  • Zhiwei Wang
  • Hongyan Du
  • Xiaowei He
  • Xin Cao

Early identification of endometrial cancer or precancerous lesions from histopathological images is crucial for precise endometrial medical care, which however is increasing hampered by the relative scarcity of pathologists. Computer-aided diagnosis (CAD) provides an automated alternative for confirming endometrial diseases with either feature-engineered machine learning or end-to-end deep learning (DL). In particular, advanced self-supervised learning alleviates the dependence of supervised learning on large-scale human-annotated data and can be used to pre-train DL models for specific classification tasks. Thereby, we develop a novel self-supervised triplet contrastive learning (SSTCL) model for classifying endometrial histopathological images. Specifically, this model consists of one online branch and two target branches. The second target branch includes a simple yet powerful augmentation module named random mosaic masking (RMM), which functions as an effective regularization by mapping the features of masked images close to those of intact ones. Moreover, we add a bottleneck Transformer (BoT) model into each branch as a self-attention module to learn the global information by considering both content information and relative distances between features at different locations. On public endometrial dataset, our model achieved four-class classification accuracies of 77. 31 ± 0. 84, 80. 87 ± 0. 48 and 83. 22 ± 0. 87% using 20, 50 and 100% labeled images, respectively. When transferred to the in-house dataset, our model obtained a three-class diagnostic accuracy of 96. 81% with 95% confidence interval of 95. 61–98. 02%. On both datasets, our model outperformed state-of-the-art supervised and self-supervised methods. Our model may help pathologists to automatically diagnose endometrial diseases with high accuracy and efficiency using limited human-annotated histopathological images.

JBHI Journal 2021 Journal Article

Variation-Aware Federated Learning With Multi-Source Decentralized Medical Image Data

  • Zengqiang Yan
  • Jeffry Wicaksana
  • Zhiwei Wang
  • Xin Yang
  • Kwang-Ting Cheng

Privacy concerns make it infeasible to construct a large medical image dataset by fusing small ones from different sources/institutions. Therefore, federated learning (FL) becomes a promising technique to learn from multi-source decentralized data with privacy preservation. However, the cross-client variation problem in medical image data would be the bottleneck in practice. In this paper, we propose a variation-aware federated learning (VAFL) framework, where the variations among clients are minimized by transforming the images of all clients onto a common image space. We first select one client with the lowest data complexity to define the target image space and synthesize a collection of images through a privacy-preserving generative adversarial network, called PPWGAN-GP. Then, a subset of those synthesized images, which effectively capture the characteristics of the raw images and are sufficiently distinct from any raw image, is automatically selected for sharing with other clients. For each client, a modified CycleGAN is applied to translate its raw images to the target image space defined by the shared synthesized images. In this way, the cross-client variation problem is addressed with privacy preservation. We apply the framework for automated classification of clinically significant prostate cancer and evaluate it using multi-source decentralized apparent diffusion coefficient (ADC) image data. Experimental results demonstrate that the proposed VAFL framework stably outperforms the current horizontal FL framework. As VAFL is independent of deep learning architectures for classification, we believe that the proposed framework is widely applicable to other medical image classification tasks.

JBHI Journal 2020 Journal Article

Bi-Modality Medical Image Synthesis Using Semi-Supervised Sequential Generative Adversarial Networks

  • Xin Yang
  • Yi Lin
  • Zhiwei Wang
  • Xin Li
  • Kwang-Ting Cheng

In this paper, we propose a bi-modality medical image synthesis approach based on sequential generative adversarial network (GAN) and semi-supervised learning. Our approach consists of two generative modules that synthesize images of the two modalities in a sequential order. A method for measuring the synthesis complexity is proposed to automatically determine the synthesis order in our sequential GAN. Images of the modality with a lower complexity are synthesized first, and the counterparts with a higher complexity are generated later. Our sequential GAN is trained end-to-end in a semi-supervised manner. In supervised training, the joint distribution of bi-modality images are learned from real paired images of the two modalities by explicitly minimizing the reconstruction losses between the real and synthetic images. To avoid overfitting limited training images, in unsupervised training, the marginal distribution of each modality is learned based on unpaired images by minimizing the Wasserstein distance between the distributions of real and fake images. We comprehensively evaluate the proposed model using two synthesis tasks based on three types of evaluate metrics and user studies. Visual and quantitative results demonstrate the superiority of our method to the state-of-the-art methods, and reasonable visual quality and clinical significance. Code is made publicly available at https://github.com/hust-linyi/Multimodal-Medical-Image-Synthesis.

AAAI Conference 2020 Conference Paper

Learning Multi-Level Dependencies for Robust Word Recognition

  • Zhiwei Wang
  • Hui Liu
  • Jiliang Tang
  • Songfan Yang
  • Gale Yan Huang
  • Zitao Liu

Robust language processing systems are becoming increasingly important given the recent awareness of dangerous situations where brittle machine learning models can be easily broken with the presence of noises. In this paper, we introduce a robust word recognition framework that captures multi-level sequential dependencies in noised sentences. The proposed framework employs a sequence-to-sequence model over characters of each word, whose output is given to a word-level bi-directional recurrent neural network. We conduct extensive experiments to verify the effectiveness of the framework. The results show that the proposed framework outperforms state-of-the-art methods by a large margin and they also suggest that character-level dependencies can play an important role in word recognition. The code of the proposed framework and the major experiments are publicly available1.

UAI Conference 1994 Conference Paper

On Axiomatization of Probabilistic Conditional Independencies

  • S. K. Michael Wong
  • Zhiwei Wang

This paper studies the connection between probabilistic conditional independence in uncertain reasoning and data dependency in relational databases. As a demonstration of the usefulness of this preliminary investigation, an alternate proof is presented for refuting the conjecture suggested by Pearl and Paz that probabilistic conditional independencies have a complete axiomatization.

UAI Conference 1993 Conference Paper

Qualitative Measures of Ambiguity

  • S. K. Michael Wong
  • Zhiwei Wang

This paper introduces a qualitative measure of ambiguity and analyses its relationship with other measures of uncertainty. Probability measures relative likelihoods, while ambiguity measures vagueness surrounding those judgments. Ambiguity is an important representation of uncertain knowledge. It deals with a different, type of uncertainty modeled by subjective probability or belief.

v2026.09.13