Arrow Research search

Author name cluster

Wenhui Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
1 author row

Possible papers

12

AAAI Conference 2026 Conference Paper

Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning

  • Chenyu Zhang
  • Lanjun Wang
  • Yiwen Ma
  • Wenhui Li
  • Guoqing Jin
  • Anan Liu

Text-to-Image (T2I) models typically deploy safety mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreaking attack methods manually design instructions for the LLM to generate adversarial prompts, which effectively exposing safety vulnerabilities of T2I models. However, existing methods have two limitations: 1) relying on manually exhaustive strategies for designing adversarial prompts, lacking a unified framework, and 2) requiring numerous queries to achieve a successful attack, limiting their practical applicability. To address this issue, we propose Reason2Attack~(R2A), which aims to enhance the effectiveness and efficiency of the LLM in jailbreaking attacks. Specifically, we first use Frame Semantics theory to systematize existing manually crafted strategies and propose a unified generation framework to generate CoT adversarial prompts step by step. Following this, we propose a two-stage LLM reasoning training framework guided by the attack process. In the first stage, the LLM is fine-tuned with CoT examples generated by the unified generation framework to internalize the adversarial prompt generation process grounded in Frame Semantics. In the second stage, we incorporate the jailbreaking task into the LLM's reinforcement learning process, guided by the proposed attack process reward function that balances prompt stealthiness, effectiveness, and length, enabling the LLM to understand T2I models and safety mechanisms. Extensive experiments on various T2I models with safety mechanisms, and commercial T2I models, show the superiority and practicality of R2A.

AAAI Conference 2026 Conference Paper

SGoT-R1: Social Graph of Thought Reasoning-Enhanced Multimodal Large Language Model for Harmful Meme Detection

  • Xiuxian Wang
  • Yuting Su
  • Wenhui Li
  • Xiaowen Wang
  • Zhuojun Li
  • Anan Liu

Internet memes serve as widely distributed multimodal social content that conveys complex ideas through metaphorical expressions, often containing harmful implications that make accurate harmful meme detection an important problem. Reasoning knowledge extracted from large language models plays a crucial role in recent advances in harmful meme detection. However, these methods only perform reasoning analysis on memes from a single opinion, ignoring that memes are essentially products of group consensus, where their true meaning interpretation highly depends on the collision and aggregation process of diverse user viewpoints. To address this problem, we propose a Social Graph of Thought Reasoning Enhancement (SGoTRE) framework for harmful meme detection. The SGoTRE contains three key steps: First, through multi-agent simulation technology, we obtain diverse chains of thought that represent the parsing logic of users from different backgrounds toward memes, authentically restoring the diversity characteristics of group cognition. Second, we construct a Social Graph of Thought (SGoT) that effectively integrates multi-chain reasoning processes and structurally expresses the consensus and diversity of viewpoints among users. Finally, we utilize the SGoT for cognitive distillation, internalizing multi-opinion reasoning logic into a single multimodal large model SGoT-R1 to achieve efficient and interpretable harmful meme detection. Experimental results show that SGoT-R1 significantly improves detection performance on mainstream datasets. Particularly on the most challenging FHM dataset, SGoT-R1 achieves an 8.9% improvement over state-of-the-art models.

JBHI Journal 2026 Journal Article

SpineVLM: A Markdown-Guided Structured Fine-Tuning Framework for Spine X-ray Report Generation

  • Dong Liu
  • Wenhui Li
  • Ning Xu
  • Guoge Han
  • Rui Hao
  • Xianzhu Liu
  • An-An Liu

Automated medical report generation in specialized fields like spine radiography is constrained by data scarcity and high annotation costs. Consequently, existing multimodal large language models (MLLMs) struggle in these settings, often missing minute, scattered spinal abnormalities. We introduce SpineVLM, a data-efficient framework for structured spine X-ray report generation. The framework is built upon the newly constructed SXRG dataset, comprising 10, 468 image-report pairs developed via a hierarchical AI-assisted annotation pipeline. To optimize learning under limited data, we propose Markdown-Guided Structured Learning (MGSL), which reformulates unconstrained free-text synthesis into a structured completion task, acting as a strong regularizer. Furthermore, an unsupervised Region-Focused Inference (RFI) module powered by foundation models (DINOv2) isolates the vertebral column to enhance the perception of subtle lesions without requiring manual spatial annotations. Evaluated on a 7B-parameter vision-language backbone, SpineVLM achieves strong performance against ten baseline multimodal models across standard linguistic metrics. In a double-blind reader study, the system achieved a diagnostic F1-score of 0. 866, comparable to specialist performance, while reducing clinical reporting time by over 41%. By open-sourcing the dataset and codebase, we provide, to our knowledge, the first quantitative benchmark for automated spine radiography report generation, together with a structured framework for this data-limited setting. All data and code will be publicly released at https://github.com/LiuDongDaniel/SpineVLM.

AAAI Conference 2026 Conference Paper

T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model

  • Chenyu Zhang
  • Tairen Zhang
  • Lanjun Wang
  • Ruidong Chen
  • Wenhui Li
  • Anan Liu

Using risky text prompts, such as pornography and violent prompts, to test the safety of text-to-image (T2I) models is a critical task. However, existing risky prompt datasets are limited in three key areas: 1) limited risky categories, 2) coarse-grained annotation, and 3) low effectiveness. To address these limitations, we introduce T2I-RiskyPrompt, a comprehensive benchmark designed for evaluating safety-related tasks in T2I models. Specifically, we first develop a hierarchical risk taxonomy, which consists of 6 primary categories and 14 fine-grained subcategories. Building upon this taxonomy, we construct a pipeline to collect and annotate risky prompts. Finally, we obtain 6,432 effective risky prompts, where each prompt is annotated with both hierarchical category labels and detailed risk reasons. Moreover, to facilitate the evaluation, we propose a reason-driven risky image detection method that explicitly aligns the MLLM with safety annotations. Based on T2I-RiskyPrompt, we conduct a comprehensive evaluation of eight T2I models, nine defense methods, five safety filters, and five attack strategies, offering nine key insights into the strengths and limitations of T2I model safety. Finally, we discuss potential applications of T2I-RiskyPrompt across various research fields.

EAAI Journal 2024 Journal Article

Dynamic graphs attention for ocean variable forecasting

  • Junhao Wang
  • Zhengya Sun
  • Chunxin Yuan
  • Wenhui Li
  • An-An Liu
  • Zhiqiang Wei
  • Bo Yin

Forecasting the ocean dynamics is a critical issue for a wide array of climate extremes and environmental crisis. The dynamic variations are traditionally approached by relying on numerical models with all the related physical processes identified beforehand. An efficient alternative forecasting approach is based on the data-driven models. Despite their potential ability in modeling spatio-temporal ocean data, they ignore the fact that the ocean variables in different spatial regions and time periods typically have ever changing influences on each other, thus cannot yield satisfactory prediction results. In this paper, we develop a novel attention based dynamic graph for the ocean variable forecasting problem, which captures both the spatial and temporal dependencies. Specifically, we employ joint self-attention to incorporate information from the spatial graph over the target region, and model the graph evolution across long-range time steps. The performance of the proposed prediction model has been examined in the Indian Ocean based on ocean grid data products datasets. Experimental results demonstrate that this model has significant forecasting capability within 12 months, compared with the numerical methods and the state-of-the-art spatio-temporal embedding baselines.

EAAI Journal 2024 Journal Article

Heterogeneous propagation graph convolution network for a recommendation system based on a knowledge graph

  • Jiawei Lu
  • Jiapeng Li
  • Wenhui Li
  • Junfeng Song
  • Gang Xiao

Recently, there has been a surge of interest in recommendation systems that leverage knowledge graphs, primarily because of their effectiveness in addressing sparsity and cold-start challenges inherent in collaborative filtering approaches. In most previous studies, researchers have focused on the way knowledge associations are encoded in knowledge graphs, but have not sufficiently highlighted the signals of collaboration that are implicit in the interaction between users and items. As a result, the learned embeddings do not provide a complete representation of the semantic information. In this paper, we describe a new model called a heterogeneous propagation graph convolution network for a recommendation system combined with a knowledge graph (HP-GCN). It adopts a heterogeneous propagation to generate user embedding representations, thereby combining encoded collaborative signals and auxiliary knowledge in knowledge graphs. Furthermore, we incorporate an attention mechanism to differentiate the contributions made by diverse neighbors as opposed to those made by users. Since most graph convolutions tend to suffer from over-smoothing when the number of convolutional layers increases, leading to insufficient utilization of high-order information, this paper uses an improved graph convolution strategy to generate item embeddings. This strategy has two different aggregation mechanisms embedding into different subgraphs, which can more fully utilize high-order information and mitigate the over-smoothing problem. Thus, we are able to efficiently prevent negative information originating from higher-order neighbors into the process of embedding learning. In extensive experiments, we applied HP-GCN to four large-scale real datasets for music, books, movies, and restaurants. The experimental outcomes revealed that HP-GCN generally surpassed the baseline methods in both recommendation accuracy and diversity, showing superior recommendation performance overall.

AAAI Conference 2023 Conference Paper

Mx2M: Masked Cross-Modality Modeling in Domain Adaptation for 3D Semantic Segmentation

  • Boxiang Zhang
  • Zunran Wang
  • Yonggen Ling
  • Yuanyuan Guan
  • Shenghao Zhang
  • Wenhui Li

Existing methods of cross-modal domain adaptation for 3D semantic segmentation predict results only via 2D-3D complementarity that is obtained by cross-modal feature matching. However, as lacking supervision in the target domain, the complementarity is not always reliable. The results are not ideal when the domain gap is large. To solve the problem of lacking supervision, we introduce masked modeling into this task and propose a method Mx2M, which utilizes masked cross-modality modeling to reduce the large domain gap. Our Mx2M contains two components. One is the core solution, cross-modal removal and prediction (xMRP), which makes the Mx2M adapt to various scenarios and provides cross-modal self-supervision. The other is a new way of cross-modal feature matching, the dynamic cross-modal filter (DxMF) that ensures the whole method dynamically uses more suitable 2D-3D complementarity. Evaluation of the Mx2M on three DA scenarios, including Day/Night, USA/Singapore, and A2D2/SemanticKITTI, brings large improvements over previous methods on many metrics.

AAAI Conference 2022 Conference Paper

Memory-Based Jitter: Improving Visual Recognition on Long-Tailed Data with Diversity in Memory

  • Jialun Liu
  • Wenhui Li
  • Yifan Sun

This paper considers deep visual recognition on long-tailed data. To make our method general, we tackle two applied scenarios, i. e. , deep classification and deep metric learning. Under the long-tailed data distribution, the most classes (i. e. , tail classes) only occupy relatively few samples and are prone to lack of within-class diversity. A radical solution is to augment the tail classes with higher diversity. To this end, we introduce a simple and reliable method named Memorybased Jitter (MBJ). We observe that during training, the deep model constantly changes its parameters after every iteration, yielding the phenomenon of weight jitters. Consequentially, given a same image as the input, two historical editions of the model generate two different features in the deeplyembedded space, resulting in feature jitters. Using a memory bank, we collect these (model or feature) jitters across multiple training iterations and get the so-called Memory-based Jitter. The accumulated jitters enhance the within-class diversity for the tail classes and consequentially improves longtailed visual recognition. With slight modifications, MBJ is applicable for two fundamental visual recognition tasks, i. e. , deep image classification and deep metric learning (on longtailed data). Extensive experiments on five long-tailed classification benchmarks and two deep metric learning benchmarks demonstrate significant improvement. Moreover, the achieved performance are on par with the state of the art on both tasks.

IJCAI Conference 2020 Conference Paper

Consistent Domain Structure Learning and Domain Alignment for 2D Image-Based 3D Objects Retrieval

  • Yuting Su
  • Yuqian Li
  • Dan Song
  • Weizhi Nie
  • Wenhui Li
  • An-An Liu

2D image-based 3D objects retrieval is a new topic for 3D objects retrieval which can be used to manage 3D data with 2D images. The goal is to search some related 3D objects when given a 2D image. The task is challenging due to the large domain gap between 2D images and 3D objects. Therefore, it is essential to consider domain adaptation problems to reduce domain discrepancy. However, most of the existing domain adaptation methods only utilize the semantic information from the source domain to predict labels in the target domain and neglect the intrinsic structure of the target domain. In this paper, we propose a domain alignment framework with consistent domain structure learning to reduce the large gap between 2D images and 3D objects. The domain structure learning module makes use of both the semantic information from the source domain and the intrinsic structure of the target domain, which provides more reliable predicted labels to the domain alignment module to better align the conditional distribution. We conducted experiments on two public datasets, MI3DOR and MI3DOR-2, and the experimental results demonstrate the proposed method outperforms the state-of-the-art methods.

IJCAI Conference 2020 Conference Paper

Hierarchical Instance Feature Alignment for 2D Image-Based 3D Shape Retrieval

  • Heyu Zhou
  • Weizhi Nie
  • Wenhui Li
  • Dan Song
  • An-An Liu

2D image-based 3D shape retrieval has become a hot research topic since its wide industrial applications and academic significance. However, existing view-based 3D shape retrieval methods are restricted by two settings, 1) learn the common-class features while neglecting the instance visual characteristics, 2) narrow the global domain variations while ignoring the local semantic variations in each category. To overcome these problems, we propose a novel hierarchical instance feature alignment (HIFA) method for this task. HIFA consists of two modules, cross-modal instance feature learning and hierarchical instance feature alignment. Specifically, we first use CNN to extract both 2D image and multi-view features. Then, we maximize the mutual information between the input data and the high-level feature to preserve as much as visual characteristics of an individual instance. To mix up the features in two domains, we enforce feature alignment considering both global domain and local semantic levels. By narrowing the global domain variations we impose the identical large norm restriction on both 2D and 3D feature-norm expectations to facilitate more transferable possibility. By narrowing the local variations we propose to minimize the distance between two centroids of the same class from different domains to obtain semantic consistency. Extensive experiments on two popular and novel datasets, MI3DOR and MI3DOR-2, validate the superiority of HIFA for 2D image-based 3D shape retrieval task.

IJCAI Conference 2018 Conference Paper

Cross-Domain 3D Model Retrieval via Visual Domain Adaption

  • Anan Liu
  • Shu Xiang
  • Wenhui Li
  • Weizhi Nie
  • Yuting Su

Recent advances in 3D capturing devices and 3D modeling software have led to extensive and diverse 3D datasets, which usually have different distributions. Cross-domain 3D model retrieval is becoming an important but challenging task. However, existing works mainly focus on 3D model retrieval in a closed dataset, which seriously constrain their implementation for real applications. To address this problem, we propose a novel crossdomain 3D model retrieval method by visual domain adaptation. This method can inherit the advantage of deep learning to learn multi-view visual features in the data-driven manner for 3D model representation. Moreover, it can reduce the domain divergence by exploiting both domainshared and domain-specific features of different domains. Consequently, it can augment the discrimination of visual descriptors for cross-domain similarity measure. Extensive experiments on two popular datasets, under three designed cross-domain scenarios, demonstrate the superiority and effectiveness of the proposed method by comparing against the state-of-the-art methods. Especially, the proposed method can significantly outperform the most recent method for cross-domain 3D model retrieval and the champion of Shrec’16 Large-Scale 3D Shape Retrieval from ShapeNet Core55.

IJCAI Conference 2018 Conference Paper

Hierarchical Graph Structure Learning for Multi-View 3D Model Retrieval

  • Yuting Su
  • Wenhui Li
  • Anan Liu
  • Weizhi Nie

3D model retrieval has been widely utilized in numerous domains, such as computer-aided design, digital entertainment and virtual reality. Recently, many graph-based methods have been proposed to address this task by using multiple views of 3D models. However, these methods are always constrained by the many-to-many graph matching for similarity measure between pair-wise models. In this paper, we propose an hierarchical graph structure learning method (HGS) for 3D model retrieval. The proposed method can decompose the complicated multi-view graph-based similarity measure into multiple single-view graph-based similarity measures. In the bottom hierarchy, we present the method for single-view graph generation and further propose the novel method for similarity measure in single-view graph by leveraging both node-wise context and model-wise context. In the top hierarchy, we fuse the similarities in single-view graphs with respect to different viewpoints to get the multi-view similarity between pair-wise models. In this way, the proposed method can avoid the difficulty in definition and computation in the traditional high-order graph. Moreover, this method is unsupervised and is independent of large-scale 3D dataset for model learning. We conduct extensive evaluation on three popular and challenging datasets. The comparison demonstrates the superiority and effectiveness of the proposed method comparing with the state of the arts. Especially, this unsupervised method can achieve competing performance against the most recent supervised & deep learning method.

v2026.09.13