Arrow Research search

Author name cluster

Jie Wu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

33 papers
2 author rows

Possible papers

33

AAAI Conference 2026 Conference Paper

Boosting Adversarial Transferability via Ensemble Non-Attention

  • Yipeng Zou
  • Qin Liu
  • Jie Wu
  • Yu Peng
  • Guo Chen
  • Hui Zhou
  • Guanghui Ye

Ensemble attacks integrate the outputs of surrogate models with diverse architectures, which can be combined with various gradient-based attacks to improve adversarial transferability. However, previous work shows unsatisfactory attack performance when transferring across heterogeneous model architectures. The main reason is that the gradient update directions of heterogeneous surrogate models differ widely, making it hard to reduce the gradient variance of ensemble models while making the best of individual model. To tackle this challenge, we design a novel ensemble attack, NAMEA, which for the first time integrates the gradients from the non-attention areas of ensemble models into the iterative gradient optimization process. Our design is inspired by the observation that the attention areas of heterogeneous models vary sharply, thus the non-attention areas of ViTs are likely to be the focus of CNNs and vice versa. Therefore, we merge the gradients respectively from the attention and non-attention areas of ensemble models so as to fuse the transfer information of CNNs and ViTs. Specifically, we pioneer a new way of decoupling the gradients of non-attention areas from those of attention areas, while merging gradients by meta-learning. Empirical evaluations on ImageNet dataset indicate that NAMEA outperforms AdaEA and SMER, the state-of-the-art ensemble attacks by an average of 15.0% and 9.6%, respectively. This work is the first attempt to explore the power of ensemble non-attention in boosting cross-architecture transferability, providing new insights into launching ensemble attacks.

EAAI Journal 2026 Journal Article

Semi-solid metal oxide slag morphology detection based on cross-modal learning

  • Jie Wu
  • Degang Xu

The current robotic approach for slag removal in metal ingot casting is deficient in its capacity to detect the morphology of oxide slag, hence failing to ensure the consistency of ingot quality. This research introduces an advanced detection method for semi-solid metal oxide slag via cross-modal learning. By learning the visual and thermal information of cross-modal, the method enables the slagging robot to intelligently perceive the morphology of the oxide slag, so as to ensure the quality of the robot slagging process. A cross-modal learning model is introduced to ascertain the thickness and spatial distribution of the oxidized slag by integrating the relevant characteristics from the visual and thermal modalities of the oxidized slag. First, a multi-scale video block approach is developed to align the modal dimensions and create a multi-scale spatiotemporal block sequence. The intra-modal characteristics and cross-modal correlations are derived from these sequences utilizing the self-attention and cross-attention mechanisms. The obtained cross-modal features are then put into a multi-head decoder to detect the location distribution and identify the thickness of the oxide slag. To evaluate the effectiveness of the proposed approach, the cross-modal learning detection method accomplished accurate identification of the state of the oxide slag, and accurate detection can be achieved even under specular reflection. In 160 field slag removal operations, the approach lowered the error rate of typical robot slag removal from 6. 8% to less than 1. 4%, considerably enhancing the grade of metal ingots after slag removal.

AAAI Conference 2026 Conference Paper

StyleBreak: Revealing Alignment Vulnerabilities in Large Audio-Language Models via Style-Aware Audio Jailbreak

  • Hongyi Li
  • Chengxuan Zhou
  • Chu Wang
  • Sicheng Liang
  • Yanting Chen
  • Qinlin Xie
  • Jiawei Ye
  • Jie Wu

Large Audio-language Models (LAMs) have recently enabled powerful speech-based interactions by coupling audio encoders with Large Language Models (LLMs). However, the security of LAMs under adversarial attacks remains underexplored, especially through audio jailbreaks that craft malicious audio prompts to bypass alignment. Existing efforts primarily rely on converting text-based attacks into speech or applying shallow signal-level perturbations, overlooking the impact of human speech’s expressive variations on LAM alignment robustness. To address this gap, we propose StyleBreak, a novel style-aware audio jailbreak framework that systematically investigates how diverse human speech attributes affect LAM alignment robustness. Specifically, StyleBreak employs a two-stage style-aware transformation pipeline that perturbs both textual content and audio to control linguistic, paralinguistic, and extralinguistic attributes. Furthermore, we develop a query-adaptive policy network that automatically searches for adversarial styles to enhance the efficiency of LAM jailbreak exploration. Extensive evaluations demonstrate that LAMs exhibit critical vulnerabilities when exposed to diverse human speech attributes. Moreover, StyleBreak achieves substantial improvements in attack effectiveness and efficiency across multiple attack paradigms, highlighting the urgent need for more robust alignment in LAMs.

ECAI Conference 2025 Conference Paper

Controllable Face Inpainting via Pseudo-Style Embedding

  • Jie Wu
  • Jiawei Jiang 0002
  • Yueqian Quan
  • Honghui Xu 0002
  • Wanjun Chen
  • Jianwei Zheng 0001

Image inpainting, a critical facet of computer vision, is in full bloom accompanied by the rapid innovation of convolution neural networks and transformers, revolutionizing the practical management of abnormity disposal, image editing, etc. Of these applications, face inpainting is more challenging due to the higher demand for semantic accuracy in key regions such as eyes and nose. Classical face inpainting methods are celebrated for their fast generation speed and refined texture details. However, they often lack the level of controllability required for complex tasks. In contrast, existing multi-modal controllable inpainting techniques offer enhanced guidance through image-text integration but tend to be time-consuming and produce suboptimal texture refinement. To address these limitations, we propose the Multi-modal Pseudo-style Embedded Transformer (MPET), a novel and efficient multi-modal inpainting algorithm that seamlessly integrates the strengths of both approaches, achieving state-of-the-art performance. Specifically, edge completion facilitates a cost-efficient and simple bridging of the contour continuity. Multi-modal pseudo-style generation amalgamates the image-text modalities, successfully embedding text features within the visual vectors, thereby culminating in the formation of pseudo-style diagrams rich in diverse attributes. On that basis, a controllable style-embedded siamese network is elaborated, effectively orchestrating the interaction among style attributes while ensuring high-precision pixel infusion. Extensive experiments on public datasets demonstrate the superiority of our approach through both quantitative and qualitative evaluations, highlighting its potential to advance the field of face inpainting.

IJCAI Conference 2025 Conference Paper

FedCPD: Personalized Federated Learning with Prototype-Enhanced Representation and Memory Distillation

  • Kaili Jin
  • Li Xu
  • Xiaoding Wang
  • Sun-Yuan Hsieh
  • Jie Wu
  • Limei Lin

Federated learning, as a distributed learning framework, aims to develop a global model while preserving client privacy. However, heterogeneity of client data leads to fairness issues and reduced performance. Techniques like parameter decoupling and prototype learning appear promising, yet challenges such as forgetting historical data and limited generalization persist. These methods also lack local insights, with locally trained features prone to overfitting, which affects generalization in global parameter aggregation. To address these challenges, we propose FedCPD, a personalized federated learning framework. FedCPD maintains historical information, reduces information loss, and increases personalization through hierarchical feature distillation and cross-layer feature fusion. Moreover, we utilize representation techniques like prototype contrastive learning and prototype alignment to capture diverse client data features, thus improving model generalization and fairness. Experiments show FedCPD outperforms state-of-the-art models, enhancing generalization by up to 10. 40% and personalization by up to 4. 90%, highlighting its effectiveness and superiority.

IJCAI Conference 2025 Conference Paper

FedHAN: A Cache-Based Semi-Asynchronous Federated Learning Framework Defending Against Poisoning Attacks in Heterogeneous Clients

  • Xiaoding Wang
  • Bin Ye
  • Li Xu
  • Lizhao Wu
  • Sun-Yuan Hsieh
  • Jie Wu
  • Limei Lin

Federated learning is vulnerable to model poisoning attacks in which malicious participants compromise the global model by altering the model updates. Current defense strategies are divided into three types: aggregation-based methods, validation dataset-based methods, and update distance-based methods. However, these techniques often neglect the challenges posed by device heterogeneity and asynchronous communication. Even upon identifying malicious clients, the global model may already be significantly damaged, requiring effective recovery strategies to reduce the attacker's impact. Current recovery methods, which are based on historical update records, are limited in environments with device heterogeneity and asynchronous communication. To address these problems, we introduce FedHAN, a reliable federated learning algorithm designed for asynchronous communication and device heterogeneity. FedHAN customizes sparse models, uses historical client updates to impute missing parameters in sparse updates, dynamically assigns adaptive weights, and combines update deviation detection with update prediction-based model recovery. Theoretical analysis indicates that FedHAN achieves favorable convergence despite unbounded staleness and effectively discriminates between benign and malicious clients. Experiments reveal that FedHAN, compared to leading methods, increases the accuracy of the model by 7. 86%, improves the detection accuracy of poisoning attacks by 12%, and enhances the recovery accuracy by 7. 26%. As evidenced by these results, FedHAN exhibits enhanced reliability and robustness in intricate and dynamic federated learning scenarios.

EAAI Journal 2025 Journal Article

Inter-graph and Intra-graph: Utilizing global financial markets and constituent stocks for stock index prediction

  • Yong Shi
  • Yunong Wang
  • Jie Wu

Stock index prediction is a significant yet difficult undertaking due to its incorporation of complex and diverse information. Following the implementation of Graph Neural Networks in financial data analysis, numerous researchers have focused on the node-level task of forecasting individual stock movements by analyzing the relationships between stocks. However, two key challenges remain: first, realizing different speeds of feature propagation among nodes in graph representation learning; second, predicting stock indices by extracting and aggregating fluctuations from constituent stocks through graph-level tasks remains unaddressed. To tackle these challenges, this paper proposes a novel spatio-temporal prediction framework combining both node-level and graph-level tasks. The framework includes two types of graphs: inter-graph and intra-graph, which combine information from the micro, meso, and macro dimensions. For the inter-graph at the node level, we introduce the Granger causality test as an innovative node filtering method, which realizes the propagation of features between nodes with different strengths and speeds in the process of graph representation learning. For the intra-graph at the graph level, we examine various graph pooling methods and pooling proportions of stock index constituents to enhance the interpretability of the results and to provide new theoretical insights for stock index prediction. In conclusion, we develop the Graph Representation Learning-based Long Short-Term Memory (GRL-LSTM) model for forecasting stock index movements, and demonstrate the superiority of our approach on four major Chinese stock markets.

AAAI Conference 2025 Conference Paper

JailPO: A Novel Black-Box Jailbreak Framework via Preference Optimization Against Aligned LLMs

  • Hongyi Li
  • Jiawei Ye
  • Jie Wu
  • Tianjie Yan
  • Chu Wang
  • Zhixin Li

Large Language Models (LLMs) aligned with human feedback have recently garnered significant attention. However, it remains vulnerable to jailbreak attacks, where adversaries manipulate prompts to induce harmful outputs. Exploring jailbreak attacks enables us to investigate the vulnerabilities of LLMs and further guides us in enhancing their security. Unfortunately, existing techniques mainly rely on handcrafted templates or generated-based optimization, posing challenges in scalability, efficiency and universality. To address these issues, we present JailPO, a novel black-box jailbreak framework to examine LLM alignment. For scalability and universality, JailPO meticulously trains attack models to automatically generate covert jailbreak prompts. Furthermore, we introduce a preference optimization-based attack method to enhance the jailbreak effectiveness, thereby improving efficiency. To analyze model vulnerabilities, we provide three flexible jailbreak patterns. Extensive experiments demonstrate that JailPO not only automates the attack process while maintaining effectiveness but also exhibits superior performance in efficiency, universality, and robustness against defenses compared to baselines. Additionally, our analysis of the three JailPO patterns reveals that attacks based on complex templates exhibit higher attack strength, whereas covert question transformations elicit riskier responses and are more likely to bypass defense mechanisms.

AAAI Conference 2025 Conference Paper

Multiple Feature Refining Network for Visual Emotion Distribution Learning

  • Qinfu Xu
  • Shaozu Yuan
  • Yiwei Wei
  • Jie Wu
  • Leiquan Wang
  • Chunlei Wu

The significance of visual emotion distribution learning (VEDL) has surged, particularly with the growing inclination to convey emotions through images. The key of VEDL lies in capturing both low- and high-level features within the same visual content, thus promoting the model for salient and subtle emotion awareness. To learn the distribution of emotions involved in images, most previous works learn coarse semantic knowledge with unbiased filtering. Consequently, they focus on the entire scene and suffer from the redundancy of semantic-irrelevant information, which diminishes the affective coherence, impeding the comprehension of emotional attributes within the treated features. In light of this, we reanalyze from the perspective of information filtering and propose a novel method called Multiple Feature Refining Network (MFRN). To minimize low-level feature redundancy, we design a wavelet-based separated frequency modeling, named Spectral Mixer, to learn invariant representations and enhance emotion saliency in low-level image features. At the higher semantic level, we design a Semantic Graph Prompt Learning for emotional semantic filtering, ensuring the purity of emotional information and providing the model with richer content semantics. Experiments conducted on three commonly used datasets have demonstrated the superiority of our MFRN model over cutting-edge methods.

NeurIPS Conference 2025 Conference Paper

On the SAC-BL Algorithm for Anomaly Detection

  • Xinsong Ma
  • Jie Wu
  • Weiwei Liu

Visual anomaly detection is significant in safety-critical and reliability-sensitive scenarios. Prior studies mainly emphasize the design and training of scoring functions, while little effort has been devoted to constructing decision rules based on these score functions. A recent work Ma et al. (2025b) highlights this issue and proposes the SAC-BL algorithm to address it. This method consists of a strong anomaly constraint (SAC) network and a betting-like (BL) algorithm serving as the decision rule. The SAC-BL algorithm can control the false discovery rate (FDR). However the performance of SAC-BL algorithm on anomalous examples, or its false positive rate (FPR), has not been thoroughly investigated. This paper provides a deeper analysis of this problem and explores how to theoretically reduce its FPR. First, we show that as the number of testing examples tends to infinity, the SAC-BL algorithm performs well on abnormal data if the scores follow the generalized Gaussian-like distribution family. But such conditions about the number of testing examples and the distribution of scores are overly restrictive for the real-world applications. So, we attempt to decrease the FPR of the SAC-BL algorithm under the condition of finite samples for practical anomaly detection. To this end, we redesign the BL algorithm by incorporating a randomization strategy and propose a novel stochastic BL (SBL) algorithm. The combination of the SAC network and the SBL algorithm yields our method, SAC-SBL. Theoretical results show that the SAC-SBL algorithm can achieve smaller FPR than SAC-BL algorithm while controlling its FDR. Finally, extensive experimental results demonstrate the superiority of our method over SAC-BL algorithm on multiple visual anomaly detection benchmarks.

NeurIPS Conference 2025 Conference Paper

PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning

  • Yizhen Zhang
  • Yang Ding
  • Shuoshuo Zhang
  • Xinchen Zhang
  • Haoling Li
  • Zhong-Zhi Li
  • Peijie Wang
  • Jie Wu

Inspired by the impressive reasoning capabilities demonstrated by reinforcement learning approaches like DeepSeek-R1, recent emerging research has begun exploring the use of reinforcement learning (RL) to enhance vision-language models (VLMs) for multimodal reasoning tasks. However, most existing multimodal reinforcement learning approaches remain limited to spatial reasoning within single-image contexts, yet still struggle to generalize to more complex and real-world scenarios involving multi-image positional reasoning, where understanding the relationships across images is crucial. To address this challenge, we propose a general reinforcement learning approach PeRL tailored for interleaved multimodal tasks, and a multi-stage strategy designed to enhance the exploration-exploitation trade-off, thereby improving learning efficiency and task performance. Specifically, we introduce permutation of image sequences to simulate varied positional relationships to explore more spatial and positional diversity. Furthermore, we design a rollout filtering mechanism for resampling to focus on trajectories that contribute most to learning optimal behaviors to exploit learned policies effectively. We evaluate our model on 5 widely-used multi-image benchmarks and 3 single-image benchmarks. Our experiments confirm that PeRL trained model consistently surpasses R1-related and interleaved VLM baselines by a large margin, achieving state-of-the-art performance on multi-image benchmarks, while preserving comparable performance on single-image tasks.

IJCAI Conference 2025 Conference Paper

RepObE: Representation Learning-Enhanced Obfuscation Encryption Modular Semantic Task Framework

  • Limei Lin
  • Jinpeng Xu
  • Xiaoding Wang
  • Liang Chen
  • Sun-Yuan Hsieh
  • Jie Wu

Model inversion and adversarial attacks in semantic communication pose risks, such as content leaks, alterations, and prediction inaccuracies, which threaten security and reliability. This paper introduces, from an attacker's viewpoint, a novel framework called RepObE (Representation Learning-Enhanced Obfuscation Encryption Modular Semantic Task Framework) to secure semantic communication. This framework employs dynamic encryption during semantic extraction and feature transmission to hinder attackers from reconstructing data through eavesdropping, thus strengthening system privacy. To combat image communication task challenges, we propose a prototype adversarial collaborative alignment training approach enhanced by representation learning. This method extracts and encodes semantic features while using dynamic perturbation and robust optimization to improve system resilience against adversarial threats. The approach ensures reliable semantic communication in complex environments, maintaining performance while countering attacks using feature obfuscation, adversarial training, and representation learning. Experimental results demonstrate that our method surpasses existing techniques by more than 2% in resisting model inversion attacks on classification tasks. Visually, our method excels with minimal decipherable images for attackers. It also shows a 3% to 5% improvement in countering adversarial attacks on classification tasks.

AAAI Conference 2025 Conference Paper

ResAdapter: Domain Consistent Resolution Adapter for Diffusion Models

  • Jiaxiang Cheng
  • Pan Xie
  • Xin Xia
  • Jiashi Li
  • Jie Wu
  • Yuxi Ren
  • Huixia Li
  • Xuefeng Xiao

Recent advancement in text-to-image models and corresponding personalized technologies enables individuals to generate high-quality and imaginative images. However, they often suffer from limitations when generating images with resolutions outside of their trained domain. To overcome this limitation, we present the resolution adapter \textbf{(ResAdapter)}, a domain-consistent adapter designed for diffusion models to generate images with unrestricted resolutions and aspect ratios. Unlike other multi-resolution generation methods that process images of static resolution with complex post-process operations, ResAdapter directly generates images with the dynamical resolution. Especially, after learning a deep understanding of pure resolution priors, ResAdapter trained on the general dataset, generates resolution-free images with personalized diffusion models while preserving their original style domain. Comprehensive experiments demonstrate that ResAdapter with only 0.5M can process images with flexible resolutions for arbitrary diffusion models. More extended experiments demonstrate that ResAdapter is compatible with other modules for image generation across a broad range of resolutions, and can be integrated into other multi-resolution model for efficiently generating higher-resolution images.

AILAW Journal 2025 Journal Article

Summarizing judicial documents: a hybrid extractive- abstractive model with legal domain knowledge

  • Yan Gao
  • Jie Wu
  • Zhengtao Liu
  • Juan Li

Abstract The automatic summarization of judgment documents is a challenging task due to their length and the dispersed nature of the important information they contain. The prevailing approach to tackling the summarization of lengthy documents involves the integration of both extractive and abstractive summarization models. However, current extractive models face challenges in capturing all essential details due to the scattered distribution of pertinent information within judgment documents. Additionally, the existing abstractive models still grapple with the problem of "hallucinations" which leads to generating inaccurate information. In our work, we proposed a novel hybrid legal summarization method that incorporates legal domain knowledge into both the extractive model and abstractive model. The method consists of two parts: (1) The rhetorical role of sentences is identified by the sentence-level sequence labeling method, and the rhetorical information is integrated into the extractive model based on WoBERT through the conditional normalization to ensure that the identification of key sentences is both precise and complete. (2) The pre-trained model RoFormer is combined with Seq2Seq to construct a long text summarization model, and the prior knowledge in the external resources and the document itself is introduced into the decoding process to improve the faithfulness and coherence of the composed summary. In addition, the contrastive learning strategy is employed during the training process to enhance the robustness of the abstractive model. Experimental results on the CAIL2020 dataset show that the proposed model is superior to the baseline methods. Furthermore, our method outperforms GPT and other LLMs in processing judgment documents.

AAAI Conference 2025 Conference Paper

Towards Multimodal Sentiment Analysis via Hierarchical Correlation Modeling with Semantic Distribution Constraints

  • Qinfu Xu
  • Yiwei Wei
  • Chunlei Wu
  • Leiquan Wang
  • Shaozu Yuan
  • Jie Wu
  • Jing Lu
  • Hengyang Zhou

Sentiment analysis is rapidly advancing by utilizing various data modalities (e.g., text, video, and audio). However, most existing techniques only learn the atomic-level features that reflect strong correlations, while ignoring more complex compositions in multimodal data. Moreover, they also neglected the incongruity in semantic distribution among modalities. In light of this, we introduce a novel Hierarchical Correlation Modeling Network (HCMNet), which enhances the multimodal sentiment analysis by exploring both the atomic-level correlations based on dynamic attention reasoning and the composition-level correlations through topological graph reasoning. In addition, we also alleviate the impact of distributional inconsistencies between modalities from both atomic-level and composition-level perspectives. Specifically, we first design an atomic-level contrastive loss that constrains the semantic distribution across modalities to mitigate the atomic-level inconsistency. Then, we design a graph optimal transport module that integrates transport flows with different graphs to constrain the composition-level semantic distribution, thus reducing the inconsistency of compositional nodes. Experiments on three public benchmark datasets have demonstrated the superiority of the proposed model over the state-of-the-art methods.

NeurIPS Conference 2024 Conference Paper

Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis

  • Yuxi Ren
  • Xin Xia
  • Yanzuo Lu
  • Jiacheng Zhang
  • Jie Wu
  • Pan Xie
  • Xing Wang
  • Xuefeng Xiao

Recently, a series of diffusion-aware distillation algorithms have emerged to alleviate the computational overhead associated with the multi-step inference process of Diffusion Models (DMs). Current distillation techniques often dichotomize into two distinct aspects: i) ODE Trajectory Preservation; and ii) ODE Trajectory Reformulation. However, these approaches suffer from severe performance degradation or domain shifts. To address these limitations, we propose Hyper-SD, a novel framework that synergistically amalgamates the advantages of ODE Trajectory Preservation and Reformulation, while maintaining near-lossless performance during step compression. Firstly, we introduce Trajectory Segmented Consistency Distillation to progressively perform consistent distillation within pre-defined time-step segments, which facilitates the preservation of the original ODE trajectory from a higher-order perspective. Secondly, we incorporate human feedback learning to boost the performance of the model in a low-step regime and mitigate the performance loss incurred by the distillation process. Thirdly, we integrate score distillation to further improve the low-step generation capability of the model and offer the first attempt to leverage a unified LoRA to support the inference process at all steps. Extensive experiments and user studies demonstrate that Hyper-SD achieves SOTA performance from 1 to 8 inference steps for both SDXL and SD1. 5. For example, Hyper-SDXL surpasses SDXL-Lightning by +0. 68 in CLIP Score and +0. 51 in Aes Score in the 1-step inference.

AAAI Conference 2024 Conference Paper

PrefAce: Face-Centric Pretraining with Self-Structure Aware Distillation

  • Siyuan Hu
  • Zheng Wang
  • Peng Hu
  • Xi Peng
  • Jie Wu
  • Hongyuan Zhu
  • Yew Soon Ong

Video-based facial analysis is important for autonomous agents to understand human expressions and sentiments. However, limited labeled data is available to learn effective facial representations. This paper proposes a novel self-supervised face-centric pretraining framework, called PrefAce, which learns transferable video facial representation without labels. The self-supervised learning is performed with an effective landmark-guided global-local tube distillation. Meanwhile, a novel instance-wise update FaceFeat Cache is built to enforce more discriminative and diverse representations for downstream tasks. Extensive experiments demonstrate that the proposed framework learns universal instance-aware facial representations with fine-grained landmark details from videos. The point is that it can transfer across various facial analysis tasks, e.g., Facial Attribute Recognition (FAR), Facial Expression Recognition (FER), DeepFake Detection (DFD), and Lip Synchronization (LS). Our framework also outperforms the state-of-the-art on various downstream tasks, even in low data regimes. Code is available at https://github.com/siyuan-h/PrefAce.

AAAI Conference 2024 Conference Paper

SyFormer: Structure-Guided Synergism Transformer for Large-Portion Image Inpainting

  • Jie Wu
  • Yuchao Feng
  • Honghui Xu
  • Chuanmeng Zhu
  • Jianwei Zheng

Image inpainting is in full bloom accompanied by the progress of convolutional neural networks (CNNs) and transformers, revolutionizing the practical management of abnormity disposal, image editing, etc. However, due to the ever-mounting image resolutions and missing areas, the challenges of distorted long-range dependencies from cluttered background distributions and reduced reference information in image domain inevitably rise, which further cause severe performance degradation. To address the challenges, we propose a novel large-portion image inpainting approach, namely the Structure-Guided Synergism Transformer (SyFormer), to rectify the discrepancies in feature representation and enrich the structural cues from limited reference. Specifically, we devise a dual-routing filtering module that employs a progressive filtering strategy to eliminate invalid noise interference and establish global-level texture correlations. Simultaneously, the structurally compact perception module maps an affinity matrix within the introduced structural priors from a structure-aware generator, assisting in matching and filling the corresponding patches of large-proportionally damaged images. Moreover, we carefully assemble the aforementioned modules to achieve feature complementarity. Finally, a feature decoding alignment scheme is introduced in the decoding process, which meticulously achieves texture amalgamation across hierarchical features. Extensive experiments are conducted on two publicly available datasets, i.e., CelebA-HQ and Places2, to qualitatively and quantitatively demonstrate the superiority of our model over state-of-the-arts.

NeurIPS Conference 2024 Conference Paper

UniFL: Improve Latent Diffusion Model via Unified Feedback Learning

  • Jiacheng Zhang
  • Jie Wu
  • Yuxi Ren
  • Xin Xia
  • Huafeng Kuang
  • Pan Xie
  • Jiashi Li
  • Xuefeng Xiao

Latent diffusion models (LDM) have revolutionized text-to-image generation, leading to the proliferation of various advanced models and diverse downstream applications. However, despite these significant advancements, current diffusion models still suffer from several limitations, including inferior visual quality, inadequate aesthetic appeal, and inefficient inference, without a comprehensive solution in sight. To address these challenges, we present UniFL, a unified framework that leverages feedback learning to enhance diffusion models comprehensively. UniFL stands out as a universal, effective, and generalizable solution applicable to various diffusion models, such as SD1. 5 and SDXL. Notably, UniFL consists of three key components: perceptual feedback learning, which enhances visual quality; decoupled feedback learning, which improves aesthetic appeal; and adversarial feedback learning, which accelerates inference. In-depth experiments and extensive user studies validate the superior performance of our method in enhancing generation quality and inference acceleration. For instance, UniFL surpasses ImageReward by 17\% user preference in terms of generation quality and outperforms LCM and SDXL Turbo by 57\% and 20\% general preference with 4-step inference.

JBHI Journal 2023 Journal Article

A Novel Deep Learning Model for Medical Report Generation by Inter-Intra Information Calibration

  • Junsan Zhang
  • Xiuxuan Shen
  • Shaohua Wan
  • Sotirios K. Goudos
  • Jie Wu
  • Ming Cheng
  • Weishan Zhang

Automatic generation of medical reports can provide diagnostic assistance to doctors and reduce their workload. To improve the quality of the generated medical reports, injecting auxiliary information through knowledge graphs or templates into the model is widely adopted in previous methods. However, they suffer from two problems: 1) The injected external information is limited in amount and difficult to adequately meet the information needs of medical report generation in content. 2) The injected external information increases the complexity of model and is hard to be reasonably integrated into the generation process of medical reports. Therefore, we propose an Information Calibrated Transformer (ICT) to address the above issues. First, we design a Precursor-information Enhancement Module (PEM), which can effectively extract numerous inter-intra report features from the datasets as the auxiliary information without external injection. And the auxiliary information can be dynamically updated with the training process. Secondly, a combination mode, which consists of PEM and our proposed Information Calibration Attention Module (ICA), is designed and embedded into ICT. In this method, the auxiliary information extracted from PEM is flexibly injected into ICT and the increment of model parameters is small. The comprehensive evaluations validate that the ICT is not only superior to previous methods in the X-Ray datasets, IU-X-Ray and MIMIC-CXR, but also successfully be extended to a CT COVID-19 dataset COV-CTR.

ECAI Conference 2023 Conference Paper

Tab-Attention: Self-Attention-Based Stacked Generalization for Imbalanced Credit Default Prediction

  • Yandan Tan
  • Hongbin Zhu
  • Jie Wu
  • Hongfeng Chai

Accurately credit default prediction faces challenges due to imbalanced data and low correlation between features and labels. Existing default prediction studies on the basis of gradient boosting decision trees (GBDT), deep learning techniques, and feature selection strategies can have varying degrees of success depending on the specific task. Motivated by this, we propose Tab-Attention, a novel self-attention-based stacked generalization method for credit default prediction. This approach ensembles the potential proprietary knowledge contributions from multi-view feature spaces, to cope with low feature correlation and imbalance. We organize multi-view feature spaces according to the latent linear or nonlinear strengths between features and labels. Meanwhile, the f1 score assists the model in imbalance training to find the optimal state for identifying minority default samples. Our Tab-Attention achieves superior Recall1 and f11 of default intention recognition than existing GBDT-based models and advanced deep learning by about 32. 92% and 16. 05% on average, respectively, while maintaining outstanding overall performance and prediction performance for non-default samples. The proposed method could ensemble essential knowledge through the self-attention mechanism, which is of great significance for a more robust future prediction system.

AAAI Conference 2022 Conference Paper

Activation Modulation and Recalibration Scheme for Weakly Supervised Semantic Segmentation

  • Jie Qin
  • Jie Wu
  • Xuefeng Xiao
  • Lujun Li
  • Xingang Wang

Image-level weakly supervised semantic segmentation (WSSS) is a fundamental yet challenging computer vision task facilitating scene understanding and automatic driving. Most existing methods resort to classification-based Class Activation Maps (CAMs) to play as the initial pseudo labels, which tend to focus on the discriminative image regions and lack customized characteristics for the segmentation task. To alleviate this issue, we propose a novel activation modulation and recalibration (AMR) scheme, which leverages a spotlight branch and a compensation branch to obtain weighted CAMs that can provide recalibration supervision and task-specific concepts. Specifically, an attention modulation module (AMM) is employed to rearrange the distribution of feature importance from the channel-spatial sequential perspective, which helps to explicitly model channelwise interdependencies and spatial encodings to adaptively modulate segmentation-oriented activation responses. Furthermore, we introduce a cross pseudo supervision for dual branches, which can be regarded as a semantic similar regularization to mutually refine two branches. Extensive experiments show that AMR establishes a new state-of-the-art performance on the PASCAL VOC 2012 dataset, surpassing not only current methods trained with the image-level of supervision but also some methods relying on stronger supervision, such as saliency label. Experiments also reveal that our scheme is plug-and-play and can be incorporated with other approaches to boost their performance. Our code is available at: https: //github. com/jieqin-ai/AMR.

AAAI Conference 2022 Conference Paper

GraphMemDialog: Optimizing End-to-End Task-Oriented Dialog Systems Using Graph Memory Networks

  • Jie Wu
  • Ian G Harris
  • Hongzhi Zhao

Effectively integrating knowledge into end-to-end taskoriented dialog systems remains a challenge. It typically requires incorporation of an external knowledge base (KB) and capture of the intrinsic semantics of the dialog history. Recent research shows promising results by using Sequence-to- Sequence models, Memory Networks, and even Graph Convolutional Networks. However, current state-of-the-art models are less effective at integrating dialog history and KB into task-oriented dialog systems in the following ways: 1. The KB representation is not fully context-aware. The dynamic interaction between the dialog history and KB is seldom explored. 2. Both the sequential and structural information in the dialog history can contribute to capturing the dialog semantics, but they are not studied concurrently. In this paper, we propose a novel Graph Memory Network (GMN) based Seq2Seq model, GraphMemDialog, to effectively learn the inherent structural information hidden in dialog history, and to model the dynamic interaction between dialog history and KBs. We adopt a modified graph attention network to learn the rich structural representation of the dialog history, whereas the context-aware representation of KB entities are learnt by our novel GMN. To fully exploit this dynamic interaction, we design a learnable memory controller coupled with external KB entity memories to recurrently incorporate dialog history context into KB entities through a multi-hop reasoning mechanism. Experiments on three public datasets show that our GraphMemDialog model achieves state-of-theart performance and outperforms strong baselines by a large margin, especially on datatests with more complicated KB information.

NeurIPS Conference 2021 Conference Paper

Revisiting Discriminator in GAN Compression: A Generator-discriminator Cooperative Compression Scheme

  • Shaojie Li
  • Jie Wu
  • Xuefeng Xiao
  • Fei Chao
  • Xudong Mao
  • Rongrong Ji

Recently, a series of algorithms have been explored for GAN compression, which aims to reduce tremendous computational overhead and memory usages when deploying GANs on resource-constrained edge devices. However, most of the existing GAN compression work only focuses on how to compress the generator, while fails to take the discriminator into account. In this work, we revisit the role of discriminator in GAN compression and design a novel generator-discriminator cooperative compression scheme for GAN compression, termed GCC. Within GCC, a selective activation discriminator automatically selects and activates convolutional channels according to a local capacity constraint and a global coordination constraint, which help maintain the Nash equilibrium with the lightweight generator during the adversarial training and avoid mode collapse. The original generator and discriminator are also optimized from scratch, to play as a teacher model to progressively refine the pruned generator and the selective activation discriminator. A novel online collaborative distillation scheme is designed to take full advantage of the intermediate feature of the teacher generator and discriminator to further boost the performance of the lightweight generator. Extensive experiments on various GAN-based generation tasks demonstrate the effectiveness and generalization of GCC. Among them, GCC contributes to reducing 80% computational costs while maintains comparable performance in image translation tasks.

IJCAI Conference 2021 Conference Paper

Weakly-Supervised Spatio-Temporal Anomaly Detection in Surveillance Video

  • Jie Wu
  • Wei Zhang
  • Guanbin Li
  • Wenhao Wu
  • Xiao Tan
  • Yingying Li
  • Errui Ding
  • Liang Lin

In this paper, we introduce a novel task, referred to as Weakly-Supervised Spatio-Temporal Anomaly Detection (WSSTAD) in surveillance video. Specifically, given an untrimmed video, WSSTAD aims to localize a spatio-temporal tube (i. e. , a sequence of bounding boxes at consecutive times) that encloses the abnormal event, with only coarse video-level annotations as supervision during training. To address this challenging task, we propose a dual-branch network which takes as input the proposals with multi-granularities in both spatial-temporal domains. Each branch employs a relationship reasoning module to capture the correlation between tubes/videolets, which can provide rich contextual information and complex entity relationships for the concept learning of abnormal behaviors. Mutually-guided Progressive Refinement framework is set up to employ dual-path mutual guidance in a recurrent manner, iteratively sharing auxiliary supervision information across branches. It impels the learned concepts of each branch to serve as a guide for its counterpart, which progressively refines the corresponding branch and the whole framework. Furthermore, we contribute two datasets, i. e. , ST-UCF-Crime and STRA, consisting of videos containing spatio-temporal abnormal annotations to serve as the benchmarks for WSSTAD. We conduct extensive qualitative and quantitative evaluations to demonstrate the effectiveness of the proposed approach and analyze the key factors that contribute more to handle this task.

NeurIPS Conference 2020 Conference Paper

Robust Sequence Submodular Maximization

  • Gamal Sallam
  • Zizhan Zheng
  • Jie Wu
  • Bo Ji

Submodularity is an important property of set functions and has been extensively studied in the literature. It models set functions that exhibit a diminishing returns property, where the marginal value of adding an element to a set decreases as the set expands. This notion has been generalized to considering sequence functions, where the order of adding elements plays a crucial role and determines the function value; the generalized notion is called sequence (or string) submodularity. In this paper, we study a new problem of robust sequence submodular maximization with cardinality constraints. The robustness is against the removal of a subset of elements in the selected sequence (e. g. , due to malfunctions or adversarial attacks). Compared to robust submodular maximization for set function, new challenges arise when sequence functions are concerned. Specifically, there are multiple definitions of submodularity for sequence functions, which exhibit subtle yet critical differences. Another challenge comes from two directions of monotonicity: forward monotonicity and backward monotonicity, both of which are important to proving performance guarantees. To address these unique challenges, we design two robust greedy algorithms: while one algorithm achieves a constant approximation ratio but is robust only against the removal of a subset of contiguous elements, the other is robust against the removal of an arbitrary subset of the selected elements but requires a stronger assumption and achieves an approximation ratio that depends on the number of the removed elements. Finally, we generalize the analyses to considering sequence functions under weaker assumptions based on approximate versions of sequence submodularity and backward monotonicity.

AAAI Conference 2020 Conference Paper

Tree-Structured Policy Based Progressive Reinforcement Learning for Temporally Language Grounding in Video

  • Jie Wu
  • Guanbin Li
  • Si Liu
  • Liang Lin

Temporally language grounding in untrimmed videos is a newly-raised task in video understanding. Most of the existing methods suffer from inferior efficiency, lacking interpretability, and deviating from the human perception mechanism. Inspired by human’s coarse-to-fine decision-making paradigm, we formulate a novel Tree-Structured Policy based Progressive Reinforcement Learning (TSP-PRL) framework to sequentially regulate the temporal boundary by an iterative refinement process. The semantic concepts are explicitly represented as the branches in the policy, which contributes to efficiently decomposing complex policies into an interpretable primitive action. Progressive reinforcement learning provides correct credit assignment via two task-oriented rewards that encourage mutual promotion within the treestructured policy. We extensively evaluate TSP-PRL on the Charades-STA and ActivityNet datasets, and experimental results show that TSP-PRL achieves competitive performance over existing state-of-the-art methods.

EAAI Journal 2019 Journal Article

A self-adaptive approach to service deployment under mobile edge computing for autonomous driving

  • Wei Xiong
  • Zhihui Lu
  • Bing Li
  • Zhao Wu
  • Bo Hang
  • Jie Wu
  • Xiaohua Xuan

Mobile edge computing for autonomous driving needs to manage heterogeneous resources and process large amounts of data or multi-purpose payload. There needs to be deploying, scheduling and migrating tasks on edge nodes to ensure the reliability of tasks or maximize the utilization of resources. However, applying autonomous learning methods on autonomous driving is exceptionally difficult, due to the complexity of multi-dimensional context and the sensitivity to hyperparameters. In this paper, we propose a learning approach to quality-of-service (QoS) prediction of services via multi-dimensional context, and develop a stable approach for service deployment that requires minimal hyperparameter tuning and a modest number of trials to learn multilayer neural network policies. This approach can automatically trades off exploration against exploitation by automatically tuning hyperparameter based on maximum entropy reinforcement learning. We then demonstrate that this approach achieves state-of-the-art performance on Autoware benchmark environments.

EAAI Journal 2019 Journal Article

Data Privacy Protection for Edge Computing of Smart City in a DIKW Architecture

  • Yucong Duan
  • Zhihui Lu
  • Zhangbing Zhou
  • Xiaobing Sun
  • Jie Wu

Current trend of shifting computing from centralized Cloud to Edge has not only empowered huge amount of IoT devices with the capability of the accumulation of individualized computing and storage as a flexible whole, but also brought along new privacy challenges originating in emerging new usage requests on the accumulated content or resources from multiple sources of various integrated devices at the Edge. In this work, we focus on modeling the privacy content of multiple sources through mapping them as resources of types of Data, Information and Knowledge in the well-known DIKW architecture. We propose to categorize content objects and relationships uniformly as typed resources of data, information, and knowledge, according to our formalized DIKW architecture composing a meta model of DIKW and extended Data Graph, Information Graph and Knowledge Graph. We further propose to categorize target privacy resources of data and information according to their presence in the modeled searching space in our DIKW architecture as explicit and implicit divisions. Thereafter we propose protection solutions according to explicit and implicit divisions for privacy target concerning typed data. The efficiency and performance potential of our processing solution originates in a multiple dimensional modeling strategy of typed data, which is modeled solely with various frequencies of various meta level dimensions.

TAAS Journal 2017 Journal Article

e-Sampling

  • Md Zakirul Alam Bhuiyan
  • Jie Wu
  • Guojun Wang
  • Tian Wang
  • Mohammad Mehedi Hassan

Sampling rate adaptation is a critical issue in many resource-constrained networked systems, including Wireless Sensor Networks (WSNs). Existing algorithms are primarily employed to detect events such as objects or physical changes at a high, low, or fixed frequency sampling usually adapted by a central unit or a sink, therefore requiring additional resource usage. Additionally, this algorithm potentially makes a network unable to capture a dynamic change or event of interest, which therefore affects monitoring quality. This article studies the problem of a fully autonomous adaptive sampling regarding the presence of a change or event. We propose a novel scheme, termed “event-sensitive adaptive sampling and low-cost monitoring (e-Sampling)” by addressing the problem in two stages, which leads to reduced resource usage (e.g., energy, radio bandwidth). First, e-Sampling provides the embedded algorithm to adaptive sampling that automatically switches between high- and low-frequency intervals to reduce the resource usage, while minimizing false negative detections. Second, by analyzing the frequency content, e-Sampling presents an event identification algorithm suitable for decentralized computing in resource-constrained networks. In the absence of an event, the “uninteresting” data is not transmitted to the sink. Thus, the energy cost is further reduced. e-Sampling can be useful in a broad range of applications. We apply e-Sampling to Structural Health Monitoring (SHM) and Fire Event Monitoring (FEM), which are typical applications of high-frequency events. Evaluation via both simulations and experiments validates the advantages of e-Sampling in low-cost event monitoring, and in effectively expanding the capacity of WSNs for high data rate applications.

IJCAI Conference 2015 Conference Paper

Offline Sketch Parsing via Shapeness Estimation

  • Jie Wu
  • Changhu Wang
  • Liqing Zhang
  • Yong Rui

In this work, we target at the problem of offline sketch parsing, in which the temporal orders of strokes are unavailable. It is more challenging than most of existing work, which usually leverages the temporal information to reduce the search space. Different from traditional approaches in which thousands of candidate groups are selected for recognition, we propose the idea of shapeness estimation to greatly reduce this number in a very fast way. Based on the observation that most of hand-drawn shapes with well-defined closed boundaries can be clearly differentiated from nonshapes if normalized into a very small size, we propose an efficient shapeness estimation method. A compact feature representation as well as its efficient extraction method is also proposed to speed up this process. Based on the proposed shapeness estimation, we present a three-stage cascade framework for offline sketch parsing. The shapeness estimation technique in this framework greatly reduces the number of false positives, resulting in a 96. 2% detection rate with only 32 candidate group proposals, which is two orders of magnitude less than existing methods. Extensive experiments show the superiority of the proposed framework over stateof-the-art works on sketch parsing in both effectiveness and efficiency, even though they leveraged the temporal information of strokes.

TCS Journal 2015 Journal Article

Randomized oblivious integral routing for minimizing power cost

  • Yangguang Shi
  • Fa Zhang
  • Jie Wu
  • Zhiyong Liu

Given an undirected network G ( V, E ) and a set of traffic requests R, the minimum power-cost routing problem requires that each R k ∈ R be routed along a single path to minimize ∑ e ∈ E ( l e ) α, where l e is the traffic load on edge e and α is a constant greater than 1. Typically, α ∈ ( 1, 3 ]. This problem is important in optimizing the energy consumption of networks. To address this problem, we propose a randomized oblivious routing algorithm. An oblivious routing algorithm makes decisions independently of the current traffic in the network. This feature enables the efficient implementation of our algorithm in a distributed manner, which is desirable for large-scale high-capacity networks. An important feature of our work is that our algorithm can satisfy the integral constraint, which requires that each traffic request R k should follow a single path. We prove that, given this constraint, no randomized oblivious routing algorithm can guarantee a competitive ratio bounded by o ( | E | α − 1 α + 1 ). By contrast, our approach provides a competitive ratio of O ( | E | α − 1 α + 1 log 2 α α + 1 ⁡ | V | ⋅ log α − 1 ⁡ D ), where D is the maximum demand of traffic requests. Furthermore, our results also hold for a more general case where the objective is to minimize ∑ e ( l e ) p, where p ≥ 1 is an arbitrary unknown parameter with a given upper bound α > 1. The theoretical results established in proving these bounds can be further generalized to a framework of designing and analyzing oblivious integral routing algorithms, which is significant for research on minimizing ∑ e ( l e ) α in specific scenarios with simplified problem settings. For instance, we prove that this framework can generate an oblivious integral routing algorithm whose competitive ratio can be bounded by O ( log α ⁡ | V | ⋅ log α − 1 ⁡ D ) and O ( log 3 α ⁡ | V | ⋅ log α − 1 ⁡ D ) on expanders and hypercubes, respectively.

AAAI Conference 2014 Conference Paper

Sketch Recognition with Natural Correction and Editing

  • Jie Wu
  • Changhu Wang
  • Liqing Zhang
  • Yong Rui

In this paper, we target at the problem of sketch recognition. We systematically study how to incorporate users’ correction and editing into isolated and full sketch recognition. This is a natural and necessary interaction in real systems such as Visio where very similar shapes exist. First, a novel algorithm is proposed to mine the prior shape knowledge for three editing modes. Second, to differentiate visually similar shapes, a novel symbol recognition algorithm is introduced by leveraging the learnt shape knowledge. Then, a novel editing detection algorithm is proposed to facilitate symbol recognition. Furthermore, both of the symbol recognizer and the editing detector are systematically incorporated into the full sketch recognition. Finally, based on the proposed algorithms, a realtime sketch recognition system is built to recognize handdrawn flowcharts and diagrams with flexible interactions. Extensive experiments show the effectiveness of the proposed algorithms.

v2026.09.13