Arrow Research search

Author name cluster

Peng Wu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

24 papers
2 author rows

Possible papers

24

AAAI Conference 2026 Conference Paper

MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation

  • Run Ling
  • Ke Cao
  • Jian Lu
  • Ao Ma
  • Haowei Liu
  • Runze He
  • Changwei Wang
  • Rongtao Xu

Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural generation, and permutation sensitivity, where the order of reference inputs causes subject distortion. In this paper, we propose MoFu, a unified framework that tackles both challenges. For scale inconsistency, we introduce Scale-Aware Modulation (SMO), an LLM-guided module that extracts implicit scale cues from the prompt and modulates features to ensure consistent subject sizes. To address permutation sensitivity, we present a simple yet effective Fourier Fusion strategy that processes the frequency information of reference features via the Fast Fourier Transform to produce a unified representation. Besides, we design a Scale-Permutation Stability Loss to jointly encourage scale-consistent and permutation-invariant generation. To further evaluate these challenges, we establish a dedicated benchmark with controlled variations in subject scale and reference permutation. Extensive experiments demonstrate that MoFu significantly outperforms existing methods in preserving natural scale, subject fidelity, and overall visual quality.

AAAI Conference 2026 Conference Paper

TargetVAU: Multimodal Anomaly-Aware Reasoning for Target Behavior Understanding in Videos

  • Lingru Zhou
  • Peng Wu
  • Manqing Zhang
  • Qingsheng Wang
  • Guansong Pang
  • Peng Wang

Understanding anomalous human behaviors at a fine-grained level remains a major challenge in complex scenarios. Existing video anomaly understanding (VAU) methods often rely on coarse frame-level cues or overlook structured modeling of individual actions, limiting their capacity for reasoning about human interactions and accountability. To address these challenges, we propose TargetVAU, a multimodal anomaly-aware reasoning framework designed for individual-level anomaly recognition and explanation. TargetVAU first extracts both global-level and human-centric visual features using a frozen Vision Transformer (ViT) encoder. An Anomaly-focused Temporal Sampler is then employed to select behaviorally informative frames via a density-aware strategy guided by predicted anomaly scores. A Spatio-Temporal Interaction Graph is constructed to explicitly model interactions among individuals across time and space. These structured representations are fused with prompt embeddings via a frozen Q-Former to form a unified semantic representation. Finally, a large language model fine-tuned with low-rank adaptation (LoRA) performs instruction-guided reasoning to identify anomalous individuals and generate natural language explanations. Extensive experiments on UCCD and HIVAU-70K demonstrate that TargetVAU significantly outperforms existing methods in both accuracy and interpretability, advancing the state of individual-level anomaly understanding in surveillance videos.

NeurIPS Conference 2025 Conference Paper

Adaptive Data-Borrowing for Improving Treatment Effect Estimation using External Controls

  • Qinwei Yang
  • Jingyi Li
  • Peng Wu

Randomized controlled trials (RCTs) often exhibit limited inferential efficiency in estimating treatment effects due to small sample sizes. In recent years, the combination of external controls has gained increasing attention as a means of improving the efficiency of RCTs. However, external controls are not always comparable to RCTs, and direct borrowing without careful evaluation can introduce substantial bias and reduce the efficiency of treatment effect estimation. In this paper, we propose a novel influence-based adaptive sample borrowing approach that effectively quantifies the "comparability'' of each sample in the external controls using influence function theory. Given a selected set of borrowed external controls, we further derive a semiparametric efficient estimator under an exchangeability assumption. Recognizing that the exchangeability assumption may not hold for all possible borrowing sets, we conduct a detailed analysis of the asymptotic bias and variance of the proposed estimator under violations of exchangeability. Building on this bias-variance trade-off, we further develop a data-driven approach to select the optimal subset of external controls for borrowing. Extensive simulations and real-world applications demonstrate that the proposed approach significantly enhances treatment effect estimation efficiency in RCTs, outperforming existing approaches.

IROS Conference 2025 Conference Paper

All-in-one Defensive Network (ADNet): Trustworthy Segmentation of Complex Maritime Environments for Unmanned Surface Vessels (USVs)

  • Yanhong Huang
  • Yuze Duan
  • Peng Wu
  • Yuanchang Liu

The visual perception system of unmanned surface vessels (USVs) is often subjected to various adversarial attacks (e. g. , lens stains, sun glare, ship painting, etc.), impacting the safety of autonomous navigation in maritime environments. To enhance the reliability and robustness of situational awareness in complex environments, we proposed a defensive model to effectively counteract multiple attacks targeting the perception system. Specifically, we first constructed a maritime instance segmentation dataset including various adversarial attack samples, with accurate annotations for the sky, water, land, ships and obstacles. To address the degradation in perception accuracy caused by adversarial attacks, we introduced a Monte Carlo-based random fusion module (MC Fusion) to enhance the adaptability of USVs in various dynamic environments. Additionally, as USVs are always equipped with onboard PC with limited computing resources, we incorporated the lightweight universal inverted bottleneck (UIB) module into the backbone to ensure effective feature extraction while reducing model parameters. Finally, we conducted comparative experiments under various adversarial attack scenarios. Our results demonstrate that, even in the presence of multiple adversarial attacks, our method improves ship detection accuracy by 13. 9% and increases the mean accuracy of segmentation masks by over 10% compared to state-of-the-art models, enhancing the safety of USVs in navigation. The source code and datasets are available at https://github.com/huangyanh/ADNet.

NeurIPS Conference 2025 Conference Paper

Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment

  • Pengfei Zhao
  • Rongbo Luan
  • Wei Zhang
  • Peng Wu
  • Sifeng He

Despite Contrastive Language–Image Pre-training (CLIP)'s remarkable capability to retrieve content across modalities, a substantial modality gap persists in its feature space. Intriguingly, we discover that off-the-shelf MLLMs (Multimodal Large Language Models) demonstrate powerful inherent modality alignment properties. While recent MLLM-based retrievers with unified architectures partially mitigate this gap, their reliance on coarse modality alignment mechanisms fundamentally limits their potential. In this work, We introduce MAPLE (Modality-Aligned Preference Learning for Embeddings), a novel framework that leverages the fine-grained alignment priors inherent in MLLM to guide cross-modal representation learning. MAPLE formulates the learning process as reinforcement learning with two key components: (1) Automatic preference data construction using off-the-shelf MLLM, and (2) a new Relative Preference Alignment (RPA) loss, which adapts Direct Preference Optimization (DPO) to the embedding learning setting. Experimental results show that our preference-guided alignment achieves substantial gains in fine-grained cross-modal retrieval, underscoring its effectiveness in handling nuanced semantic distinctions.

EAAI Journal 2025 Journal Article

Hybrid and multiple ensemble metamodel-based evaluation for operating tunnel performance in three-dimensional spatially variable soils

  • Ning Tian
  • Jinsong Huang
  • Jian Chen
  • Kaiwei Tian
  • Peng Wu

In recent years, the Random Finite Element Method (RFEM) has gained prominence in geotechnical engineering for assessing the inherent spatial variability in the mechanical properties of both natural and processed soils. Nevertheless, RFEM often demands more extensive computational resources than deterministic finite element analysis, as it is coupled with Monte-Carlo simulations (MCS). To mitigate this computational burden, metamodeling techniques have emerged as a popular approach. This paper proposes a novel and hybrid Support Vector Regression (SVR) metamodel by fusing the RFEM analysis. The metamodel can efficiently generate the original finite element method predicted quantities with limited training by utilizing input random field features, which encapsulate high-dimensional information pertaining to spatially variable soil stiffness parameters. Furthermore, based on ensemble learning, the Bagging and Adaboost algorithms were used to develop a multiple SVR (M-SVR) ensemble learning metamodel to enhance prediction reliability. Simultaneously, considering the limitation that machine learning prediction can only provide a single value, the prediction results with confidence intervals based on Bagging ensemble algorithms were also developed to quantify the uncertainty of machine learning predictions in regression analysis. The consistency between SVR and M-SVR predictions and RFEM calculations is demonstrated through a problem involving the failure probability evaluation of tunnel longitudinal performance induced by ground surface surcharge in three-dimensional spatially variable soils. The substantial improvement in efficiency with the adoption of the SVR and M-SVR, as compared to RFEM, underscores the immense potential of machine learning algorithms in conducting geotechnical reliability analyses involved with spatial variability.

NeurIPS Conference 2025 Conference Paper

Learning Counterfactual Outcomes Under Rank Preservation

  • Peng Wu
  • Haoxuan Li
  • Chunyuan Zheng
  • Yan Zeng
  • Jiawei Chen
  • Yang Liu
  • Ruocheng Guo
  • Kun Zhang

Counterfactual inference aims to estimate the counterfactual outcome at the individual level given knowledge of an observed treatment and the factual outcome, with broad applications in fields such as epidemiology, econometrics, and management science. Previous methods rely on a known structural causal model (SCM) or assume the homogeneity of the exogenous variable and strict monotonicity between the outcome and exogenous variable. In this paper, we propose a principled approach for identifying and estimating the counterfactual outcome. We first introduce a simple and intuitive rank preservation assumption to identify the counterfactual outcome without relying on a known structural causal model. Building on this, we propose a novel ideal loss for theoretically unbiased learning of the counterfactual outcome and further develop a kernel-based estimator for its empirical estimation. Our theoretical analysis shows that the rank preservation assumption is not stronger than the homogeneity and strict monotonicity assumptions, and shows that the proposed ideal loss is convex, and the proposed estimator is unbiased. Extensive semi-synthetic and real-world experiments are conducted to demonstrate the effectiveness of the proposed method.

IJCAI Conference 2025 Conference Paper

Optimal Policy Adaptation Under Covariate Shift

  • Xueqing Liu
  • Qinwei Yang
  • Zhaoqing Tian
  • Ruocheng Guo
  • Peng Wu

Transfer learning of prediction models has been extensively studied, while the corresponding policy learning approaches are rarely discussed. In this paper, we propose principled approaches for learning the optimal policy in the target domain by leveraging two datasets: one with full information from the source domain and the other from the target domain with only covariates. First, in the setting of covariate shift, we formulate the problem from a perspective of causality and present the identifiability assumptions for the reward induced by a given policy. Then, we derive the efficient influence function and the semiparametric efficiency bound for the reward. Based on this, we construct a doubly robust and semiparametric efficient estimator for the reward and then learn the optimal policy by optimizing the estimated reward. Moreover, we theoretically analyze the bias and the generalization error bound for the learned policy. Furthermore, in the presence of both covariate and concept shifts, we propose a novel sensitivity analysis method to evaluate the robustness of the proposed policy learning approach. Extensive experiments demonstrate that the approach not only estimates the reward more accurately but also yields a policy that closely approximates the theoretically optimal policy.

AAAI Conference 2025 Conference Paper

VarCMP: Adapting Cross-Modal Pre-Training Models for Video Anomaly Retrieval

  • Peng Wu
  • Wanshun Su
  • Xiangteng He
  • Peng Wang
  • Yanning Zhang

Video anomaly retrieval (VAR) aims to retrieve pertinent abnormal or normal videos from collections of untrimmed and long videos through cross-modal requires such as textual descriptions and synchronized audios. Cross-modal pre-training (CMP) models, by pre-training on large-scale cross-modal pairs, e.g., image and text, can learn the rich associations between different modalities, and this cross-modal association capability gives CMP an advantage in conventional retrieval tasks. Inspired by this, how to utilize the robust cross-modal association capabilities of CMP in VAR to search crucial visual component from these untrimmed and long videos becomes a critical research problem. Therefore, this paper proposes a VAR method based on CMP models, named VarCMP. First, a unified hierarchical alignment strategy is proposed to constrain the semantic and spatial consistency between video and text, as well as the semantic, temporal, and spatial consistency between video and audio. It fully leverages the efficient cross-modal association capabilities of CMP models by considering cross-modal similarities at multiple granularities, enabling VarCMP to achieve effective all-round information matching for both video-text and video-audio VAR tasks. Moreover, to further solve the problem of untrimmed and long video alignment, an anomaly-biased weighting is devised in the fine-grained alignment, which identifies key segments in untrimmed long videos using anomaly priors, giving them more attention, thereby discarding irrelevant segment information, and achieving more accurate matching with cross-modal queries. Extensive experiments demonstrates high efficacy of VarCMP in both video-text and video-audio VAR tasks, achieving significant improvements on both text-video (UCFCrime-AR) and audio-video (XDViolence-AR) datasets against the best competitors by 5.0% and 5.3% R@1.

EAAI Journal 2024 Journal Article

An adaptive few-shot fault diagnosis method based on virtual samples generated by fault characteristics of rotating machines

  • Peng Wu
  • Gongye Yu
  • Qianqian Yu
  • Pengqi Wang
  • Yongming Han
  • Bo Ma

The existing intelligent diagnostic methods based on the machine learning achieve good diagnostic results under the condition of large amount of the failure data available. However, in most cases, the diagnosis model is difficult to be constructed under the few-shot problem, which only normal data of the equipment is available. Therefore, a novel Adaptive Diagnosis with Fault Characteristics (ADFC) method, which builds the diagnosis model by generating personalized virtual fault samples, is proposed in the paper. The fault common characteristics are obtained based on the failure mechanism and the law of transmission paths, and the personalized fault virtual samples are generated by combining health state data contain the individual characteristics of the equipment. Then, the frequency domain features reflecting the operating status of the equipment are extracted as training samples. Finally, the fault diagnosis model based on the Convolutional Neural Network (CNN) is constructed for the equipment. The ADFC method is validated by the rotating machinery data, including the public data, the laboratory data and the field application data. The results indicate that the ADFC method achieved an average diagnostic accuracy of 92. 16%, and the accuracy has been improved by at least 0. 78% compared to the comparison method.

NeurIPS Conference 2024 Conference Paper

Learning the Optimal Policy for Balancing Short-Term and Long-Term Rewards

  • Qinwei Yang
  • Xueqing Liu
  • Yan Zeng
  • Ruocheng Guo
  • Yang Liu
  • Peng Wu

Learning the optimal policy to balance multiple short-term and long-term rewards has extensive applications across various domains. Yet, there is a noticeable scarcity of research addressing policy learning strategies in this context. In this paper, we aim to learn the optimal policy capable of effectively balancing multiple short-term and long-term rewards, especially in scenarios where the long-term outcomes are often missing due to data collection challenges over extended periods. Towards this goal, the conventional linear weighting method, which aggregates multiple rewards into a single surrogate reward through weighted summation, can only achieve sub-optimal policies when multiple rewards are related. Motivated by this, we propose a novel decomposition-based policy learning (DPPL) method that converts the whole problem into subproblems. The DPPL method is capable of obtaining optimal policies even when multiple rewards are interrelated. Nevertheless, the DPPL method requires a set of preference vectors specified in advance, posing challenges in practical applications where selecting suitable preferences is non-trivial. To mitigate this, we further theoretically transform the optimization problem in DPPL into an $\varepsilon$-constraint problem, where $\varepsilon$ represents the minimum acceptable levels of other rewards while maximizing one reward. This transformation provides intuitive into the selection of preference vectors. Extensive experiments are conducted on the proposed method and the results validate the effectiveness of the method.

AAAI Conference 2024 Conference Paper

VadCLIP: Adapting Vision-Language Models for Weakly Supervised Video Anomaly Detection

  • Peng Wu
  • Xuerong Zhou
  • Guansong Pang
  • Lingru Zhou
  • Qingsen Yan
  • Peng Wang
  • Yanning Zhang

The recent contrastive language-image pre-training (CLIP) model has shown great success in a wide range of image-level tasks, revealing remarkable ability for learning powerful visual representations with rich semantics. An open and worthwhile problem is efficiently adapting such a strong model to the video domain and designing a robust video anomaly detector. In this work, we propose VadCLIP, a new paradigm for weakly supervised video anomaly detection (WSVAD) by leveraging the frozen CLIP model directly without any pre-training and fine-tuning process. Unlike current works that directly feed extracted features into the weakly supervised classifier for frame-level binary classification, VadCLIP makes full use of fine-grained associations between vision and language on the strength of CLIP and involves dual branch. One branch simply utilizes visual features for coarse-grained binary classification, while the other fully leverages the fine-grained language-image alignment. With the benefit of dual branch, VadCLIP achieves both coarse-grained and fine-grained video anomaly detection by transferring pre-trained knowledge from CLIP to WSVAD task. We conduct extensive experiments on two commonly-used benchmarks, demonstrating that VadCLIP achieves the best performance on both coarse-grained and fine-grained WSVAD, surpassing the state-of-the-art methods by a large margin. Specifically, VadCLIP achieves 84.51% AP and 88.02% AUC on XD-Violence and UCF-Crime, respectively. Code and features are released at https://github.com/nwpu-zxr/VadCLIP.

EAAI Journal 2023 Journal Article

A data-driven distributionally newsvendor problem for edge-cloud collaboration in intelligent manufacturing systems

  • Cheng-hu Yang
  • Xiao-li Su
  • Peng Wu

In intelligent manufacturing systems, the industrial informatics has features of multi-source, multi-noise, and time series. It is difficult for small and medium enterprises (SMEs) to directly exploit the enormous amounts of data due to the limited budgets and computing capabilities. Edge intelligence is a key technique to power intelligent manufacturing systems and provide knowledge transferred from the cloud to SMEs at the edges. To address edge-cloud collaboration issue, we propose a refined data-driven distributionally robust newsvendor model based on φ-divergence measures and imprecise Dirichlet models (DRN-IDM). We construct new distributional uncertainty sets by effectively integrating local censored demand data and cloud knowledge, which helps SMEs to make intelligent production decisions and reduce significant decision deviations, even under a small censored data set. In particular, the novel demand uncertainty sets can depict the distance between distributions and probability intervals. Then, we transform the DRN-IDM model into a convex optimization model that is amenable to algorithmic implementation. Additionally, based on the coefficient of variation of limited historical data, we propose an adaptive demand information fusion procedure to achieve excellent synergy effect from cloud knowledge. We also validate the effectiveness of the DRN-IDM model and the practicability of adaptive procedure using extensive numerical studies with both simulated and real-life data. Furthermore, we measure the relative expected value of cloud knowledge and investigate the effect of censored demand samples. Our results verify the effectiveness condition of the DRN-IDM model and indicate that cloud knowledge can improve the precision and robustness of SMEs’ production decisions with small-scale censored data. Interestingly, the verified adaptive procedure can be applied in the learning criteria design of metaheuristics in intelligent manufacturing systems, and the reconstructed uncertainty set can narrow the search space to improve the convergence performance of algorithms.

AAAI Conference 2023 Conference Paper

Accurate Fairness: Improving Individual Fairness without Trading Accuracy

  • Xuran Li
  • Peng Wu
  • Jing Su

Accuracy and individual fairness are both crucial for trustworthy machine learning, but these two aspects are often incompatible with each other so that enhancing one aspect may sacrifice the other inevitably with side effects of true bias or false fairness. We propose in this paper a new fairness criterion, accurate fairness, to align individual fairness with accuracy. Informally, it requires the treatments of an individual and the individual's similar counterparts to conform to a uniform target, i.e., the ground truth of the individual. We prove that accurate fairness also implies typical group fairness criteria over a union of similar sub-populations. We then present a Siamese fairness in-processing approach to minimize the accuracy and fairness losses of a machine learning model under the accurate fairness constraints. To the best of our knowledge, this is the first time that a Siamese approach is adapted for bias mitigation. We also propose fairness confusion matrix-based metrics, fair-precision, fair-recall, and fair-F1 score, to quantify a trade-off between accuracy and individual fairness. Comparative case studies with popular fairness datasets show that our Siamese fairness approach can achieve on average 1.02%-8.78% higher individual fairness (in terms of fairness through awareness) and 8.38%-13.69% higher accuracy, as well as 10.09%-20.57% higher true fair rate, and 5.43%-10.01% higher fair-F1 score, than the state-of-the-art bias mitigation techniques. This demonstrates that our Siamese fairness approach can indeed improve individual fairness without trading accuracy. Finally, the accurate fairness criterion and Siamese fairness approach are applied to mitigate the possible service discrimination with a real Ctrip dataset, by on average fairly serving 112.33% more customers (specifically, 81.29% more customers in an accurately fair way) than baseline models.

AAAI Conference 2023 Conference Paper

CowClip: Reducing CTR Prediction Model Training Time from 12 Hours to 10 Minutes on 1 GPU

  • Zangwei Zheng
  • Pengtai Xu
  • Xuan Zou
  • Da Tang
  • Zhen Li
  • Chenguang Xi
  • Peng Wu
  • Leqi Zou

The click-through rate (CTR) prediction task is to predict whether a user will click on the recommended item. As mind-boggling amounts of data are produced online daily, accelerating CTR prediction model training is critical to ensuring an up-to-date model and reducing the training cost. One approach to increase the training speed is to apply large batch training. However, as shown in computer vision and natural language processing tasks, training with a large batch easily suffers from the loss of accuracy. Our experiments show that previous scaling rules fail in the training of CTR prediction neural networks. To tackle this problem, we first theoretically show that different frequencies of ids make it challenging to scale hyperparameters when scaling the batch size. To stabilize the training process in a large batch size setting, we develop the adaptive Column-wise Clipping (CowClip). It enables an easy and effective scaling rule for the embeddings, which keeps the learning rate unchanged and scales the L2 loss. We conduct extensive experiments with four CTR prediction networks on two real-world datasets and successfully scaled 128 times the original batch size without accuracy loss. In particular, for CTR prediction model DeepFM training on the Criteo dataset, our optimization framework enlarges the batch size from 1K to 128K with over 0.1% AUC improvement and reduces training time from 12 hours to 10 minutes on a single V100 GPU. Our code locates at github.com/bytedance/LargeBatchCTR.

NeurIPS Conference 2023 Conference Paper

Fairly Recommending with Social Attributes: A Flexible and Controllable Optimization Approach

  • Jinqiu Jin
  • Haoxuan Li
  • Fuli Feng
  • Sihao Ding
  • Peng Wu
  • Xiangnan He

Item-side group fairness (IGF) requires a recommendation model to treat different item groups similarly, and has a crucial impact on information diffusion, consumption activity, and market equilibrium. Previous IGF notions only focus on the direct utility of the item exposures, i. e. , the exposure numbers across different item groups. Nevertheless, the item exposures also facilitate utility gained from the neighboring users via social influence, called social utility, such as information sharing on the social media. To fill this gap, this paper introduces two social attribute-aware IGF metrics, which require similar user social attributes on the exposed items across the different item groups. In light of the trade-off between the direct utility and social utility, we formulate a new multi-objective optimization problem for training recommender models with flexible trade-off while ensuring controllable accuracy. To solve this problem, we develop a gradient-based optimization algorithm and theoretically show that the proposed algorithm can find Pareto optimal solutions with varying trade-off and guaranteed accuracy. Extensive experiments on two real-world datasets validate the effectiveness of our approach.

EAAI Journal 2023 Journal Article

Group decision making with hesitant fuzzy linguistic preference relations based on multiplicative DEA cross-efficiency and stochastic acceptability analysis

  • Jingmiao Song
  • Peng Wu
  • Jinpei Liu
  • Huayou Chen

The aim of this paper is to investigate a novel approach to group decision making (GDM) based on multiplicative DEA cross-efficiency and stochastic acceptability analysis with hesitant fuzzy linguistic preference relations (HFLPRs), which can avoid information distortion and obtain more credible decision-making results. First, a transform function is defined, which can extract effective information of the hesitant fuzzy linguistic term set (HFLTS) sufficiently. Then, an optimization model is proposed to derive the occurring probabilities of the elements in the HFLTS. It can avoid the normalization process of HFLPRs and simulate the element-selection process of decision makers (DMs). Moreover, a Maximum Log DEA Cross Efficiency model is proposed to evaluate the relative efficiency of each alternative from HFLPR. We further develop the stochastic weight space acceptability analysis method to solve the GDM problem and a step-by-step procedure is presented. Finally, numerical examples are given to illustrate the validity and applicability of the proposed method. This is the first attempt of employing the multiplicative DEA cross-efficiency to the GDM with HFLPRs.

EAAI Journal 2023 Journal Article

MemSeg: A semi-supervised method for image surface defect detection using differences and commonalities

  • Minghui Yang
  • Peng Wu
  • Hui Feng

High-accuracy and real-time semi-supervised image surface defect detection is extensively needed in industrial scenarios. However, existing methods do not provide a good balance between accuracy and speed of defect detection, so this paper proposes an end-to-end memory-based segmentation network (MemSeg) to better accomplish this task. Considering the small intra-class variance of products in the same production line, from the perspective of differences and commonalities, MemSeg introduces artificially simulated abnormal samples and memory samples to assist the model learning. In the training phase, MemSeg explicitly learns the potential differences between normal and simulated abnormal images to obtain a robust classification hyperplane. At the same time, inspired by the mechanism of human memory, MemSeg uses a memory pool to store the general patterns of normal samples. By comparing the similarities and differences between input samples and memory samples in the memory pool to give effective guesses about abnormal regions; In the inference phase, MemSeg directly determines the abnormal regions of the input image in an end-to-end approach. Simple but high-performance, MemSeg achieves state-of-the-art (SOTA) performance on MVTec AD datasets with AUC scores of 99. 56% and 98. 84% at the image level and pixel level, respectively, while also meeting the real-time requirements in industrial scenarios.

AAAI Conference 2023 Conference Paper

Multiple Robust Learning for Recommendation

  • Haoxuan Li
  • Quanyu Dai
  • Yuru Li
  • Yan Lyu
  • Zhenhua Dong
  • Xiao-Hua Zhou
  • Peng Wu

In recommender systems, a common problem is the presence of various biases in the collected data, which deteriorates the generalization ability of the recommendation models and leads to inaccurate predictions. Doubly robust (DR) learning has been studied in many tasks in RS, with the advantage that unbiased learning can be achieved when either a single imputation or a single propensity model is accurate. In this paper, we propose a multiple robust (MR) estimator that can take the advantage of multiple candidate imputation and propensity models to achieve unbiasedness. Specifically, the MR estimator is unbiased when any of the imputation or propensity models, or a linear combination of these models is accurate. Theoretical analysis shows that the proposed MR is an enhanced version of DR when only having a single imputation and propensity model, and has a smaller bias. Inspired by the generalization error bound of MR, we further propose a novel multiple robust learning approach with stabilization. We conduct extensive experiments on real-world and semi-synthetic datasets, which demonstrates the superiority of the proposed approach over state-of-the-art methods.

NeurIPS Conference 2023 Conference Paper

Removing Hidden Confounding in Recommendation: A Unified Multi-Task Learning Approach

  • Haoxuan Li
  • Kunhan Wu
  • Chunyuan Zheng
  • Yanghao Xiao
  • Hao Wang
  • Zhi Geng
  • Fuli Feng
  • Xiangnan He

In recommender systems, the collected data used for training is always subject to selection bias, which poses a great challenge for unbiased learning. Previous studies proposed various debiasing methods based on observed user and item features, but ignored the effect of hidden confounding. To address this problem, recent works suggest the use of sensitivity analysis for worst-case control of the unknown true propensity, but only valid when the true propensity is near to the nominal propensity within a finite bound. In this paper, we first perform theoretical analysis to reveal the possible failure of previous approaches, including propensity-based, multi-task learning, and bi-level optimization methods, in achieving unbiased learning when hidden confounding is present. Then, we propose a unified multi-task learning approach to remove hidden confounding, which uses a few unbiased ratings to calibrate the learned nominal propensities and nominal error imputations from biased data. We conduct extensive experiments on three publicly available benchmark datasets containing a fully exposed large-scale industrial dataset, validating the effectiveness of the proposed methods in removing hidden confounding.

IJCAI Conference 2022 Conference Paper

On the Opportunity of Causal Learning in Recommendation Systems: Foundation, Estimation, Prediction and Challenges

  • Peng Wu
  • Haoxuan Li
  • Yuhao Deng
  • Wenjie Hu
  • Quanyu Dai
  • Zhenhua Dong
  • Jie Sun
  • Rui Zhang

Recently, recommender system (RS) based on causal inference has gained much attention in the industrial community, as well as the states of the art performance in many prediction and debiasing tasks. Nevertheless, a unified causal analysis framework has not been established yet. Many causal-based prediction and debiasing studies rarely discuss the causal interpretation of various biases and the rationality of the corresponding causal assumptions. In this paper, we first provide a formal causal analysis framework to survey and unify the existing causal-inspired recommendation methods, which can accommodate different scenarios in RS. Then we propose a new taxonomy and give formal causal definitions of various biases in RS from the perspective of violating the assumptions adopted in causal analysis. Finally, we formalize many debiasing and prediction tasks in RS, and summarize the statistical and machine learning-based causal estimation methods, expecting to provide new research opportunities and perspectives to the causal RS community.

AAAI Conference 2022 Short Paper

Optimizing Global Influenza Surveillance for Locations with Deficient Data (Student Abstract)

  • Songwei Shan
  • Qi Tan
  • Yiu Chung Lau
  • Zhanwei Du
  • Eric H.Y. Lau
  • Peng Wu
  • Benjamin J. Cowling

For better monitoring and controlling influenza, WHO has launched FluNet (recently integrated to FluMART) to provide a unified platform for participating countries to routinely collect influenza-related syndromic, epidemiological and virological data. However, the reported data were incomplete. We propose a novel surveillance system based on data from multiple sources to accurately assess the epidemic status of different countries, especially for those with missing surveillance data in some periods. The proposed method can automatically select a small set of reliable and informative indicators for assessing the underlying epidemic status and proper supporting data to train the predictive model. Our proactive selection method outperforms three other out-of-box methods (linear regression, multilayer perceptron, and long-short term memory) to make accurate predictions.

EAAI Journal 2021 Journal Article

Unsupervised anomaly detection for underwater gliders using generative adversarial networks

  • Peng Wu
  • Catherine A. Harris
  • Georgios Salavasidis
  • Alvaro Lorenzo-Lopez
  • Izzat Kamarudzaman
  • Alexander B. Phillips
  • Giles Thomas
  • Enrico Anderlini

An effective anomaly detection system is critical for marine autonomous systems operating in complex and dynamic marine environments to reduce operational costs and achieve concurrent large-scale fleet deployments. However, developing an automated fault detection system remains challenging for several reasons including limited data transmission via satellite services. Currently, most anomaly detection for marine autonomous systems, such as underwater gliders, rely on intensive analysis by pilots. This study proposes an unsupervised anomaly detection system using bidirectional generative adversarial networks guided by assistive hints for marine autonomous systems with time series data collected by multiple sensors. In this study, the anomaly detection system for a fleet of underwater gliders is trained on two healthy deployment datasets and tested on other nine deployment datasets collected by a selection of vehicles operating in a range of locations and environmental conditions. The system is successfully applied to detect anomalies in the nine test deployments, which include several different types of anomalies as well as healthy behaviour. Also, a sensitivity study of the data decimation settings suggests the proposed system is robust for Near Real-Time anomaly detection for underwater gliders.

v2026.09.13