Arrow Research search

Author name cluster

Zhen Chen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

24 papers
2 author rows

Possible papers

24

EAAI Journal 2026 Journal Article

A risk assessment framework for online transactions via Graph Neural Networks and efficient probabilistic prediction

  • Jicai Chang
  • Xuejing Fu
  • Zhen Chen
  • Li Pan
  • Shijun Liu

Transactions are integral to daily life, but the occurrence of abnormal behaviors can lead to significant risks. Online transaction risk is characterized by the accumulation of abnormal behaviors, where their frequency surpasses a predefined threshold, resulting in measurable probabilities and consequences. Therefore, the assessment of online transaction risk heavily depends on probabilistic predictions of the accumulated frequency of abnormal behaviors, presenting two major challenges. Firstly, abnormal behaviors across different instances (e. g. , behavior types, product categories, regions, and platforms) exhibit temporal correlations, such as co-occurrence and concomitance, which most probabilistic models fail to identify and utilize effectively. Additionally, these models do not fully address the real-time demands. To address these challenges, we propose a novel risk assessment framework based on Graph Neural Network (GNN) and probabilistic prediction, named GNN-Probformer. The framework uses Dynamic Time Warping to capture temporal correlations between abnormal behavior frequency sequences and constructs a graph structure through clustering. It then employs Graph Neural Networks to aggregate features and learn representations through a novel embedding module. A sparse self-attention mechanism and an efficient encoder–decoder architecture are incorporated to further enhance performance, while probabilistic predictions are generated through Monte Carlo sampling and cumulative distribution functions. Experimental results on a real-world dataset demonstrate that GNN-Probformer achieves substantial performance gains, with a 15% reduction in normalized deviation. At the 90th percentile, it further reduces normalized quantile loss by 15% and improves the F1-score by 16%, while also reducing training time and inference time by 47% and 38%, respectively.

EAAI Journal 2026 Journal Article

Artificial intelligence-empowered point cloud technologies for rapid resilience assessment of urban infrastructure under environment extremes: A comprehensive review

  • Yunchao Tang
  • Zhen Chen
  • Yanzhen Hu
  • Qixing Liu
  • Lei Yin

Natural disasters such as earthquakes, hurricanes, and floods present significant challenges to the inspection and assessment of damaged urban infrastructure, often hindering rapid emergency response and recovery. Recent advances in point cloud technology, especially when combined with artificial intelligence, are transforming post-disaster damage assessment by enabling fast, high-resolution, and multi-dimensional data acquisition across complex urban environments. This review systematically categorizes and synthesizes current research on the use of Artificial Intelligence(AI)-empowered point cloud technologies for resilience assessment of urban infrastructure under environmental extremes. The approaches are organized according to data availability, including methods that compare pre- and post-disaster point cloud data as well as those that rely solely on post-disaster data for rapid damage evaluation. The review further extends to the assessment of a wide range of civil infrastructure, such as bridges, roads, utility networks, and lifeline systems, highlighting how different platforms and sensor types are selected based on specific disaster scenarios and assessment objectives. For each category of method and infrastructure type, the advantages, limitations, and distinctive features of point cloud-based assessment techniques are thoroughly analyzed. Finally, the paper discusses current challenges and emerging trends in intelligent, automated, and multi-source point cloud solutions, and outlines key directions for future research to enhance the resilience and sustainability of cities facing environmental extremes.

AAAI Conference 2026 Conference Paper

FlowAnyTime: Efficient Fine-tuning with Intra-Inter Frame Distillation for All-Weather Optical Flow Estimation

  • Zixu Wang
  • Hongye Chen
  • Xiaochun Zou
  • Congxuan Zhang
  • Zhen Chen
  • Xinbo Zhao

Motion estimation in degraded scenes has long been a significant challenge, primarily attributed to substantial scene variations and insufficient training data. Existing approaches typically address this limitation by incorporating additional training strategies or modifying network architectures within conventional frameworks. However, these solutions not only require cumbersome training procedures or additional modal inputs, but also lack generalization capabilities. To address this problem, we propose a unified optical flow estimation framework specifically designed for degraded scenes. In this work, we employ large-scale pre-trained optical flow foundation models as both teacher and student networks. Our objective is to compensate for feature incompleteness during image degradation through pre-trained large models. Subsequently, we leverage supervised signals for fine-tuning and introduce an intra-inter frame distillation method to enable the student network to adapt to diverse cross-domain scenarios. Our proposed methodology provides deeper insights into learning style-invariant features from these learnable fine-tuning layers. Extensive experiments demonstrate that our approach achieves superior generalization performance and state-of-the-art results in degraded scenes (including low-light, rain, fog and other conditions) while requiring minimal training resources.

AAAI Conference 2026 Conference Paper

Fragile by Design: On the Limits of Adversarial Defenses in Personalized DreamBooth Generation

  • Zhen Chen
  • Yi Zhang
  • Xiangyu Yin
  • Chengxuan Qin
  • Xingyu Zhao
  • Xiaowei Huang
  • Wenjie Ruan

Personalized AI applications such as DreamBooth enable the generation of customized content from user images, but also raise significant privacy concerns, particularly the risk of facial identity leakage. Recent defense mechanisms like Anti-DreamBooth attempt to mitigate this risk by injecting adversarial perturbations into user photos to prevent successful personalization. However, we identify two critical yet overlooked limitations of these methods. First, the adversarial examples often exhibit perceptible artifacts such as conspicuous patterns or stripes, making them easily detectable as manipulated content. Second, the perturbations are highly fragile, as even a simple, non-learned filter can effectively remove them, thereby restoring the model's ability to memorize and reproduce user identity. To investigate this vulnerability, we propose a novel evaluation framework, AntiDB_Purify, to systematically evaluate existing defenses under realistic purification threats, including both traditional image filters and adversarial purification. Results reveal that none of the current methods maintains their protective effectiveness under such threats. These findings highlight that current defenses offer a false sense of security and underscore the urgent need for more imperceptible and robust protections to safeguard user identity in personalized generation.

EAAI Journal 2026 Journal Article

Reinforcement learning joint control method for strip thickness-crown based on implicit weight contraction

  • Yue Huang
  • Zhen Chen
  • Di Zhou
  • Ershun Pan

Thickness-crown joint control in hot-rolled strip exhibits strongly coupled, multivariable, and time-varying dynamics. Mechanism-based coupled modeling is both complex and difficult to parameterize accurately. Conventional proportional integral (PI) controllers are constrained by simple structure, extensive manual tuning, and limited ability to handle strong coupling and long-horizon performance metrics. Offline reinforcement learning (RL) can learn long-horizon optimization policies from historical operation data and thus has promise for this task. But distributional shift, value-estimation bias, and training instability impede its safe and robust industrial deployment. To address these challenges, we propose an implicit-weighted offline RL control method (IWC) for high-precision thickness-crown regulation. We extend the contraction-mapping condition theory to RL, introduce a novel implicit-weight design, and combine it with a density-ratio correction to mitigate estimation bias caused by distributional shift. Validation on four industrial datasets shows that IWC yields substantial improvements in control accuracy, stability, and robustness relative to PI controllers and classical RL methods, offering a safer and more efficient intelligent solution for actual industrial control. The code will be available at https: //github. com/EtsuHuang/IWC.

EAAI Journal 2025 Journal Article

Cross-domain representation learning with causal invariance

  • Dianlong You
  • Bingxin Liu
  • Dongyan Wang
  • Xiaoyi Ge
  • Zhen Chen
  • Xindong Wu

Domain generalization (DG) aims at learning models from several source domains that generalize effectively to unseen target domains. Existing methods are difficult to learn causal-invariant representations, whose core challenge is to remove spurious and redundant features from potential representations. To this end, we propose a novel cross Domains-Invariant Representation leaRning model with Causal Invariance mechanism ( DIR CI 2 ) for domain generalization with the three-fold ideas: (1) breaking spurious correlations between spurious and causal representations via Fourier-based data augmentation; (2) mining domain-invariant representations by injecting causal-invariant mechanism into multiple cross-domain environments; (3) removing the redundant features through mutual information regularization. To evaluate performance, we conduct extensive experiments with state-of-the-art related algorithms on benchmark datasets. Besides, a case study on cross-domain image fields is conducted to elaborate on our DIR CI 2 model’s effectiveness. The results show that our DIR CI 2 significantly outperforms its rivals. The code is released at https: //github. com/youdianlong/DIR2CI. git.

EAAI Journal 2025 Journal Article

Early prediction of lithium-ion battery health in electric vehicles using small-sample learning and hybrid models

  • Changsheng Zhao
  • Lei Yao
  • Yanqiu Xiao
  • Guangzhen Cui
  • Changhui Qu
  • Huilin Dai
  • Zhen Chen
  • Tiansi Wang

Data collection during electric vehicle operation is costly, limited, and often of low quality, making early prediction of battery health a challenging task. This paper presents an early prediction method for lithium-ion battery health using small-sample learning. Features are categorized into four types, and a comprehensive correlation calculation method is introduced to address feature fluctuation. An innovative hybrid model, incorporating an attention mechanism and Bayesian optimization, is designed to enhance feature contribution and optimize hyperparameters. Experimental results show that with only 10% of the training data and 20 iterations of optimization, the method achieves an average prediction error of less than 3%, demonstrating high accuracy and efficiency in small-sample data learning and multi-feature selection.

EAAI Journal 2025 Journal Article

High-order complementary cloud application programming interface recommendation with logical reasoning for incremental development

  • Zhen Chen
  • Denghui Xie
  • Xiaolong Wang
  • Dianlong You
  • Limin Shen

Cloud application programming interface, as the best carrier for service delivery, data exchange, and capability replication, has been an indispensable element of innovation in today’s app-driven world. However, it is difficult for developers to select the suitable one when facing the sea of cloud application programming interfaces. Existing researches focus on generating single-function and high-quality recommendation lists, while ignoring developers’ needs for high-order complementary cloud application programming interfaces in incremental development. In this paper, we present a high-order complementary cloud application programming interface recommendation with logical reasoning. Firstly, we conduct data analysis to demonstrate the necessity of recommending high-order complementary cloud application programming interfaces and the existence of substitute noise. Secondly, a logical reasoning network is designed using projection, intersection, and negation three logic operators, wherein high-order complementary relations are mined and substitute noises are eliminated. Then, the cloud application programming interface base vector that is complementary but not substitute to the query set is generated, and Kullback–Leibler divergence is subsequently introduced to generate complementary recommendation results. Finally, experimental results demonstrate the superiority of our approach in low-, high-, and hybrid-order complementary recommendation scenarios, and there is a significant increase in hit rate, normalize discounted cumulative gain, mean reciprocal rank, and substitute degree by 11. 43%/4. 86%, 10. 08%/4. 28%, 7. 50%/2. 67%, and 36. 33%/32. 35% on ProgrammableWeb and Huawei AppGallery datasets respectively. The proposed approach is not only more likely to produce diversified results that meet developers’ needs but also help providers better formulate pricing strategies to achieve combined sales and improve revenue.

EAAI Journal 2025 Journal Article

Lightweight defect detection network based on steel strip raw images

  • Yue Huang
  • Zhen Chen
  • Zhaoxiang Chen
  • Di Zhou
  • Ershun Pan

Efficient defect detection on hot rolled steel strips is important for industrial production. However, existing defect detection methods are not lightweight enough and lack detection ability for actual steel production environments. A lightweight detection framework based on raw steel strip defect images is proposed to address this gap. A new strip surface defect detection dataset comprising 2650 raw images with five defect types is established. Considering the unique features of strip images, a new strip image augmentation strategy is employed to enhance training sample diversity. Then, a novel lightweight model is introduced. The model consists of the Location Enhanced Ghost Network (LEG-Net) and the Refine Grouped Spatial Network (RGS-Net). The LEG-Net incorporates Ghost modules and a new Location Enhanced Attention. The lightweight backbone effectively reduces the number of parameters. The RGS-Net neck part consists of a slim neck and Efficient Channel Attention. The RGS-Net increases the extraction of channel information and then realizes the defect recognition between different defect types and backgrounds by adjusting the spatial receptive field mechanism. The mean Average Precision (mAP) accuracy of the proposed lightweight model is 76. 9% and the speed is 33 Frames Per Second (FPS). Compared to existing models on three datasets, the proposed network delivers superior detection accuracy while incurring lower computational costs. It effectively identifies challenging defects, such as small targets and large fuzzy samples. Furthermore, the proposed framework closely resembles the actual steel production environment, thereby advancing the industrial application of intelligent defect detection.

ICML Conference 2025 Conference Paper

POQD: Performance-Oriented Query Decomposer for Multi-vector retrieval

  • Yaoyang Liu
  • Junlin Li
  • Yinjun Wu
  • Zhen Chen

Although Multi-Vector Retrieval (MVR) has achieved the state of the art on many information retrieval (IR) tasks, its performance highly depends on how to decompose queries into smaller pieces, say phrases or tokens. However, optimizing query decomposition for MVR performance is not end-to-end differentiable. Even worse, jointly solving this problem and training the downstream retrieval-based systems, say RAG systems could be highly inefficient. To overcome these challenges, we propose Performance-Oriented Query Decomposer (POQD), a novel query decomposition framework for MVR. POQD leverages one LLM for query decomposition and searches the optimal prompt with an LLM-based optimizer. We further propose an end-to-end training algorithm to alternatively optimize the prompt for query decomposition and the downstream models. This algorithm can achieve superior MVR performance at a reasonable training cost as our theoretical analysis suggests. POQD can be integrated seamlessly into arbitrary retrieval-based systems such as Retrieval-Augmented Generation (RAG) systems. Extensive empirical studies on representative RAG-based QA tasks show that POQD outperforms existing query decomposition strategies in both retrieval performance and end-to-end QA accuracy. POQD is available at https: //github. com/PKU-SDS-lab/POQD-ICML25.

NeurIPS Conference 2025 Conference Paper

TANDEM: Bi-Level Data Mixture Optimization with Twin Networks

  • Jiaxing Wang
  • Deping Xiang
  • Jin Xu
  • Mingyang Yi
  • Guoqiang Gong
  • Zicheng Zhang
  • Haoran Li
  • Pengzhang Liu

The capabilities of large language models (LLMs) significantly depend on training data drawn from various domains. Optimizing domain-specific mixture ratios can be modeled as a bi-level optimization problem, which we simplify into a single-level penalized form and solve with twin networks: a proxy model trained on primary data and a dynamically updated reference model trained with additional data. Our proposed method, Twin Networks for bi-level DatA mixturE optiMization (TANDEM), measures the data efficacy through the difference between the twin models and up-weights domains that benefit more from the additional data. TANDEM provides theoretical guarantees and wider applicability, compared to prior approaches. Furthermore, our bi-level perspective suggests new settings to study domain reweighting such as data-restricted scenarios and supervised fine-tuning, where optimized mixture ratios significantly improve the performance. Extensive experiments validate TANDEM's effectiveness in all scenarios.

AAAI Conference 2025 Conference Paper

U-KAN Makes Strong Backbone for Medical Image Segmentation and Generation

  • Chenxin Li
  • Xinyu Liu
  • Wuyang Li
  • Cheng Wang
  • Hengyu Liu
  • Yifan Liu
  • Zhen Chen
  • Yixuan Yuan

U-Net has become a cornerstone in various visual applications such as image segmentation and diffusion probability models. While numerous innovative designs and improvements have been introduced by incorporating transformers or MLPs, the networks are still limited to linearly modeling patterns as well as the deficient interpretability. To address these challenges, our intuition is inspired by the impressive results of the Kolmogorov-Arnold Networks (KANs) in terms of accuracy and interpretability, which reshape the neural network learning via the stack of non-linear learnable activation functions derived from the Kolmogorov-Anold representation theorem. Specifically, in this paper, we explore the untapped potential of KANs in improving backbones for vision tasks. We investigate, modify and re-design the established U-Net pipeline by integrating the dedicated KAN layers on the tokenized intermediate representation, termed U-KAN. Rigorous medical image segmentation benchmarks verify the superiority of UKAN by higher accuracy even with less computation cost. We further delved into the potential of U-KAN as an alternative U-Net noise predictor in diffusion models, demonstrating its applicability in generating task-oriented model architectures.

JBHI Journal 2025 Journal Article

Unified Multi-Modal Diagnostic Framework With Reconstruction Pre-Training and Heterogeneity-Combat Tuning

  • Yupei Zhang
  • Li Pan
  • Qiushi Yang
  • Tan Li
  • Zhen Chen

Medical multi-modal pre-training has revealed promise in computer-aided diagnosis by leveraging large-scale unlabeled datasets. However, existing methods based on masked autoencoders mainly rely on data-level reconstruction tasks, but lack high-level semantic information. Furthermore, two significant heterogeneity challenges hinder the transfer of pre-trained knowledge to downstream tasks, i. e. , the distribution heterogeneity between pre-training data and downstream data, and the modality heterogeneity within downstream data. To address these challenges, we propose a Unified Medical Multi-modal Diagnostic (UMD) framework with tailored pre-training and downstream tuning strategies. Specifically, to enhance the representation abilities of vision and language encoders, we propose the Multi-level Reconstruction Pre-training (MR-Pretrain) strategy, including a feature-level and data-level reconstruction, which guides models to capture the semantic information from masked inputs of different modalities. Moreover, to tackle two kinds of heterogeneities during the downstream tuning, we present the heterogeneity-combat downstream tuning strategy, which consists of a Task-oriented Distribution Calibration (TD-Calib) and a Gradient-guided Modality Coordination (GM-Coord). In particular, TD-Calib fine-tunes the pre-trained model regarding the distribution of downstream datasets, and GM-Coord adjusts the gradient weights according to the dynamic optimization status of different modalities. Extensive experiments on five public medical datasets demonstrate the effectiveness of our UMD framework, which remarkably outperforms existing approaches on three kinds of downstream tasks.

ICLR Conference 2025 Conference Paper

UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic Segmentation

  • Tao Zhang
  • Jinyong Wen
  • Zhen Chen
  • Kun Ding 0001
  • Shiming Xiang
  • Chunhong Pan

Pre-training techniques significantly enhance the performance of semantic segmentation tasks with limited training data. However, the efficacy under a large domain gap between pre-training (e.g. RGB) and fine-tuning (e.g. infrared) remains underexplored. In this study, we first benchmark the infrared semantic segmentation performance of various pre-training methods and reveal several phenomena distinct from the RGB domain. Next, our layerwise analysis of pre-trained attention maps uncovers that: (1) There are three typical attention patterns (local, hybrid, and global); (2) Pre-training tasks notably influence pattern distribution across layers; (3) The hybrid pattern is crucial for semantic segmentation as it attends to both nearby and foreground elements; (4) The texture bias impedes model generalization in infrared tasks. Building on these insights, we propose UNIP, a UNified Infrared Pre-training framework, to enhance the pre-trained model performance. This framework uses the hybrid-attention distillation NMI-HAD as the pre-training target, a large-scale mixed dataset InfMix for pre-training, and a last-layer feature pyramid network LL-FPN for fine-tuning. Experimental results show that UNIP outperforms various pre-training methods by up to 13.5% in average mIoU on three infrared segmentation tasks, evaluated using fine-tuning and linear probing metrics. UNIP-S achieves performance on par with MAE-L while requiring only 1/10 of the computational cost. Furthermore, with fewer parameters, UNIP significantly surpasses state-of-the-art (SOTA) infrared or RGB segmentation methods and demonstrates the broad potential for application in other modalities, such as RGB and depth. Our code is available at https://github.com/casiatao/UNIP.

AAAI Conference 2025 Conference Paper

Universal Domain Adaptive Object Detection via Dual Probabilistic Alignment

  • Yuanfan Zheng
  • Jinlin Wu
  • Wuyang Li
  • Zhen Chen

Domain Adaptive Object Detection (DAOD) transfers knowledge from a labeled source domain to an unannotated target domain under closed-set assumption. Universal DAOD (UniDAOD) extends DAOD to handle open-set, partial-set, and closed-set domain adaptation. In this paper, we first unveil two issues: domain-private category alignment is crucial for global-level features, and the domain probability heterogeneity of features across different levels. To address these issues, we propose a novel Dual Probabilistic Alignment (DPA) framework to model domain probability as Gaussian distribution, enabling the heterogeneity domain distribution sampling and measurement. The DPA consists of three tailored modules: the Global-level Domain Private Alignment (GDPA), the Instance-level Domain Shared Alignment (IDSA), and the Private Class Constraint (PCC). GDPA utilizes the global-level sampling to mine domain-private category samples and calculate alignment weight through a cumulative distribution function to address the global-level private category alignment. IDSA utilizes instance-level sampling to mine domain-shared category samples and calculates alignment weight through Gaussian distribution to conduct the domain-shared category domain alignment to address the feature heterogeneity. The PCC aggregates domain-private category centroids between feature and probability spaces to mitigate negative transfer. Extensive experiments demonstrate that our DPA outperforms state-of-the-art UniDAOD and DAOD methods across various datasets and scenarios, including open, partial, and closed sets.

EAAI Journal 2025 Journal Article

Unsupervised motion-based anomaly detection with graph attention networks for industrial robots labeling

  • Jinrui Han
  • Zhen Chen
  • Di Zhou
  • Bing Hu
  • Tangbin Xia
  • Ershun Pan

As automated labeling on products in intelligent manufacturing grows in importance, detecting anomalies in the end-effectors used for industrial robots labeling is essential for maintaining production line stability and efficiency. Considering the distinct characteristics of specific movements in the labeling process, different motions, such as moving, labeling and rolling, contribute different effects to end-effector abnormalities. It's challenging to distinguish between normal and anomalies, instead of treating all motions as a homogeneous whole. Also, real-world industrial scenarios often lack the sufficient data on abnormal states, and resource constraints limit computation efficiency. In view of this, this paper aims to develop a task-specific anomaly detection solution tailored to the distinct motions of industrial robots labeling. To achieve this goal, an unsupervised, motion-based anomaly detection framework is proposed. The raw sensor signals from each motion are segmented and a group of encoder networks are employed to extract latent representations for each motion. Then, these motions are modeled as nodes in a graph, where a feature fusion module based on a Graph Attention Network (GAT) captures the interrelationships between them. A memory-augmented reconstruction module with multi-scale skip connections enhances the model's ability to detect anomalies. Finally, an anomaly detection module identifies abnormal states of the end-effector. Experimental validations are conducted on a dataset from real-world steel coil labeling task. The results show that the proposed framework can achieve an average performance of 98. 24% with an inference time of 15 ms, also demonstrating the effectiveness of its structural design and key modules.

EAAI Journal 2024 Journal Article

A data-driven approach to full-field stress reconstruction of ship hull structure using deep learning

  • Chao Sun
  • Zhen Chen
  • Junan Yi
  • Dongyang Li

Reconstructing full-field stress distribution has extensive engineering applications in design optimization and structure health monitoring. This article develops a data-driven approach for efficient and accurate reconstruction of stress fields in ship hull structure. The method integrates numerical simulation with conditional generative adversarial network (cGAN) to infer full-field responses based on stress values of limited monitoring points. The network architecture, which consists of a generator and a discriminator, is optimized through adversarial training. Based on the training database derived from finite element analysis (FEA), full-field von Mises stress distribution of the inner bottom plate of an oil tanker is reconstructed. The discrete stresses obtained from FEA are utilized as input to simulate on-board sensor data used in cGAN-based model. According to the stresses comparison between cGAN-based model and finite element model, it is observed that a sparse arrangement of monitoring points enables accurate reconstruction of the full-field stresses. It is hence shown that this method provides a potential alternative for field monitoring.

TCS Journal 2024 Journal Article

Computational task offloading algorithm based on deep reinforcement learning and multi-task dependency

  • Xiaoqi Zhang
  • Tengxiang Lin
  • Cheng-Kuan Lin
  • Zhen Chen
  • Hongju Cheng

Edge computing is an emerging promising computing paradigm, which can significantly reduce the service latency by moving computing and storage demands to the edge of the network. Resource-constrained edge servers may fail to process multiple tasks simultaneously when several time-delay-sensitive and computationally demanding tasks are offloaded to only one edge server, and results in some issues such as high task processing costs. In this paper, we introduce a novel idea by dividing one task into several sub-tasks via the dependencies within the task and then offloading the sub-tasks to other edge servers in light of high concurrency for synchronization to minimize the total cost of task processing. To address the challenge of task dependencies and adaptation to dynamic scenes, we propose a Multi-Task Dependency Offloading Algorithm (MTDOA) based on deep reinforcement learning. The task offloading decision is modeled as a Markov decision process, and then a graph attention network is applied to extract the dependency information of different tasks, while LSTM and DQN are combined to deal with sequential problems. The simulation results show that the proposed MTDOA has better convergence ability compared with the baseline algorithms.

EAAI Journal 2024 Journal Article

Dynamic time scales ensemble framework for similarity-based remaining useful life prediction under multiple failure modes

  • Yuhui Xu
  • Tangbin Xia
  • Dong Wang
  • Zhen Chen
  • Ershun Pan
  • Lifeng Xi

In modern industry, the stochastic degradation of mechanical equipment typically involves multiple failure modes, which heavily affects the reliability of the remaining useful life (RUL) prediction. The similarity-based methods have been widely deployed in RUL prediction due to their flexibility, but it is still challenging to accurately identify similar degradation trajectories under varying failure modes. The obstacles lie in the interference of reference trajectories under different degradation states and the insufficiency of measuring trajectory trends. Therefore, this paper proposes a dynamic scales ensemble method based on the mean removal Canberra distance with failure identification (FI-MRC-DSE) for similarity-based prognosis. Firstly, a gated recurrent unit autoencoder network is employed to adaptively extract failure features from multi-dimensional monitoring data to support the targeted selection of reference trajectories. Then, the similarity matching is performed based on the proposed MRC distance instead of the commonly used Euclidean distance, enhancing the perception of degradation trends. Finally, the matching results across multiple time scales, which are dynamically determined by the instance's degradation state, are integrated to obtain the predicted RUL. It effectively overcomes the insufficient utilization of trajectory caused by the single time scale. In the experiments, the superiority of our developed similarity-based FI-MRC-DSE method is demonstrated by comparison with the state-of-the-art similarity-based methods. The effectiveness analyses and the ablation study show that all three key components contribute to accurate prognosis under multiple failure modes.

NeurIPS Conference 2024 Conference Paper

Flaws can be Applause: Unleashing Potential of Segmenting Ambiguous Objects in SAM

  • Chenxin Li
  • Yuzhi Huang
  • Wuyang Li
  • Hengyu Liu
  • Xinyu Liu
  • Qing Xu
  • Zhen Chen
  • Yue Huang

As the vision foundation models like the Segment Anything Model (SAM) demonstrate potent universality, they also present challenges in giving ambiguous and uncertain predictions. Significant variations in the model output and granularity can occur with simply subtle changes in the prompt, contradicting the consensus requirement for the robustness of a model. While some established works have been dedicated to stabilizing and fortifying the prediction of SAM, this paper takes a unique path to explore how this flaw can be inverted into an advantage when modeling inherently ambiguous data distributions. We introduce an optimization framework based on a conditional variational autoencoder, which jointly models the prompt and the granularity of the object with a latent probability distribution. This approach enables the model to adaptively perceive and represent the real ambiguous label distribution, taming SAM to produce a series of diverse, convincing, and reasonable segmentation outputs controllably. Extensive experiments on several practical deployment scenarios involving ambiguity demonstrates the exceptional performance of our framework. Project page: \url{https: //a-sa-m. github. io/}.

AAAI Conference 2024 Conference Paper

KGTS: Contrastive Trajectory Similarity Learning over Prompt Knowledge Graph Embedding

  • Zhen Chen
  • Dalin Zhang
  • Shanshan Feng
  • Kaixuan Chen
  • Lisi Chen
  • Peng Han
  • Shuo Shang

Trajectory similarity computation serves as a fundamental functionality of various spatial information applications. Although existing deep learning similarity computation methods offer better efficiency and accuracy than non-learning solutions, they are still immature in trajectory embedding and suffer from poor generality and heavy preprocessing for training. Targeting these limitations, we propose a novel framework named KGTS based on knowledge graph grid embedding, prompt trajectory embedding, and unsupervised contrastive learning for improved trajectory similarity computation. Specifically, we first embed map grids with a GRot embedding method to vigorously grasp the neighbouring relations of grids. Then, a prompt trajectory embedding network incorporates the resulting grid embedding and extracts trajectory structure and point order information. It is trained by unsupervised contrastive learning, which not only alleviates the heavy preprocessing burden but also provides exceptional generality with creatively designed strategies for positive sample generation. The prompt trajectory embedding adopts a customized prompt paradigm to mitigate the gap between the grid embedding and the trajectory embedding. Extensive experiments on two real-world trajectory datasets demonstrate the superior performance of KGTS over state-of-the-art methods.

NeurIPS Conference 2024 Conference Paper

TARP-VP: Towards Evaluation of Transferred Adversarial Robustness and Privacy on Label Mapping Visual Prompting Models

  • Zhen Chen
  • Yi Zhang
  • Fu Wang
  • Xingyu Zhao
  • Xiaowei Huang
  • Wenjie Ruan

Adversarial robustness and privacy of deep learning (DL) models are two widely studied topics in AI security. Adversarial training (AT) is an effective approach to improve the robustness of DL models against adversarial attacks. However, while models with AT demonstrate enhanced robustness, they become more susceptible to membership inference attacks (MIAs), thus increasing the risk of privacy leakage. This indicates a negative trade-off between adversarial robustness and privacy in general deep learning models. Visual prompting is a novel model reprogramming (MR) technique used for fine-tuning pre-trained models, achieving good performance in vision tasks, especially when combined with the label mapping technique. However, the performance of label-mapping-based visual prompting (LM-VP) under adversarial attacks and MIAs lacks evaluation. In this work, we regard the MR of LM-VP as a unified entity, referred to as the LM-VP model, and take a step toward jointly evaluating the adversarial robustness and privacy of LM-VP models. Experimental results show that the choice of pre-trained models significantly affects the white-box adversarial robustness of LM-VP, and standard AT even substantially degrades its performance. In contrast, transfer AT-trained LM-VP achieves a good trade-off between transferred adversarial robustness and privacy, a finding that has been consistently validated across various pre-trained models.

AAAI Conference 2021 Conference Paper

Diagnose Like A Pathologist: Weakly-Supervised Pathologist-Tree Network for Slide-Level Immunohistochemical Scoring

  • Zhen Chen
  • Jun Zhang
  • Shuanlong Che
  • Junzhou Huang
  • Xiao Han
  • Yixuan Yuan

The immunohistochemistry (IHC) test of biopsy tissue is crucial to develop targeted treatment and evaluate prognosis for cancer patients. The IHC staining slide is usually digitized into the whole-slide image (WSI) with gigapixels for quantitative image analysis. To perform a whole image prediction (e. g. , IHC scoring, survival prediction, and cancer grading) from this kind of high-dimensional image, algorithms are often developed based on multi-instance learning (MIL) framework. However, the multi-scale information of WSI and the associations among instances are not well explored in existing MIL based studies. Inspired by the fact that pathologists jointly analyze visual fields at multiple powers of objective for diagnostic predictions, we propose a Pathologist-Tree Network (PTree-Net) to sparsely model the WSI efficiently in multi-scale manner. Specifically, we propose a Focal-Aware Module (FAM) that can approximately estimate diagnosis-related regions with an extractor trained using the thumbnail of WSI. With the initial diagnosis-related regions, we hierarchically model the multi-scale patches in a tree structure, where both the global and local information can be captured. To explore this tree structure in an end-to-end network, we propose a patch Relevance-enhanced Graph Convolutional Network (RGCN) to explicitly model the correlations of adjacent parent-child nodes, accompanied by patch relevance to exploit the implicit contextual information among distant nodes. In addition, tree-based self-supervision is devised to improve representation learning and suppress irrelevant instances adaptively. Extensive experiments are performed on a large-scale IHC HER2 dataset. The ablation study confirms the effectiveness of our design, and our approach outperforms state-of-the-art by a large margin.

AAAI Conference 2017 Conference Paper

CatchÕEm All: Locating Multiple Diffusion Sources in Networks with Partial Observations

  • Kai Zhu
  • Zhen Chen
  • Lei Ying

This paper studies the problem of locating multiple diffusion sources in networks with partial observations. We propose a new source localization algorithm, named Optimal-Jordan- Cover (OJC). The algorithm first extracts a subgraph using a candidate selection algorithm that selects source candidates based on the number of observed infected nodes in their neighborhoods. Then, in the extracted subgraph, OJC finds a set of nodes that “cover” all observed infected nodes with the minimum radius. The set of nodes is called the Jordan cover, and is regarded as the set of diffusion sources. Considering the heterogeneous susceptible-infected-recovered (SIR) diffusion in the Erdős-Rényi (ER) random graph, we prove that OJC can locate all sources with probability one asymptotically with partial observations. OJC is a polynomial-time algorithm in terms of network size. However, the computational complexity increases exponentially in m, the number of sources. We further propose a low-complexity heuristic based on the K-Means for approximating the Jordan cover, named Approximate-Jordan- Cover (AJC). Simulations on random graphs and real networks demonstrate that both AJC and OJC significantly outperform other heuristic algorithms.

v2026.09.13