Arrow Research search

Author name cluster

Wen Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

19 papers
2 author rows

Possible papers

19

EAAI Journal 2026 Journal Article

A risk-aware driving agent with Long Short-Term Memory and Time-to-Collision co-decision and multi-modal fusion in urban simulation

  • Xu Zhao
  • Yu Gao
  • Wen Liu
  • Lianpeng Li
  • Suxian Zhang
  • Zijun Wang

We present a risk-aware driving agent for urban freeway settings that combines Long Short-Term Memory (LSTM) motion prediction with a dynamic Time-to-Collision (TTC) safety gate. Light Detection and Ranging (LiDAR), point-cloud radar, and camera inputs are fused in a unified interface; the decision layer predicts lead-vehicle acceleration and combines it with TTC that reflects road adhesion, braking capability, and obstacle confidence. A finite-state machine with hysteresis maps TTC to four discrete risk actions: safe, warning, prepare, and automatic emergency braking (AEB). The enhanced version adds multi-target Kalman tracking, adaptive TTC parameters, and physics-aware acceleration limits. The TTC gate remains explicit. A Transformer-based fusion backbone and a deep Q-network (DQN) planner process heterogeneous sensor features. In this engineering artificial intelligence application for autonomous driving and intelligent transportation systems, the implemented artificial intelligence modules (including LSTM prediction, attention-based multi-modal fusion, and DQN planning) operate under the same TTC-based safety logic so that risk metrics and braking bounds remain explicit. Experiments cover a 60-second, 1. 9-kilometer free-flow run and two scripted hazards: an aggressive cut-in (minimum headway 0. 64 meters, minimum TTC 0. 07 s, warning/prepare/automatic emergency braking occupancy 48. 8%, with 10. 8% automatic emergency braking) and a pedestrian crossing (minimum pedestrian gap 3. 1 meters, warning/prepare/automatic emergency braking occupancy 48. 0%). A TTC-parameter sweep and a TTC-only ablation show that the current thresholds are conservative in benign cases and that calibration should focus on sharper hazards. The agent design and logs provide a transparent baseline for longitudinal risk-aware control.

TCS Journal 2026 Journal Article

An approximation algorithm for the asymmetric profitable tour problem with submodular penalties

  • Zhenzhen Pang
  • Wen Liu
  • Bo Hou

In this paper, we study the asymmetric profitable tour problem with submodular penalties. The objective is to find a directed tour that visits a subset of vertices such that the length of the tour plus the penalty for the rejected vertex set, which is determined by a submodular function, is minimized. Following the framework as in the Bienstock et al. ’s algorithm [1], we design an approximation algorithm for this problem. Our algorithm achieves an approximation ratio of n ( 1 + ⌈ log ( n ) ⌉ ), where n is the number of vertices.

TCS Journal 2026 Journal Article

Approximation algorithm for k-Product uncapacitated facility location problem with submodular penalties

  • Chenzheng Feng
  • Wen Liu
  • Gengsheng Zhang
  • Bo Hou

In this paper, we consider the k -product uncapacitated facility location problem with submodular penalties. In this problem, we are given a set of demand points where clients are located, a set of potential locations where unlimited capacity facilities can be opened and a set of k different kinds of products. Each open facility can only supply one kind of product, and its open cost is determined by the product it supplies. There is a service cost between each pair of locations. Assume these costs of service are metric. Each client is either supplied k different kinds of products by a set of k different open facilities or completely rejected and a rejection cost has to be paid, which is determined by a submodular function. The objective is to minimize the total cost, including the cost of opening facilities at sites, the service cost for providing products to clients from the open facilities, and the penalty cost of the set of the rejected clients. Based on the LP rounding technique, we propose a ( 2 k + 2 ) -approximation algorithm for this problem.

EAAI Journal 2025 Journal Article

World model-based reinforcement learning for autonomous ship safe collision avoidance

  • Kangjie Zheng
  • Xinyu Zhang
  • Wen Liu
  • Jialin Ma
  • Jinlong Cui

Autonomous ships encounter substantial challenges in dynamic collision avoidance within complex maritime environments, characterized by high sampling requirements, costly trial-and-error processes, and limited adaptability. Conventional artificial intelligence methods such as reinforcement learning, which depend heavily on predefined datasets and fixed training scenarios, often fail to cope with rapidly changing obstacles and uncertain navigation conditions. This study introduces a novel framework, named the world model-based reinforcement learning for autonomous ship safe collision avoidance, which leverages world model for safe and adaptive maritime navigation. Specifically, this study is the first to apply and improve world model theory for maritime collision avoidance. It integrates a variational auto-encoder with a transformer-based sequence model to compress high-dimensional observations into a latent space and predict future states, rewards, and terminal conditions. Based on this latent dynamics, a tree search algorithm is applied to optimize local trajectories. The combination of tree-search-derived and actor policies helps improve stability and safety. Additionally, a multi-objective loss function, along with randomized training scenarios, enhances the framework’s robustness and ability to generalize across different marine environments. Experimental results demonstrate that the proposed artificial intelligence-driven approach achieves reliable collision avoidance and adaptability under complex conditions, while reducing reliance on extensive real-world data.

EAAI Journal 2024 Journal Article

Data-driven deformation prediction and control for existing tunnels below shield tunneling

  • Zongbao Feng
  • Jingyi Wang
  • Wen Liu
  • Tiejun Li
  • Xianguo Wu
  • Pengxin Zhao

In order to effectively and accurately control the existing adjacent tunnel deformation caused by shield adjacent undercrossing construction, a hybrid intelligent framework for predicting and analyzing existing tunnel deformation based on a Bayesian optimization light gradient boosting machine (BO-LGBM) and a digital twin system is established in this paper. The BO-LGBM model forecasts five current tunnel deformation control targets while identifying crucial features influencing adjacent tunnel deformation. This prediction model for existing tunnel deformation under shield tunneling is subject to a comparative assessment of five machine learning algorithms, with its interpretability analyzed utilizing the Shapley Additive exPlanations (SHAP) algorithm. Furthermore, a digital twin system is established to furnish a visual platform for making informed decisions regarding the control of existing tunnel deformations. The applicability and validity of the proposed method are tested via a case study from the Wuhan Metro. The results indicate that: (1) A digital twin system is established, the existing tunnel deformations are monitored, and construction parameters are recommended for real-time adjustment. (2) The proposed BO-LGBM framework can realize rapid and accurate deformation prediction, and the goodness of fit (R 2 ) range of the five targets is 0. 951–0. 977. (3) An early warning standard is established, corresponding treatment measures are proposed for existing tunnel deformation control, and control measures are proposed and adopted in this case. The proposed prediction and control framework for deformation of existing tunnels adjacent to shield tunneling provides a timely reference for ensuring safe tunneling operations.

ICLR Conference 2024 Conference Paper

DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior

  • Jingxiang Sun
  • Bo Zhang
  • Ruizhi Shao
  • Lizhen Wang 0002
  • Wen Liu
  • Zhenda Xie
  • Yebin Liu

We present DreamCraft3D, a hierarchical 3D content generation method that produces high-fidelity and coherent 3D objects. We tackle the problem by leveraging a 2D reference image to guide the stages of geometry sculpting and texture boosting. A central focus of this work is to address the consistency issue that existing works encounter. To sculpt geometries that render coherently, we perform score distillation sampling via a view-dependent diffusion model. This 3D prior, alongside several training strategies, prioritizes the geometry consistency but compromises the texture fidelity. We further propose bootstrapped score distillation to specifically boost the texture. We train a personalized diffusion model, Dreambooth, on the augmented renderings of the scene, imbuing it with 3D knowledge of the scene being optimized. The score distillation from this 3D-aware diffusion prior provides view-consistent guidance for the scene. Notably, through an alternating optimization of the diffusion prior and 3D scene representation, we achieve mutually reinforcing improvements: the optimized 3D scene aids in training the scene-specific diffusion model, which offers increasingly view-consistent guidance for 3D optimization. The optimization is thus bootstrapped and leads to substantial texture boosting. With tailored 3D priors throughout the hierarchical generation, DreamCraft3D generates coherent 3D objects with photorealistic renderings, advancing the state-of-the-art in 3D content generation.

AAAI Conference 2024 Conference Paper

PM-INR: Prior-Rich Multi-Modal Implicit Large-Scale Scene Neural Representation

  • Yiying Yang
  • Fukun Yin
  • Wen Liu
  • Jiayuan Fan
  • Xin Chen
  • Gang Yu
  • Tao Chen

Recent advancements in implicit neural representations have contributed to high-fidelity surface reconstruction and photorealistic novel view synthesis. However, with the expansion of the scene scale, such as block or city level, existing methods will encounter challenges because traditional sampling cannot cope with the cubically growing sampling space. To alleviate the dependence on filling the sampling space, we explore using multi-modal priors to assist individual points to obtain more global semantic information and propose a priorrich multi-modal implicit neural representation network, Pm-INR, for the outdoor unbounded large-scale scene. The core of our method is multi-modal prior extraction and crossmodal prior fusion modules. The former encodes codebooks from different modality inputs and extracts valuable priors, while the latter fuses priors to maintain view consistency and preserve unique features among multi-modal priors. Finally, feature-rich cross-modal priors are injected into the sampling regions to allow each region to perceive global information without filling the sampling space. Extensive experiments have demonstrated the effectiveness and robustness of our method for outdoor unbounded large-scale scene novel view synthesis, which outperforms state-of-the-art methods in terms of PSNR, SSIM, and LPIPS.

YNICL Journal 2023 Journal Article

Altered gray matter volumes and plasma IL-6 level in major depressive disorder patients with suicidal ideation

  • Yingrui Guo
  • Xiaowei Jiang
  • Linna Jia
  • Yue Zhu
  • Xinyu Han
  • Yifan Wu
  • Wen Liu
  • Wenhui Zhao

BACKGROUNDS: Suicidal ideation (SI) is one of the most serious consequences of major depressive disorder (MDD). Understanding the unique mechanism of MDD with SI (MDD + S) is crucial for treatment development. While abundant research has studied MDD, past studies have not reached a consensus on the mechanism of MDD + S. The study aimed to investigate the abnormalities of the gray matter volumes (GMVs) and plasma IL-6 level in MDD + S to further reveal the mechanism of MDD + S. METHODS: We tested the plasma IL-6 level using Luminex multifactor assays and collected the Structural Magnetic Resonance Imaging (SMRI) data from 34 healthy controls (HCs), 36 MDD patients without SI (MDD - S) and 34 MDD + S patients. We performed a partial correlation between the GMVs of the brain regions with significant differences and plasma IL-6 level with age, sex, medication, scores of HAMD-17 and HAMA as the covariates. RESULTS: Compared with HCs and MDD - S, MDD + S had significantly decreased GMVs in the left cerebellum Crus I/II and significantly increased plasma IL-6 level; compared with HCs, both the MDD + S and MDD - S had significantly decreased GMVs in right precentral and postcentral gyri. No significant correlation was found between the GMVs and the plasma IL-6 level in the MDD + S and MDD - S, respectively. While the GMVs of the right precentral and postcentral gyri negatively correlated with the level of IL-6 in the whole MDD (r = -0.28, P = 0.03). The GMVs of the left cerebellum Crus I/II (r = -0.47, P = 0.02), and the right precentral and postcentral gyri (r = -0.42, P = 0.04) negatively correlated with the level of IL-6 in HCs. CONCLUSION: The altered GMVs and the plasma IL-6 level may provide a scientific basis to understand the pathophysiological mechanisms of MDD + S.

YNICL Journal 2023 Journal Article

Gray matter volume reduction in orbitofrontal cortex correlated with plasma glial cell line-derived neurotrophic factor (GDNF) levels within major depressive disorder

  • Yifan Wu
  • Lingtao Kong
  • Anqi Yang
  • Kaiqi Xin
  • Yihui Lu
  • Xintong Yan
  • Wen Liu
  • Yue Zhu

BACKGROUND: Major depressive disorder (MDD) is a severe mental disorder characterized by reduced gray matter volume (GMV). To date, the pathogenesis of MDD remains unclear, but neurotrophic factors play an essential role in the pathophysiological alterations of MDD during disease development. In particular, plasma glial cell line-derived neurotrophic factor (GDNF) has been suggested as a potential biomarker that may be associated with disease activity and neurological progression in MDD. Our study investigated whether plasma GDNF levels in MDD patients and healthy controls (HCs) are correlated with GMV alterations. METHODS: We studied 54 MDD patients and 48 HCs. The effect of different diagnoses on whole-brain GMV was investigated using ANOVA (Analysis of Variance). The threshold of significance was p < 0.05, and Gaussian random-field (GRF) correction for error was used. All analyses were controlled for covariates such as ethnicity, handedness, age, and gender that could affect GMV. RESULT: Compared with the HC group, the GMV in the MDD group was significantly reduced in the right inferior orbitofrontal cortex (OFC), and plasma GDNF levels were significantly higher in the MDD group than in the HC group. In the right inferior OFC, the GDNF levels were positively correlated with GMV reduction in the MDD group, whereas in the HC group, a negative correlation was observed between GDNF levels and GMV reduction. CONCLUSION: Although increased production of GDNF in MDD may help repair neural damage in brain regions associated with brain disease, its repairing effects may be interfered with and hindered by underlying neuroinflammatory processes.

NeurIPS Conference 2023 Conference Paper

Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representation

  • Zibo Zhao
  • Wen Liu
  • Xin Chen
  • Xianfang Zeng
  • Rui Wang
  • Pei Cheng
  • Bin Fu
  • Tao Chen

We present a novel alignment-before-generation approach to tackle the challenging task of generating general 3D shapes based on 2D images or texts. Directly learning a conditional generative model from images or texts to 3D shapes is prone to producing inconsistent results with the conditions because 3D shapes have an additional dimension whose distribution significantly differs from that of 2D images and texts. To bridge the domain gap among the three modalities and facilitate multi-modal-conditioned 3D shape generation, we explore representing 3D shapes in a shape-image-text-aligned space. Our framework comprises two models: a Shape-Image-Text-Aligned Variational Auto-Encoder (SITA-VAE) and a conditional Aligned Shape Latent Diffusion Model (ASLDM). The former model encodes the 3D shapes into the shape latent space aligned to the image and text and reconstructs the fine-grained 3D neural fields corresponding to given shape embeddings via the transformer-based decoder. The latter model learns a probabilistic mapping function from the image or text space to the latent shape space. Our extensive experiments demonstrate that our proposed approach can generate higher-quality and more diverse 3D shapes that better semantically conform to the visual or textural conditional inputs, validating the effectiveness of the shape-image-text-aligned space for cross-modality 3D shape generation.

NeurIPS Conference 2023 Conference Paper

MotionGPT: Human Motion as a Foreign Language

  • Biao Jiang
  • Xin Chen
  • Wen Liu
  • Jingyi Yu
  • Gang Yu
  • Tao Chen

Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multimodal data, such as motion, remains challenging and untouched so far. Fortunately, human motion displays a semantic coupling akin to human language, often perceived as a form of body language. By fusing language data with large-scale motion models, motion-language pre-training that can enhance the performance of motion-related tasks becomes feasible. Driven by this insight, we propose MotionGPT, a unified, versatile, and user-friendly motion-language model to handle multiple motion-relevant tasks. Specifically, we employ the discrete vector quantization for human motion and transfer 3D motion into motion tokens, similar to the generation process of word tokens. Building upon this "motion vocabulary", we perform language modeling on both motion and text in a unified manner, treating human motion as a specific language. Moreover, inspired by prompt learning, we pre-train MotionGPT with a mixture of motion-language data and fine-tune it on prompt-based question-and-answer tasks. Extensive experiments demonstrate that MotionGPT achieves state-of-the-art performances on multiple motion tasks including text-driven motion generation, motion captioning, motion prediction, and motion in-between.

NeurIPS Conference 2023 Conference Paper

PDF: Point Diffusion Implicit Function for Large-scale Scene Neural Representation

  • Yuhan Ding
  • Fukun Yin
  • Jiayuan Fan
  • Hui Li
  • Xin Chen
  • Wen Liu
  • Chongshan Lu
  • Gang Yu

Recent advances in implicit neural representations have achieved impressive results by sampling and fusing individual points along sampling rays in the sampling space. However, due to the explosively growing sampling space, finely representing and synthesizing detailed textures remains a challenge for unbounded large-scale outdoor scenes. To alleviate the dilemma of using individual points to perceive the entire colossal space, we explore learning the surface distribution of the scene to provide structural priors and reduce the samplable space and propose a Point Diffusion implicit Function, PDF, for large-scale scene neural representation. The core of our method is a large-scale point cloud super-resolution diffusion module that enhances the sparse point cloud reconstructed from several training images into a dense point cloud as an explicit prior. Then in the rendering stage, only sampling points with prior points within the sampling radius are retained. That is, the sampling space is reduced from the unbounded space to the scene surface. Meanwhile, to fill in the background of the scene that cannot be provided by point clouds, the region sampling based on Mip-NeRF 360 is employed to model the background representation. Expensive experiments have demonstrated the effectiveness of our method for large-scale scene novel view synthesis, which outperforms relevant state-of-the-art baselines.

YNICL Journal 2022 Journal Article

Altered dynamic amplitude of low-frequency fluctuation between bipolar type I and type II in the depressive state

  • Wen Liu
  • Xiaowei Jiang
  • Zijing Deng
  • Linna Jia
  • Qikun Sun
  • Lingtao Kong
  • Feng Wu
  • Yanqing Tang

BACKGROUND: Bipolar disorder is a chronic and highly recurrent mental disorder that can be classified as bipolar type I (BD I) and bipolar type II (BD II). BD II is sometimes taken as a milder form of BD I or even doubted as an independent subtype. However, the fact that symptoms and severity differ in patients with BD I and BD II suggests different pathophysiologies and underlying neurobiological mechanisms. In this study, we aimed to explore the shared and unique functional abnormalities between subtypes. METHODS: The dynamic amplitude of low-frequency fluctuation (dALFF) was performed to compare 31 patients with BD I, 32 with BD II, and 79 healthy controls (HCs). Global dALFF was calculated using sliding-window analysis. Group differences in dALFF among the 3 groups were compared using analysis of covariance (ANCOVA), with covariates of age, sex, years of education, and mean FD, and Bonferroni correction was applied for post hoc analysis. Pearson and Spearman's correlations were conducted between clusters with significant differences and clinical features in the BD I and BD II groups, after which false error rate (FDR) was used for correction. RESULTS: We found a significant decrease in dALFF values in BD patients compared with HCs in the following brain regions: the bilateral-side inferior frontal gyrus (including the triangular, orbital, and opercular parts), inferior temporal gyrus, the medial part of the superior frontal gyrus, middle frontal gyrus, anterior cingulum, insula gyrus, lingual gyrus, calcarine gyrus, precuneus gyrus, cuneus gyrus, left-side precentral gyrus, postcentral gyrus, inferior parietal gyrus, superior temporal pole gyrus, middle temporal gyrus, middle occipital gyrus, superior occipital gyrus and right-side fusiform gyrus, parahippocampal gyrus, hippocampus, middle cingulum, orbital part of the medial frontal gyrus and superior frontal gyrus. Unique alterations in BD I were observed in the right-side supramarginal gyrus and postcentral gyrus. In addition, dALFF values in BD II were significantly higher than those in BD I in the right superior temporal gyrus and middle temporal gyrus. The variables of dALFF correlated with clinical characteristics differently according to the subtypes, but no correlations survived after FDR correction. LIMITATIONS: Our study was cross-sectional. Most of our patients were on medication, and the sample was limited. CONCLUSIONS: Our findings demonstrated neurobiological characteristics of BD subtypes, providing evidence for BD II as an independent existence, which could be the underlying explanation for the specific symptoms and/or severity and point to potential biomarkers for the differential diagnosis of bipolar subtypes.

NeurIPS Conference 2022 Conference Paper

Coordinates Are NOT Lonely - Codebook Prior Helps Implicit Neural 3D representations

  • Fukun Yin
  • Wen Liu
  • Zilong Huang
  • Pei Cheng
  • Tao Chen
  • Gang Yu

Implicit neural 3D representation has achieved impressive results in surface or scene reconstruction and novel view synthesis, which typically uses the coordinate-based multi-layer perceptrons (MLPs) to learn a continuous scene representation. However, existing approaches, such as Neural Radiance Field (NeRF) and its variants, usually require dense input views (i. e. 50-150) to obtain decent results. To relive the over-dependence on massive calibrated images and enrich the coordinate-based feature representation, we explore injecting the prior information into the coordinate-based network and introduce a novel coordinate-based model, CoCo-INR, for implicit neural 3D representation. The cores of our method are two attention modules: codebook attention and coordinate attention. The former extracts the useful prototypes containing rich geometry and appearance information from the prior codebook, and the latter propagates such prior information into each coordinate and enriches its feature representation for a scene or object surface. With the help of the prior information, our method can render 3D views with more photo-realistic appearance and geometries than the current methods using fewer calibrated images available. Experiments on various scene reconstruction datasets, including DTU and BlendedMVS, and the full 3D head reconstruction dataset, H3DS, demonstrate the robustness under fewer input views and fine detail-preserving capability of our proposed method.

AAAI Conference 2021 Conference Paper

Appearance-Motion Memory Consistency Network for Video Anomaly Detection

  • Ruichu Cai
  • Hao Zhang
  • Wen Liu
  • Shenghua Gao
  • Zhifeng Hao

Abnormal event detection in the surveillance video is an essential but the challenging task and many methods have been proposed to deal with this problem. The previous methods either only considers the appearance information or directly integrate the results of appearance and motion information without considering their endogenous consistency semantic explicitly. Inspired by the rule that humans identify the abnormal frames from multi-modality signals, we propose an Appearance-Motion Memory Consistency Network (AMMC-Net). Our method first makes full use of the prior knowledge of appearance and motion signals to capture the correspondence between them in the high-level feature space explicitly. Then, it combines the multi-view features to obtain a more essential and robust feature representation of regular events, which can significantly increase the gap between an abnormal and a regular event. In the anomaly detection phase, we further introduce a commit error in the latent space joint with the prediction error in pixel space to enhance the detection accuracy. Solid experimental results on various standard datasets validate the effectiveness of our approach.

TCS Journal 2021 Journal Article

Approximation algorithms for the submodular edge cover problem with submodular penalties

  • Xin Wang
  • Suogang Gao
  • Bo Hou
  • Lidong Wu
  • Wen Liu

In this paper, we consider the submodular edge cover problem with submodular penalties. In this problem, we are given an undirected graph G = ( V, E ) with vertex set V and edge set E. Assume the covering cost function c: 2 E → R + and the penalty function p: 2 V → R + are both submodular with p non-decreasing, c ( ∅ ) = 0 and p ( ∅ ) = 0. The goal of the submodular edge cover problem with submodular penalties is to select an edge subset to cover some vertices and penalize the vertex subset containing uncovered vertices such that the total cost of covering and penalty is minimized. For this problem, we first give a 2Δ-approximation algorithm by using a primal-dual technique, where Δ is the maximal degree of the graph G. Then we transform this problem into a submodular set cover problem, and by applying a known result for the submodular set cover problem we conclude that there is an approximation algorithm with an approximation ratio Δ + 1.

TCS Journal 2021 Journal Article

On approximation algorithm for the edge metric dimension problem

  • Yufei Huang
  • Bo Hou
  • Wen Liu
  • Lidong Wu
  • Stephen Rainwater
  • Suogang Gao

In this paper, we study the edge metric dimension problem (EMDP). We establish a potential function and give a corresponding greedy algorithm with approximation ratio 1 + ln ⁡ n + ln ⁡ ( log 2 ⁡ n ), where n is the number of vertices in the graph G.

TCS Journal 2020 Journal Article

An approximation algorithm for the k-prize-collecting multicut on a tree problem

  • Xin Hou
  • Wen Liu
  • Bo Hou

In this paper, we consider the k-prize-collecting multicut on a tree (k-PCM(T)) problem. In this problem, we are given an undirected tree T = ( V, E ), a set of m distinct pairs of vertices P = { ( s 1, t 1 ), …, ( s m, t m ) } and a parameter k with k ≤ m. Every edge in E has a nonnegative cost c e. Every pair ( s i, t i ) in P has a nonnegative penalty cost π i. Our goal is to find a multicut M that separates at least k pairs in P such that the total cost, including the edge cost of the multicut M and the penalty cost of the pairs not separated by M, is minimized. This problem generalizes the well-known multicut on a tree problem. Our main work is to present a ( 4 + ε ) -approximation algorithm for the k-PCM(T) problem via the methods of primal-dual and Lagrangean relaxation, where ε is any fixed positive number.

IJCAI Conference 2019 Conference Paper

Margin Learning Embedded Prediction for Video Anomaly Detection with A Few Anomalies

  • Wen Liu
  • Weixin Luo
  • Zhengxin Li
  • Peilin Zhao
  • Shenghua Gao

Classical semi-supervised video anomaly detection assumes that only normal data are available in the training set because of the rare and unbounded nature of anomalies. It is obviously, however, these infrequently observed abnormal events can actually help with the detection of identical or similar abnormal events, a line of thinking that motivates us to study open-set supervised anomaly detection with only a few types of abnormal observed events and many normal events available. Under the assumption that normal events can be well predicted, we propose a Margin Learning Embedded Prediction (MLEP) framework. There are three features in MLEP- based open-set supervised video anomaly detection: i) we customize a video prediction framework that favors the prediction of normal events and distorts the prediction of abnormal events; ii) The margin learning framework learns a more compact normal data distribution and enlarges the margin between normal and abnormal events. Since abnormal events are unbounded, our framework consequently helps with the detection of abnormal events, even for anomalies that have never been previously observed. Therefore, our framework is suitable for the open-set supervised anomaly detection setting; iii) our framework can readily handle both frame-level and video-level anomaly annotations. Considering that video-level anomaly detection is more easily annotated in practice and that anomaly detection with a few anomalies is a more practical setting, our work thus pushes the application of anomaly detection towards real scenarios. Extensive experiments validate the effectiveness of our framework for anomaly detection.

v2026.09.13