Arrow Research search

Author name cluster

Mingyang Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

19 papers
1 author row

Possible papers

19

AAAI Conference 2026 Conference Paper

DcSplat: Dual-Constraint Human Gaussian Splatting with Latent Multi-View Consistency

  • Tengfei Xiao
  • Yue Wu
  • Zhigang Gao
  • Yongzhe Yuan
  • Can Qin
  • Hao Li
  • Mingyang Zhang

Human Novel View Synthesis (HNVS) aims to synthesize photorealistic human images from novel viewpoints given observations from known views. Despite significant advances achieved by existing methods such as NeRF, diffusion models, and 3DGS, they still face substantial challenges in achieving stable modeling from a single image. In this paper, we introduce Dual-Constraint Human Gaussian Splatting (DcSplat), a novel, simple, and efficient 3D Gaussian-based framework for single-view 3D human reconstruction. To address occlusion-induced texture missing and depth ambiguities, we introduce two key components: a Latent Multi-View Consistency Constraint Mechanism and a Geometric Constraint Module. The former employs a Latent-space Appearance Transformer (LatentFormer) to learn semantically coherent, view-consistent appearance priors via SMPL-guided pseudo-view fusion. The latter refines noisy SMPL-based depth through a U-Net-like structure conditioned on latent appearance features. These two modules are jointly optimized to generate high-quality Gaussian parameters in a unified latent space. Extensive experiments demonstrate that DcSplat outperforms existing SOTA methods in both geometry and texture quality, while achieving fast inference and lower computational cost.

EAAI Journal 2026 Journal Article

Uncertainty-aware vessel trajectory prediction for heterogeneous data fusion in internet of things-driven smart waterways

  • Yuxu Lu
  • Kaisen Yang
  • Dong Yang
  • Haifeng Ding
  • Jinxian Weng
  • Mingyang Zhang

The rapid growth of inland waterway transportation has brought significant challenges in ensuring navigational safety and operational efficiency. Internet of things (IoT)-driven smart waterways systems present innovative solutions to enhance safety, efficiency, and sustainability in inland navigation. Nevertheless, integrating automatic identification system (AIS) and closed-circuit television (CCTV) data remains challenging due to AIS’s low and inconsistent update frequency. It leads to temporal misalignment with real-time video streams, especially in dense or complex scenarios. This work proposes a novel framework to address these challenges with vessel trajectory prediction and heterogeneous data fusion. For trajectory prediction, a spatio-temporal Transformer-based encoder extracts vessel motion features, while a variational autoencoder is employed to capture uncertainties in future trajectories. A gated recurrent unit-driven decoder iteratively predicts high-precision and robust trajectories. The framework aligns AIS-predicted trajectories with CCTV-detected targets using probabilistic matching. It calculates spatio-temporal similarities by combining AIS trajectory distributions with CCTV Gaussian representations. Spatio-temporal similarity computation and optimal assignment-based matching ensure robust and precise correspondences. Extensive experiments on real-world datasets demonstrate that our method improves overall trajectory prediction accuracy by approximately 5. 0% compared to the traditional baseline and achieves a matching accuracy of 82. 45% in dense traffic scenarios. The proposed framework exhibits excellent robustness under high-density and complex navigation conditions, providing strong support for the practical deployment of IoT-driven smart waterway systems.

AAAI Conference 2026 Conference Paper

Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training

  • Ruicheng Zhang
  • Jun Zhou
  • Zunnan Xu
  • Zihao Liu
  • Jiehui Huang
  • Mingyang Zhang
  • Yu Sun
  • Xiu Li

Trajectory-Guided image-to-video (I2V) generation aims to synthesize videos that adhere to user-specified motion instructions. Existing methods typically rely on computationally expensive fine-tuning on scarce annotated datasets. Although some zero-shot methods attempt to trajectory control in the latent space, they may yield unrealistic motion by neglecting 3D perspective and creating a misalignment between the manipulated latents and the network's noise predictions. To address these challenges, we introduce Zo3T, a novel zero-shot test-time-training framework for trajectory-guided generation with three core innovations: First, we incorporate a 3D-Aware Kinematic Projection, leveraging inferring scene depth to derive perspective-correct affine transformations for target regions. Second, we introduce Trajectory-Guided Test-Time LoRA, a mechanism that dynamically injects and optimizes ephemeral LoRA adapters into the denoising network alongside the latent state. Driven by a regional feature consistency loss, this co-adaptation effectively enforces motion constraints while allowing the pre-trained model to locally adapt its internal representations to the manipulated latent, thereby ensuring generative fidelity and on-manifold adherence. Finally, we develop Guidance Field Rectification, which refines the denoising evolutionary path by optimizing the conditional guidance field through a one-step lookahead strategy, ensuring efficient generative progression towards the target trajectory. Zo3T significantly enhances 3D realism and motion accuracy in trajectory-controlled I2V generation, demonstrating superior performance over existing training-based and zero-shot approaches.

EAAI Journal 2025 Journal Article

A hybrid deep learning method for the real-time prediction of collision damage consequences in operational conditions

  • Mingyang Zhang
  • Hongdong Wang
  • Fabien Conti
  • Teemu Manderbacka
  • Heikki Remes
  • Spyros Hirdaris

Ship collisions can result in catastrophic outcomes, necessitating effective real-time collision risk assessment methods for proactive risk management. These methods need to rapidly evaluate both the probability of collision and the potential damage dimensions (length, height, and penetration) in real conditions. Existing frameworks often underestimate collision damage consequences during operational risk assessments. This paper presents a hybrid deep learning approach for the real-time prediction of collision damage dimensions under real ship operation conditions. Collision scenarios are identified using Automatic Identification System (AIS) data, with damage extents simulated through the Super Element (SE) method. A comprehensive database of collision scenarios and corresponding damage assessments is developed, sourced from realistic operational data of Ro-Pax ship in the Gulf of Finland. The deep learning model is trained and validated using this dataset, ensuring the model's relevance and practical applicability. Extensive comparative analyses and generalization tests demonstrate the high accuracy of the model in predicting ship collision damages in diverse ship operational conditions. In addition, traditional simulation methods for evaluating damage dimensions require approximately 10 min, whereas the trained deep learning model reduces the time to less than 0. 1 s, enabling real-time potential collision consequence assessment in real operational conditions. The proposed model may provide significant insights for ship operators, enhancing ship safety and supporting intelligent decision-making in ship operations.

AAAI Conference 2025 Conference Paper

AdvDisplay: Adversarial Display Assembled by Thermoelectric Cooler for Fooling Thermal Infrared Detectors

  • Hao Li
  • Fanggao Wan
  • Yue Su
  • Yue Wu
  • Mingyang Zhang
  • Maoguo Gong

When the current physical adversarial patches cannot deceive thermal infrared detectors, the existing techniques implement adversarial attacks from scratch, such as digital patch generation, material production, and physical deployment. Besides, it is difficult to finely regulate infrared radiation. To address these issues, this paper designs an adversarial thermal display (AdvDisplay ) by assembling thermoelectric coolers (TECs) as an array. Specifically, to reduce the gap between patches in the physical and digital worlds and decrease the power of AdvDisplay device, heat transfer loss and electric power loss are designed to guide the patch optimization. In addition, a precise temperature control scheme for AdvDisplay is proposed based on proportional-integral-derivative (PID) control. Due to the accurate temperature regulation and the reusability of AdvDisplay, our method is able to improve the attack success rate and the efficiency of physical deployments. Extensive experimental results indicate that the proposed method possesses superior adversarial effectiveness compared to other methods and demonstrates strong robustness in physical attacks.

AAAI Conference 2025 Conference Paper

Channel Merging: Preserving Specialization for Merged Experts

  • Mingyang Zhang
  • Jing Liu
  • Ganggui Ding
  • Linlin Ou
  • Xinyi Yu
  • Bohan Zhuang

Lately, the practice of utilizing task-specific fine-tuning has been implemented to improve the performance of large language models (LLM) in subsequent tasks. Through the integration of diverse LLMs, the overall competency of LLMs is significantly boosted. Nevertheless, traditional ensemble methods are notably memory-intensive, necessitating the simultaneous loading of all specialized models into GPU memory. To address the inefficiency, model merging strategies have emerged, merging all LLMs into one model to reduce the memory footprint during inference. Despite these advances, model merging often leads to parameter conflicts and performance decline as the number of experts increases. Previous methods to mitigate these conflicts include post-pruning and partial merging. However, both approaches have limitations, particularly in terms of performance and storage efficiency when merged experts increase. To address these challenges, we introduce Channel Merging, a novel strategy designed to minimize parameter conflicts while enhancing storage efficiency. This method initially clusters and merges channel parameters based on their similarity to form several groups offline. By ensuring that only highly similar parameters are merged within each group, it significantly reduces parameter conflicts. During inference, we can instantly look up the expert parameters from the merged groups, preserving specialized knowledge. Our experiments demonstrate that Channel Merging consistently delivers high performance, matching unmerged models in tasks like English and Chinese reasoning, mathematical reasoning, and code generation. Moreover, it obtains results comparable to model ensemble with just 53% parameters when used with a task-specific router.

AAAI Conference 2025 Conference Paper

Decoupling Scattering: Pseudo-Label Guided NeRF for Scenes with Scattering Media

  • Mingyang Zhang
  • Junkang Zhang
  • Faming Fang
  • Guixu Zhang

Neural Radiance Fields (NeRF) has been widely used in computer vision and graphics, achieving impressive results in novel view synthesis and multi-view 3D reconstruction. However, despite its excellent performance under ideal conditions, NeRF struggles in challenging environments such as hazy, foggy, and underwater scenes, primarily due to the difficulty in decoupling objects from the scattering medium. To mitigate this limitation, we proposed a novel approach for NeRF in scenes with scattering media. Specifically, we leverage pseudo-labels during the early stage of training to guide NeRF in decoupling the densities of objects and the scattering medium, guiding the model toward a more appropriate search space. Furthermore, we introduce a Cyclical Progressive Dimensional Optimization Strategy (CPDOS) that focuses on optimizing a single or a few variables during specific periods. Experimental results demonstrate that our method can effectively simulate hazy and underwater scenes, accurately decouple the scattering medium from objects, estimate atmospheric parameters, and outperform existing methods in novel view synthesis and image restoration tasks.

AAAI Conference 2025 Conference Paper

MUCD: Unsupervised Point Cloud Change Detection via Masked Consistency

  • Yue Wu
  • Zhipeng Wang
  • Yongzhe Yuan
  • Maoguo Gong
  • Hao Li
  • Mingyang Zhang
  • Wenping Ma
  • Qiguang Miao

3D Change Detection (3DCD) has gradually become another research hotspot after image change detection. Recent works focus on using artificial labels for supervised or weakly-supervised training of siamese networks to segment changed points. However, labeling every points of multi-temporal point clouds is very expensive and time-consuming. In addition, these works lack effective self-supervised signals, and existing self-supervised signals often fail to capture sufficiently rich change information. To solve this problem, we assume that the powerful representation of 3D objects should model the consistency information of unchanged regions and distinguish different objects. Based on this assumption, we propose a new unsupervised framework called MUCD to learn change information of multi-temporal point clouds through bidirectional optimization of change segmentor and feature extractor. The training of network is divided into two stages. We first design a foreknowledge point contrastive loss based on the characteristics of the 3DCD task to initialize the feature extractor, and then propose a masked consistency loss to further learn the shared geometric information of unchanged regions in the multi-temporal point clouds, utilizing it as a free and powerful supervised signal to train a change segmentor. In the inference stage, only the segmentor is used to take multi-temporal point clouds as input and produce change segmentation result. Extensive experiments are conducted on SLPCCD and Urb3DCD, two real-world datasets of streets and urban buildings, to verify that our proposed unsupervised method is highly competitive and even outperforms supervised methods in scenes where semantic information changes occur, exhibiting better performance in generalization ability and robustness.

NeurIPS Conference 2025 Conference Paper

PointTruss: K-Truss for Point Cloud Registration

  • Yue Wu
  • Jun Jiang
  • Yongzhe Yuan
  • Maoguo Gong
  • Qiguang Miao
  • Hao Li
  • Mingyang Zhang
  • Wenping Ma

Point cloud registration is a fundamental task in 3D computer vision. Recent advances have shown that graph-based methods are effective for outlier rejection in this context. However, existing clique-based methods impose overly strict constraints and are NP-hard, making it difficult to achieve both robustness and efficiency. While the k-core reduces computational complexity, which only considers node degree and ignores higher-order topological structures such as triangles, limiting its effectiveness in complex scenarios. To overcome these limitations, we introduce the $k$-truss from graph theory into point cloud registration, leveraging triangle support as a constraint for inlier selection. We further propose a consensus voting-based low-scale sampling strategy to efficiently extract the structural skeleton of the point cloud prior to $k$-truss decomposition. Additionally, we design a spatial distribution score that balances coverage and uniformity of inliers, preventing selections that concentrate on sparse local clusters. Extensive experiments on KITTI, 3DMatch, and 3DLoMatch demonstrate that our method consistently outperforms both traditional and learning-based approaches in various indoor and outdoor scenarios, achieving state-of-the-art results.

EAAI Journal 2025 Journal Article

Ship fuel consumption prediction based on transfer learning: Models and applications

  • Xi Luo
  • Mingyang Zhang
  • Yi Han
  • Ran Yan
  • Shuaian Wang

Data-driven fuel consumption rate (FCR) prediction models largely depend on the amount of training data, which can be scarce for new ships with limited operating time. To tackle this issue, we implement three transfer learning strategies to leverage knowledge from another seven container ships to construct artificial neural network (ANN)-based FCR prediction models for a target ship with limited data. Numerical experiments reveal that the ANN models incorporating the three transfer strategies outperform the model trained solely on the target ship data, reducing mean absolute percentage error by 12. 57%, 6. 44%, and 16. 03%, respectively. This study also investigates the impacts of target dataset size on the performance of transfer strategies using ship FCR prediction as an example, revealing that the smaller amount of available data, the greater improvement in prediction accuracy using the transfer strategy. These insights contribute to the development of effective operational solutions for enhancing ship energy efficiency and promoting sustainable shipping practices.

EAAI Journal 2024 Journal Article

A data mining-then-predict method for proactive maritime traffic management by machine learning

  • Zhao Liu
  • Wanli Chen
  • Cong Liu
  • Ran Yan
  • Mingyang Zhang

Proactive traffic management is increasingly critical in maritime intelligent transportation systems. Central to this is maritime traffic forecasting, which leverages specific structures and properties of the problem. This study focuses on the traffic dynamics within convergent areas of inland waterways and proposes a method based on data mining followed by prediction using Automatic Identification System (AIS) data. This approach addresses uncertainties in ship voyage destinations and optimizes predictions for temporary stops in inland waterways. AIS data is processed to depict complete ship motion trajectories, grouping them into trajectory sets based on shared origin, destination, and route. These groups help represent maritime traffic patterns using the entrance and exit points of channels and the boundaries of the study area. Additionally, a stop detection model is applied to these trajectories to identify nodes within maritime traffic networks. A decision tree algorithm is then employed to train a classifier for predicting traffic patterns. The method was validated in the convergent area of the Yangtze River and the Hanjiang River, demonstrating effective pattern extraction from inland maritime traffic and high accuracy in predicting single ship trajectories, achieving a 96. 7% accuracy rate and 80. 9% precision. The findings suggest that the proposed method (1) effectively extracts and predicts traffic patterns, (2) identifies congestion in convergent waters, and (3) supports traffic management strategies.

EAAI Journal 2024 Journal Article

A deep learning method for the prediction of ship fuel consumption in real operational conditions

  • Mingyang Zhang
  • Nikolaos Tsoulakos
  • Pentti Kujala
  • Spyros Hirdaris

In recent years, the European Commission and the International Maritime Organization (IMO) implemented various operational measures and policies to reduce ship fuel consumption and related emissions. The effectiveness of these measures relies upon developing accurate predictive models encompassing the influence of real operational conditions. This paper presents a deep learning method for the prediction of ship fuel consumption. The method utilizes big data analytics from sensors, voyage reporting and hydrometeorological data, comprising of 266 variables made available following sea trials of a Kamsarmax bulk carrier of Laskaridis Shipping Co. Ltd. A variable importance estimation model using a Decision Tree (DT) is used to understand the underlying relationships in the available dataset. Consequently, a deep learning model is developed to understand the influence of sailing speed, heading, displacement/draft, trim, weather, sea conditions, etc. on ship fuel consumption (SFC). This is achieved by incorporating attention mechanism into Bi-directional Long Short-Term Memory (Bi-LSTM) network. The potential of the new method is demonstrated by training data streams corresponding to real ship fuel consumption rates as well as internal and external operational conditions. A comprehensive comparison with existing methods indicates that the Bi-LSTM with attention mechanism presents the best fit when using high frequency data. It is concluded that subject to further testing and validation the method could be used for the development of decision support systems for monitoring environmentally sustainable ship operations.

EAAI Journal 2024 Journal Article

A hybrid deep learning method for the prediction of ship time headway using automatic identification system data

  • Quandang Ma
  • Xu Du
  • Cong Liu
  • Yuting Jiang
  • Zhao Liu
  • Zhe Xiao
  • Mingyang Zhang

Ship Time Headway (STH) is used in maritime navigation to describe the time interval between the arrivals of two consecutive ships in the same water area. This measurement may offer a straightforward way to gauge the frequency of ship traffic and the likelihood of congestion in a particular area. STH is an important factor in understanding and managing the dynamics of ship movements in busy waterways. This paper introduces a hybrid deep learning method for predicting STH in time domain. The method integrates the Seasonal-Trend Decomposition using Loess (STL), Multi-head Self-Attention (MSA) mechanism into Long Short-Term Memory (LSTM) neural network. The STH dataset was extracted from the Automatic Identification System (AIS) through ship trajectory spatial motion, and the seasonal, trend and residual components of the decomposition were then determined from the STH dataset using the STL algorithms. MSA-LSTM is adopted to comprehensively capture the evolving patterns of STH from the sequence. Comparison studies with existing methods demonstrate the accuracy and robustness of the predictions provided by this method, indicating that the proposed method outperforms other models in terms of prediction performance and learning capabilities. By predicting STH, the method offers potential to assist maritime traffic managers and navigators in assessing ship flow, thereby enabling them to make informed decisions on navigation safety and efficiency.

AAAI Conference 2024 Conference Paper

Enhancing Hyperspectral Images via Diffusion Model and Group-Autoencoder Super-resolution Network

  • Zhaoyang Wang
  • Dongyang Li
  • Mingyang Zhang
  • Hao Luo
  • Maoguo Gong

Existing hyperspectral image (HSI) super-resolution (SR) methods struggle to effectively capture the complex spectral-spatial relationships and low-level details, while diffusion models represent a promising generative model known for their exceptional performance in modeling complex relations and learning high and low-level visual features. The direct application of diffusion models to HSI SR is hampered by challenges such as difficulties in model convergence and protracted inference time. In this work, we introduce a novel Group-Autoencoder (GAE) framework that synergistically combines with the diffusion model to construct a highly effective HSI SR model (DMGASR). Our proposed GAE framework encodes high-dimensional HSI data into low-dimensional latent space where the diffusion model works, thereby alleviating the difficulty of training the diffusion model while maintaining band correlation and considerably reducing inference time. Experimental results on both natural and remote sensing hyperspectral datasets demonstrate that the proposed method is superior to other state-of-the-art methods both visually and metrically.

JBHI Journal 2024 Journal Article

Few-Shot Class-Incremental Learning for Medical Time Series Classification

  • Le Sun
  • Mingyang Zhang
  • Benyou Wang
  • Prayag Tiwari

Continuously analyzing medical time series as new classes emerge is meaningful for health monitoring and medical decision-making. Few-shot class-incremental learning (FSCIL) explores the classification of few-shot new classes without forgetting old classes. However, little of the existing research on FSCIL focuses on medical time series classification, which is more challenging to learn due to its large intra-class variability. In this paper, we propose a framework, the Meta self-Attention Prototype Incrementer (MAPIC) to address these problems. MAPIC contains three main modules: an embedding encoder for feature extraction, a prototype enhancement module for increasing inter-class variation, and a distance-based classifier for reducing intra-class variation. To mitigate catastrophic forgetting, MAPIC adopts a parameter protection strategy in which the parameters of the embedding encoder module are frozen at incremental stages after being trained in the base stage. The prototype enhancement module is proposed to enhance the expressiveness of prototypes by perceiving inter-class relations using a self-attention mechanism. We design a composite loss function containing the sample classification loss, the prototype non-overlapping loss, and the knowledge distillation loss, which work together to reduce intra-class variations and resist catastrophic forgetting. Experimental results on three different time series datasets show that MAPIC significantly outperforms state-of-the-art approaches by 27. 99%, 18. 4%, and 3. 95%, respectively.

NeurIPS Conference 2024 Conference Paper

On provable privacy vulnerabilities of graph representations

  • Ruofan Wu
  • Guanhua Fang
  • Mingyang Zhang
  • Qiying Pan
  • Tengfei Liu
  • Weiqiang Wang

Graph representation learning (GRL) is critical for extracting insights from complex network structures, but it also raises security concerns due to potential privacy vulnerabilities in these representations. This paper investigates the structural vulnerabilities in graph neural models where sensitive topological information can be inferred through edge reconstruction attacks. Our research primarily addresses the theoretical underpinnings of similarity-based edge reconstruction attacks (SERA), furnishing a non-asymptotic analysis of their reconstruction capacities. Moreover, we present empirical corroboration indicating that such attacks can perfectly reconstruct sparse graphs as graph size increases. Conversely, we establish that sparsity is a critical factor for SERA's effectiveness, as demonstrated through analysis and experiments on (dense) stochastic block models. Finally, we explore the resilience of private graph representations produced via noisy aggregation (NAG) mechanism against SERA. Through theoretical analysis and empirical assessments, we affirm the mitigation of SERA using NAG. In parallel, we also empirically delineate instances wherein SERA demonstrates both efficacy and deficiency in its capacity to function as an instrument for elucidating the trade-off between privacy and utility.

TIST Journal 2023 Journal Article

You Are How You Use Apps: User Profiling Based on Spatiotemporal App Usage Behavior

  • Tong Li
  • Yong Li
  • Mingyang Zhang
  • Sasu Tarkoma
  • Pan Hui

Mobile apps have become an indispensable part of people’s daily lives. Users determine what apps to use and when and where to use them based on their tastes, interests, and personal demands, depending on their personality traits. This article aims to infer user profiles from their spatiotemporal mobile app usage behavior. Specifically, we first transform mobile app usage records into a heterogeneous graph. On the graph, nodes represent users, apps, locations, and time slots. Edges describe the co-occurrence of entities in usage records. We then develop a multi-relational heterogeneous graph attention network (MRel-HGAN), an end-to-end system for user profiling. MRel-HGAN first adopts a neighbor sampling strategy based on bootstrapping to sample heavily connected neighbors of a fixed size for each node. Next, we design a relational graph convolutional operation and a multi-relational attention operation. Through such modules, MRel-HGAN can generate node embedding by sufficiently leveraging the rich semantic information of the multi-relational structure in the mobile app usage graph. Experimental results on real-world mobile app usage datasets show the effectiveness and superiority of our MRel-HGAN in the user profiling task for attributes of gender and age.

IJCAI Conference 2020 Conference Paper

Multi-View Joint Graph Representation Learning for Urban Region Embedding

  • Mingyang Zhang
  • Tong Li
  • Yong Li
  • Pan Hui

The increasing amount of urban data enable us to investigate urban dynamics, assist urban planning, and eventually, make our cities more livable and sustainable. In this paper, we focus on learning an embedding space from urban data for urban regions. For the first time, we propose a multi-view joint learning model to learn comprehensive and representative urban region embeddings. We first model different types of region correlations based on both human mobility and inherent region properties. Then, we apply a graph attention mechanism in learning region representations from each view of the built correlations. Moreover, we introduce a joint learning module that boosts the region embedding learning by sharing cross-view information and fuses multi-view embeddings by learning adaptive weights. Finally, we exploit the learned embeddings in the downstream applications of land usage classification and crime prediction in urban areas with real-world data. Extensive experiment results demonstrate that by exploiting our proposed joint learning model, the performance is improved by a large margin on both tasks compared with the state-of-the-art methods.

IJCAI Conference 2019 Conference Paper

A Decomposition Approach for Urban Anomaly Detection Across Spatiotemporal Data

  • Mingyang Zhang
  • Tong Li
  • Hongzhi Shi
  • Yong Li
  • Pan Hui

Urban anomalies such as abnormal flow of crowds and traffic accidents could result in loss of life or property if not handled properly. Detecting urban anomalies at the early stage is important to minimize the adverse effects. However, urban anomaly detection is difficult due to two challenges: a) the criteria of urban anomalies varies with different locations and time; b) urban anomalies of different types may show different signs. In this paper, we propose a decomposing approach to address these two challenges. Specifically, we decompose urban dynamics into the normal component and the abnormal component. The normal component is merely decided by spatiotemporal features, while the abnormal component is caused by anomalous events. Then, we extract spatiotemporal features and estimate the normal component accordingly. At last, we derive the abnormal component to identify anomalies. We evaluate our method using both real-world and synthetic datasets. The results show our method can detect meaningful events and outperforms state-of-the-art anomaly detecting methods by a large margin.

v2026.09.13