Arrow Research search

Author name cluster

Jie Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

85 papers
2 author rows

Possible papers

85

EAAI Journal 2026 Journal Article

A high-performance memetic algorithm for integrated process planning and shop floor scheduling considering heterogeneous automated guided vehicle transportation system

  • Yang Zhou
  • Guiliang Gong
  • Xiahui Liu
  • Zhipeng Yuan
  • Hongbo Huang
  • Ao He
  • Jie Li

Integrating transportation decisions into Integrated Process Planning and Shop Scheduling (IPPS) enables tighter coordination between machining operations and material handling, leading to more efficient production schedules. However, most existing IPPS studies assume homogeneous or unlimited transportation resources, and thus overlook the practical scheduling complexity introduced by heterogeneous and capacity-limited automated guided vehicle (AGV) fleets. This study proposes an IPPS model with a heterogeneous AGV transportation system (IPPSHT), in which process route selection, operation sequencing, machine assignment, and AGV allocation are jointly optimized under functional compatibility constraints between AGVs and machines. To solve the resulting bi-objective problem of minimizing makespan and total energy consumption across machining and transportation stages, we develop a high-performance memetic algorithm that combines tailored encoding/decoding, dominance-based selection, and problem-aware local improvement. Extensive experiments on 60 benchmark instances and an engineering case study demonstrate that the proposed method consistently outperforms three representative multi-objective evolutionary algorithms in terms of solution quality and statistical significance, indicating its effectiveness for complex manufacturing settings with heterogeneous material-handling resources.

AAAI Conference 2026 Conference Paper

CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World

  • Yating Yu
  • Congqi Cao
  • Zhaoying Wang
  • Weihua Meng
  • Jie Li
  • Yuxin Li
  • Zihao Wei
  • Zhongpei Shen

How far are deep models from real-world video anomaly understanding (VAU)? Current works typically emphasize detecting unexpected occurrences deviating from normal patterns or comprehending anomalous events with interpretable descriptions. However, they exhibit only a superficial comprehension of real-world anomalies, with limited breadth in complex principles and subtle contexts that distinguish the anomalies from normalities, e.g., climbing cliffs with safety gear vs. without it. To this end, we introduce CueBench, the first of its kind Benchmark, devoted to Context-aware video anomalies within a Unified Evaluation framework. We comprehensively establish an event-centric hierarchical taxonomy that anchors two core event types: 14 conditional and 18 absolute anomaly events, defined by their refined semantics from diverse contexts across 174 scenes and 198 attributes. Based on this, we propose to unify and benchmark context-aware VAU with various challenging tasks across recognition, temporal grounding, detection, and anticipation. It also serves as a rigorous and fair probing evaluation suite for generalized and specialized vision-language models (VLMs) across both generative and discriminative paradigms. To address the challenges underlying CueBench, we further develop Cue-R1 based on R1-style reinforcement fine-tuning with verifiable, task-aligned, and hierarchy-refined rewards in a unified generative manner. Extensive results on CueBench reveal that, existing VLMs are still far from satisfactory real-world anomaly understanding, while our Cue-R1 surpasses these state-of-the-art approaches by over 24% on average.

AAAI Conference 2026 Conference Paper

DIFFA: Large Language Diffusion Models Can Listen and Understand

  • Jiaming Zhou
  • Hongjie Chen
  • Shiwan Zhao
  • Jian Kang
  • Jie Li
  • Enzhi Wang
  • Yujie Guo
  • Haoqin Sun

Recent advances in large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, large language diffusion models have emerged as a promising alternative to the autoregressive paradigm, offering improved controllability, bidirectional context modeling, and robust generation. However, their application to the audio modality remains underexplored. In this work, we introduce DIFFA, the first diffusion-based large audio-language model designed to perform spoken language understanding. DIFFA integrates a frozen diffusion language model with a lightweight dual-adapter architecture that bridges speech understanding and natural language reasoning. We employ a two-stage training pipeline: first, aligning semantic representations via an ASR objective; then, learning instruction-following abilities through synthetic audio-caption pairs automatically generated by prompting LLMs. Despite being trained on only 960 hours of ASR and 127 hours of synthetic instruction data, DIFFA demonstrates competitive performance on major benchmarks, including MMSU, MMAU, and VoiceBench, outperforming several autoregressive open-source baselines. Our results reveal the potential of large language diffusion models for efficient and scalable audio understanding, opening a new direction for speech-driven AI.

YNIMG Journal 2026 Journal Article

Dual neural mechanisms of sustained response inhibition: Right-lateralized core control and left-lateralized adaptive support

  • Liyue Lin
  • Jiahe Sun
  • Wei Xiong
  • Jiayi Zhao
  • Yishu Chen
  • Min Li
  • Xi Li
  • Ruyan Jiao

Sustained inhibition is critical for adaptive behavioral control in complex environments, yet its neural underpinnings remain poorly understood. Using a sequence-selective stop-signal task, we hypothesized that increasing stop-signal delay (SSD) would elicit distinct behavioral dynamics and recruit dissociable neural systems supporting different stages of sustained inhibition. fMRI data from 26 participants across three SSD conditions (no-, short-, and long-delay), combined with an activation likelihood estimation (ALE) meta-analysis of 64 classical stop-signal studies, enabled us to contrast sustained and transient inhibition. Key behavioral results revealed a systematic decrease in response time (RT) with increasing SSD, while inhibition execution time (IT/ET) followed an inverted U-shaped pattern. Neuroimaging findings demonstrated that sustained inhibition engages a significantly broader bilateral network compared to transient inhibition, spanning prefrontal, parietal, and subcortical regions. Further analyses revealed a dual-mechanism inhibitory network in which early go-stop competition recruited bilateral inhibitory regions, with the left hemisphere providing complementary support, whereas late-stage emergency stopping relied primarily on a right-dominant prefrontal pathway. Together, these findings establish sustained inhibition as a distinct and dynamically organized control process, providing a novel framework for understanding how the brain flexibly regulates behavior under evolving temporal demands.

AAAI Conference 2026 Conference Paper

FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich Training

  • Fuhan Cai
  • Yong Guo
  • Jie Li
  • Wenbo Li
  • Jian Chen
  • Xiangzhong Fang

Recent advancements in text-to-image (T2I) generation have led to the emergence of highly expressive models such as diffusion transformers (DiTs), exemplified by FLUX. However, their massive parameter sizes lead to slow inference, high memory usage, and poor deployability. Existing acceleration methods (e.g., single-step distillation and attention pruning) often suffer from significant performance degradation and incur substantial training costs. To address these limitations, we propose FastFLUX, an architecture-level pruning framework designed to enhance the inference efficiency of FLUX. At its core is the Block-wise Replacement with Linear Layers (BRLL) method, which replaces structurally complex residual branches in ResBlocks with lightweight linear layers while preserving the original shortcut connections for stability. Furthermore, we introduce Sandwich Training (ST), a localized fine-tuning strategy that leverages LoRA to supervise neighboring blocks, mitigating performance drops caused by structural replacement. Experiments show that our FastFLUX maintains high image quality under both qualitative and quantitative evaluations, while significantly improving inference speed, even with 20% of the hierarchy pruned.

JBHI Journal 2026 Journal Article

Few-Shot Class-Incremental Learning With Dynamic Prototype Refinement for Brain Activity Classification

  • Lei Cao
  • Hao Li
  • Yilin Dong
  • Tianyu Liu
  • Jie Li

The brain-computer interface (BCI) system facilitates efficient communication and control, with Electroencephalography (EEG) signals as a vital component. Traditional EEG signal classification, based on static deep-learning models, presents a challenge when new classes of the subject’s brain activity emerge. The goal is to develop a model that can recognize new few-shot classes while preserving its ability to discriminate between existing ones. This scenario is referred to as Few-Shot Class-Incremental Learning (FSCIL). This work introduces IncrementEEG, a novel framework meticulously designed to tackle the distinct challenges of FSCIL in EEG-based brain activity classification, focusing specifically on emotion recognition and steady-state visual evoked potential (SSVEP). Our work analyzes the role of additive angular margin loss in improving the model’s discrimination capabilities. The proposed method is designed to demonstrate robustness in open-world conditions and adaptability to new tasks. Furthermore, we introduce a prototype refinement module comprising a prototype augmentation block and an update block. The prototype augmentation block in the deep feature space preserves the decision boundary for prior tasks, and the prototype update block utilizes a shared embedding space to compute the relation matrix for bootstrapping prototype updates. Extensive experiments conducted across multiple datasets show the superior performance of the IncrementEEG framework compared to state-of-the-art methods. The proposed method advances FSCIL brain activity classification, offering promising potential for applications in Brain-Computer Interface systems.

AAAI Conference 2026 Conference Paper

FIND: A Simple Yet Effective Baseline for Diffusion-Generated Image Detection

  • Jie Li
  • Yingying Feng
  • Chi Xie
  • Jie Hu
  • Lei Tan
  • Jiayi Ji

The remarkable realism of images generated by diffusion models poses critical detection challenges. Current methods utilize reconstruction error as a discriminative feature, exploiting the observation that real images exhibit higher reconstruction errors when processed through diffusion models. However, these approaches require costly reconstruction computations and depend on specific diffusion models, making their performance highly model-dependent. We identify a fundamental difference: real images are more difficult to fit with Gaussian distributions compared to synthetic ones. In this paper, we propose Forgery Identification via Noise Disturbance (FIND), a novel method that requires only a simple binary classifier. It eliminates reconstruction by directly targeting the core distributional difference between real and synthetic images. Our key operation is to add Gaussian noise to real images during training and label these noisy versions as synthetic. This step allows the classifier to focus on the statistical patterns that distinguish real from synthetic images. We theoretically prove that the noise-augmented real images resemble diffusion-generated images in their ease of Gaussian fitting. Furthermore, simply by adding noise, they still retain visual similarity to the original images, highlighting the most discriminative distribution-related features. The proposed FIND improves performance by 11.7% on the GenImage benchmark while running 126x faster than existing methods. By removing the need for auxiliary diffusion models and reconstruction, it offers a practical, efficient, and generalizable way to detect diffusion-generated content.

TMLR Journal 2026 Journal Article

Improving Foundation Model Group Robustness with Auxiliary Sentence Embeddings

  • Sisuo Lyu
  • Hong Liu
  • Jie Li
  • Yan Teng
  • Yingchun Wang

This paper addresses the critical challenge of mitigating group-based biases in vision-language foundation models, a pressing issue for ensuring trustworthy AI deployment. We introduce DoubleCCA, a novel and computationally efficient framework that systematically enriches textual representations to enhance group robustness. Our key innovation is to leverage an auxiliary large sentence embedding model to capture diverse semantic perspectives, counteracting biased representations induced by limited training data. To this end, we propose a two-stage Canonical Correlation Analysis (DoubleCCA) technique: first, aligning augmented and original embeddings in a shared space; second, reconstructing invariant features to align with visual representations, thus enhancing the model's group robustness. We further propose a simple sentence augmentation approach, which aims to improve the robustness of CCA-induced subspaces. Our method is simple to implement and can be easily integrated into existing models, making it a practical solution for improving the robustness of vision-language foundation models to group-based biases. The experiments on a variety of datasets demonstrate that our method outperforms existing methods in terms of both performance and robustness. Our code is available at https://github.com/sisuolv/doublecca.

EAAI Journal 2026 Journal Article

Predicting stress in two-phase random materials and super-resolution method for stress images by embedding physical information

  • Tengfei Xing
  • Xiaodan Ren
  • Jie Li

Stress analysis is essential in material design. In materials with complex microstructures, such as two-phase random materials (TRMs), failure is typically associated with stress concentration at phase interfaces that govern mechanical performance. Existing Image Super-Resolution (ISR) methods are mainly data-driven and rely on supervised learning, where the achievable magnification of stress images is tightly constrained by the resolution of the training dataset. This limits their applicability for TRMs, where high-resolution stress information at phase interfaces is essential but often unavailable. In practical engineering, limited pixel resolution in microstructure images further constrains stress image clarity and hinders the observation of stress concentration regions. To address this research gap, we propose a stress prediction framework tailored for TRMs that combines microstructural information with physics-based constraints. First, the framework employs a Multiple Compositions U-net (MC U-net) to predict stress in low-resolution material microstructures. By incorporating phase interface information, the MC U-net effectively reduces prediction errors at phase interfaces. Secondly, a Mixed Physics-Informed Neural Network (MPINN)-based stress ISR method (SRMPINN) is introduced. Unlike conventional ISR methods, SRMPINN leverages physical constraints to achieve stress image super-resolution without requiring paired high-resolution training images, enabling stress images to be generated at substantially high magnification factors, including non-integer scales, with magnification ratios not restricted by the training dataset. Finally, transfer learning is applied to perform stress analysis on TRMs with different loading states and anisotropy. The results demonstrate that the proposed framework achieves high accuracy, generalization, and flexibility, particularly in resolving stress concentrations at phase interfaces.

AAAI Conference 2026 Conference Paper

Revealing the Invisible: Latent Structure Modeling for Semantically Consistent Cloud Removal

  • Jingwei Xin
  • Kai Guo
  • Jie Li
  • Nannan Wang

Cloud removal (CR) in remote sensing imagery is a critical yet challenging task due to complex cloud patterns and diverse underlying ground structures. Despite recent progress in generative models such as diffusion models, CR remains limited by their inadequate capability to perceive and reconstruct structured information beneath cloud-covered areas. In this work, we propose a Visibility-guided Semantic Estimation and Reconstruction network for cloud removal (VISER-CR), which reformulates CR as a structure-guided completion problem. Specifically, VISER-CR explicitly models cloud interference via spatial masking, encouraging the model to reason beyond pixel-level appearance and enhance scene-level structural understanding. Moreover, to further improve the representation of structural information, we introduce Patch Saliency Encoding, a self-guided mechanism that implicitly models structural alignment among patches, significantly enhancing clustering consistency and semantic separability in the latent space. This adaptive mechanism guides the network to focus on learning and reconstructing structurally important regions, thereby reducing redundancy and improving overall cloud removal performance. Extensive experiments on multiple benchmark datasets demonstrate the superior effectiveness of our method.

EAAI Journal 2026 Journal Article

Robust guaranteed neural learning-based output tracking control for uncertain nonlinear systems: An uncertainty feedback compensation method

  • Chengbo Dai
  • Jie Li
  • Zhenlong Wu
  • Yan Li
  • Donghai Li

The proven efficacy of neural network-based control schemes has spurred their application to physical systems. However, ensuring performance robustness when deploying such controllers in uncertain physical environments remains a significant challenge. This article proposes an uncertainty feedback compensation framework to guarantee the performance robustness of neural learning-based output tracking control for uncertain nonlinear systems. Active Disturbance Rejection Control (ADRC) is incorporated as an ancillary compensator, requiring only that the varying rate of uncertainty be bounded. A single critic network-based output tracking control is then developed by constructing an augmented nominal model and adaptive dynamic programming (ADP), while ADRC operates in parallel to compensate for general uncertainties in real time. Furthermore, a desired dynamic equation-based parameter tuning rule is proposed to configure ADRC for effective tracking using nominal model information. The convergence of neural network weights is established via Lyapunov analysis, and the closed-loop stability and performance robustness are further demonstrated by analyzing the boundedness of ADRC’s estimation and tracking errors under general system uncertainties. Finally, the effectiveness of the proposed method is validated through both numerical simulations and practical experiments, demonstrating substantial improvements in the safety and practical applicability of neural learning-based control.

AAAI Conference 2026 Conference Paper

Shrinking the Teacher: An Adaptive Teaching Paradigm for Asymmetric EEG-Vision Alignment

  • Lukun Wu
  • Jie Li
  • Ziqi Ren
  • Kaifan Zhang
  • Xinbo Gao

Decoding visual features from EEG signals is a central challenge in neuroscience, with cross-modal alignment as the dominant approach. We argue that the relationship between visual and brain modalities is fundamentally asymmetric, characterized by two critical gaps: a Fidelity Gap (stemming from EEG's inherent noise and signal degradation, vs. vision's high-fidelity features) and a Semantic Gap (arising from EEG's shallow conceptual representation, vs. vision's rich semantic depth). Previous methods often overlook this asymmetry, forcing alignment between the two modalities as if they were equal partners and thereby leading to poor generalization. To address this, we propose the adaptive teaching paradigm. This paradigm empowers the ``teacher" modality (vision) to dynamically shrink and adjust its knowledge structure under task guidance, tailoring its semantically dense features to match the ``student" modality (EEG)'s capacity. We implement this paradigm with the ShrinkAdapter, a simple yet effective module featuring a residual-free design and a bottleneck structure. Through extensive experiments, we validate the underlying rationale and effectiveness of our paradigm. Our method achieves a top-1 accuracy of 60.2% on the zero-shot brain-to-image retrieval task, surpassing previous state-of-the-art methods by a margin of 9.8%. Our work introduces a new perspective for asymmetric alignment: the teacher must shrink and adapt to bridge the vision-brain gap.

EAAI Journal 2026 Journal Article

Structure-aware context-enhanced and dual-path synergistic decoding network for atrophic gastritis segmentation

  • Xiaojuan Liu
  • Yilong Liu
  • Yongjun Zhu
  • Fang Huang
  • Jie Li
  • Shanxiong Chen

Chronic atrophic gastritis (CAG) is a critical precancerous lesion in the initiation and progression of gastric cancer. In white-light endoscopic images, CAG typically exhibits weak structural patterns, low contrast, and blurred boundaries, accompanied by complex background interference, which poses substantial challenges for accurate pixel-level lesion segmentation. Existing automatic segmentation methods remain limited in global context modeling, fine-grained structural recovery, and consistent constraint of weak boundaries. To address these challenges, we propose a novel Structure-aware Context-enhanced and Dual-path Synergistic Decoding Network (SCD-Net). The proposed framework incorporates three key designs: (1) a Multi-scale Context Enhancement (MACE) module that strengthens multi-scale contextual modeling to compensate for insufficient semantic information under low-contrast and complex background conditions; (2) a Dual-path Synergistic (DPS) decoder that collaboratively models global semantic representations and local detailed features, enabling more reliable structural recovery and boundary refinement in regions with blurred lesion boundaries and irregular morphologies; (3) a Structure-aware Composite Loss (SACLoss) function that constrains the predictions from the perspectives of regional consistency and structural integrity, guiding the model to produce more stable and continuous segmentation results in scenarios involving weak boundaries and multi-scale lesions. We evaluate SCD-Net on a real-world clinical CAG dataset against 14 state-of-the-art methods and further assess its generalization on four polyp datasets and one skin lesion segmentation dataset. SCD-Net consistently achieves superior performance, improving mean Intersection over Union (mIoU) scores by 1. 3% on the CAG dataset and delivering stable and competitive segmentation performance in cross-dataset evaluations, which demonstrates its strong generalization capability and robustness.

AAAI Conference 2026 Conference Paper

TokenPowerBench: Benchmarking the Power Consumption of LLM Inference

  • Chenxu Niu
  • Wei Zhang
  • Jie Li
  • Yongjian Zhao
  • Tongyang Wang
  • Xi Wang
  • Yong Chen

Large language model (LLM) services now answer billions of queries per day, and industry reports show that inference, not training, accounts for more than 90% of total power consumption. However, existing benchmarks focus on either training/fine-tuning or performance of inference and provide little support for power consumption measurement and analysis of inference. We introduce TokenPowerBench, the first lightweight and extensible benchmark designed for LLM-inference power consumption studies. The benchmark combines a declarative configuration interface covering model choice, prompt set, and inference engine, a measurement layer that captures GPU-, node-, and system-level power without specialized power meters, and a phase-aligned metrics pipeline that attributes energy to the prefill and decode stages of every request. These elements make it straightforward to explore the power consumed by an LLM inference run; furthermore, by varying batch size, context length, parallelism strategy and quantization, users can quickly assess how each setting affects joules per token and other energy-efficiency metrics. We evaluate TokenPowerBench on four of the most widely used model series (Llama, Falcon, Qwen, and Mistral). Our experiments cover from 1 billion parameters up to the frontier-scale Llama3-405B model. Furthermore, we release TokenPowerBench as open source to help users to measure power consumption, forecast operating expenses, and meet sustainability targets when deploying LLM services.

JBHI Journal 2026 Journal Article

Towards Cognitive Impairment Screening in Elderly Communities with Audio-Visual Modal Disentangled Representation Learning

  • Rui Feng
  • Hongbin Chen
  • Yihao Yao
  • Liuyu Wu
  • Tao Liang
  • Wentao Xiang
  • Jie Li
  • Chu Kiong Loo

Alzheimer's disease (AD) is pressing global health concerns, for which early diagnosis is critical to effective intervention. However, conventional approaches, including neuropsychological assessments and neuroimaging techniques, are resource-intensive and impractical for community-level screening. In contrast, artificial intelligence-driven behavioral analyses, including speech pattern and facial expression recognition, have demonstrated considerable potential for scalable and non invasive cognitive assessment. This work presents a community-oriented intelligent screening system for cognitive impairment screening in elderly populations. As a foundation, we introduce CIR-AV, the first large-scale Mandarin-based multimodal dataset for cognitive impairment recognition in Chinese older adults, encompassing 574 community-dwelling participants with comprehensive facial expression and speech data. Building upon this resource, we propose DiVA, a disentangled audio-visual fusion framework that decomposes multimodal features into shared and specific representations. A trajectory constrained mechanism enhances representation purity, while a cross-modal attention-based dynamic fusion (CMF) module adaptively balances modality contributions, ensuring robust performance under real-world conditions. Experimental results demonstrate that DiVA achieves an AUC of 78. 66% at the segment level and an accuracy of 79. 46% at the subject level, significantly outperforming state-of the-art methods. With its cost-efficient and scalable de sign, it is well-suited for large-scale community screening, providing a practical solution for early dementia detection in resource-limited settings with considerable social and economic value.

AAAI Conference 2026 Conference Paper

WenetSpeech-Yue: A Large-Scale Cantonese Speech Corpus with Multi-dimensional Annotation

  • Longhao Li
  • Zhao Guo
  • Hongjie Chen
  • Yuhang Dai
  • Ziyu Zhang
  • Hongfei Xue
  • Tianlun Zuo
  • Chengyou Wang

The development of speech understanding and generation has been significantly accelerated by the availability of large-scale, high-quality speech datasets. Among these, ASR and TTS are regarded as the most established and fundamental tasks. However, for Cantonese (Yue Chinese), spoken by approximately 84.9 million native speakers worldwide, limited annotated resources have hindered progress and resulted in suboptimal ASR and TTS performance. To address this challenge, we propose WenetSpeech-Pipe, an integrated pipeline for building large-scale speech corpus with multi-dimensional annotation tailored for speech understanding and generation. Based on this pipeline, we release WenetSpeech-Yue, the first large-scale Cantonese speech corpus with multi-dimensional annotation for ASR and TTS, covering 21,800 hours across 10 domains with annotations including ASR transcription, text confidence, speaker identity, age, gender, speech quality scores, among other annotations. We also release WSYue-eval, a comprehensive Cantonese benchmark with two components: WSYue-ASR-eval, a manually annotated set for evaluating ASR on short and long utterances, code-switching, and diverse acoustic conditions, and WSYue-TTS-eval, with base and coverage subsets for standard and generalization testing. Experimental results show that models trained on WenetSpeech-Yue achieve competitive results against state-of-the-art (SOTA) Cantonese ASR and TTS systems, including commercial and LLM-based models, highlighting the value of our dataset and pipeline.

EAAI Journal 2025 Journal Article

A faster heterogeneous parallel computing method for Tucker decomposition

  • Xiaosong Peng
  • Laurence T. Yang
  • Jie Li
  • Wenjun Jiang

Artificial intelligence (AI) technology is developing explosively in application fields such as speech recognition, semantic understanding, and computer vision. In particular, the new generation of knowledge enhancement large language models have gradually become the infrastructure of social productivity. With the development of large models towards multi-mode, the input data and model parameters characterized by tensors are becoming increasingly large. This puts high demands on computation and storage. Tucker decomposition obtains the optimal low-rank representation of the natural tensor by factor matrices and a core tensor, reducing storage and computation requirements in big data and artificial intelligence applications. However, the existing Tucker decomposition methods show limited computational speed and convergence performance. In this paper, a general heterogeneous computing framework of the Tucker decomposition method is proposed, which analyzes the row independence of factor matrices and the column independence of Kruskal matrices, and updates the Kruskal matrices instead of the core tensor in a column-wise manner and factor matrices in a row-wise manner respectively to reduce the storage overhead in computation. Further, the proposed method employs a heterogeneous computing platform to speed up the computing bottlenecks and takes full advantage of fine-grained parallel optimization technology to improve memory access efficiency. The experimental results show that the computation speed achieves a 3. 1 to 75. 4 times improvement compared to the latest methods. Among all the methods, it exhibits the best convergence performance.

IROS Conference 2025 Conference Paper

Asynchronous Harmony-based Decentralized Auctions Method for Scalable UAV Swarm

  • Runfeng Chen
  • Jie Li
  • Yiting Chen
  • Yuchong Huang
  • Zehao Xiong

Unmanned aerial vehicle (UAV) swarms find extensive applications in diverse fields, including search and rescue, logistics delivery, and environmental surveillance, necessitating meticulous task and temporal scheduling to meet intricate spatiotemporal requirements. A market-based strategy emerges as a suitable option for self-organizing swarm coordination. However, the consensus mechanisms employed by most market-based algorithms necessitate synchronous communication, leading to waiting times. Researchers have turned to asynchronous approaches for enhanced efficiency, yet the communication burden of existing asynchronous methods escalates swiftly with the growth of the swarm size. Therefore, this paper proposes an Asynchronous Harmony-based Decentralized Auctions (AHDA) method for networked UAV swarm to reduce the communication load and scheduling time required by a market-based approach. First, proximity communication is proposed to reduce the broadcast range and content of UAVs. Second, new conflict resolution protocols are designed to eliminate task conflict between UAVs faster. Third, propagation rules are designed to limit the scope of task information diffusion. Ultimately, it brings a decrease in communication load and scheduling time because it is expected to achieve the minimum requirement of no task conflict between UAVs, rather than swarm scheduling consistency. Monte Carlo simulations spanning 32 to 128 UAVs demonstrate that compared with the Asynchronous Consensus-Based Bundle Algorithm (ACBBA), the proposed AHDA achieves reductions of up to 70. 16% in transmitted messages, 75. 78% in communication traffic, and 63. 12% in scheduling time.

IROS Conference 2025 Conference Paper

Automated UAV-based Wind Turbine Blade Inspection: Blade Stop Angle Estimation and Blade Detail Prioritized Exposure Adjustment

  • Yichuan Shi
  • Hao Liu
  • Haowen Zheng
  • Haowen Yu
  • Xianqi Liang
  • Jie Li
  • Minmin Ma
  • Ximin Lyu

Unmanned aerial vehicles (UAVs) are critical in the automated inspection of wind turbine blades. Nevertheless, several issues persist in this domain. Firstly, existing inspection platforms encounter challenges in meeting the demands of automated inspection tasks and scenarios. Moreover, current blade stop angle estimation methods are vulnerable to environmental factors, restricting their robustness. Additionally, there is an absence of real-time blade detail prioritized exposure adjustment during capture, where lost details cannot be restored through post-optimization. To address these challenges, we introduce a platform and two approaches. Initially, a UAV inspection platform is presented to meet the automated inspection requirements. Subsequently, a Fermat point based blade stop angle estimation approach is introduced, achieving higher precision and success rates. Finally, we propose a blade detail prioritized exposure adjustment approach to ensure appropriate brightness and preserve details during image capture. Extensive tests, comprising over 120 flights across 10 wind turbine models in 5 operational wind farms, validate the effectiveness of the proposed approaches in enhancing inspection autonomy.

IROS Conference 2025 Conference Paper

BoRe-Depth: Self-Supervised Monocular Depth Estimation with Boundary Refinement for Embedded Systems

  • Juan Li
  • Sheng Zhang
  • Chang Liu
  • Jie Li
  • Xu Zhang

Depth estimation is one of the key technologies for realizing 3D perception in unmanned systems. Monocular depth estimation has been widely researched because of its low-cost advantage, but the existing methods face the challenges of poor depth estimation performance and blurred object boundaries on embedded systems. In this paper, we propose a novel monocular depth estimation model, BoRe-Depth, which contains only 8. 7M parameters. It can accurately estimate depth maps on embedded systems and significantly improves boundary quality. Firstly, we design an Enhanced Feature Adaptive Fusion Module (EFAF) which adaptively fuses depth features to enhance boundary detail representation. Secondly, we integrate semantic knowledge into the encoder to improve the object recognition and boundary perception capabilities. Finally, BoRe-Depth is deployed on NVIDIA Jetson Orin, and runs efficiently at 50. 7 FPS. We demonstrate that the proposed model significantly outperforms previous lightweight models on multiple challenging datasets, and we provide detailed ablation studies for the proposed methods. The code is available at https://github.com/liangxiansheng093/BoRe-Depth.

IROS Conference 2025 Conference Paper

Bridging the Reality Gap: Communication-Aware Task Allocation with Multi-Objective Asynchronous Policy Learning

  • Zehao Xiong
  • Yu Zhou
  • Yexun Xi
  • Yizhe Cao
  • Chang Wang
  • Jie Li

Distributed task allocation in the UAV swarm is sensitive to excessive communication overhead and frequent transmissions. Combining reinforcement learning and task allocation demonstrates great potential in enhancing algorithm performance and optimizing communication. However, existing studies rely on ideal communication assumptions and the nonphysical environment, making training and validation impractical in applying networked swarms. This paper proposes the Communication-Aware Task Allocation, which aims to train a gating mechanism policy to coordinate the transmission timing, improving robustness and timelessness of the task allocation. First, the policy learning problem is formalized as a POMDP, for which the channel access and other features are designed for observations, actions are inter-agent adaptive gating mechanisms, and the shared reward reflects global task conflicts. Second, to address the asynchronous learning under the CTDE, an asynchronous experience collection and splicing method is proposed to align trajectories. Then, the MOCPPO is proposed, which combines a primal-dual operator with proximal policy optimization, updating the optimal Lagrange multiplier and strategy parameters to simultaneously minimize task conflicts and communication overhead. Finally, sim-to-real experiments are conducted in the HIL environment, and results illustrate the best trade-off optimization of the proposed method over all state-of-the-art approaches.

AAAI Conference 2025 Conference Paper

Contrasting Adversarial Perturbations: The Space of Harmless Perturbations

  • Lu Chen
  • Shaofeng Li
  • Benhao Huang
  • Fan Yang
  • Zheng Li
  • Jie Li
  • Yuan Luo

Existing works have extensively studied adversarial examples, which are minimal perturbations that can mislead the output of deep neural networks (DNNs) while remaining imperceptible to humans. However, in this work, we reveal the existence of a harmless perturbation space, in which perturbations drawn from this space, regardless of their magnitudes, leave the network output unchanged when applied to inputs. Essentially, the harmless perturbation space emerges from the usage of non-injective functions (linear or non-linear layers) within DNNs, enabling multiple distinct inputs to be mapped to the same output. For linear layers with input dimensions exceeding output dimensions, any linear combination of the orthogonal bases of the nullspace of the parameter consistently yields no change in their output. For non-linear layers, the harmless perturbation space may expand, depending on the properties of the layers and input samples. Inspired by this property of DNNs, we solve for a family of general perturbation spaces that are redundant for the DNN's decision, and can be used to hide sensitive data and serve as a means of model identification. Our work highlights the distinctive robustness of DNNs (i.e., consistency under large magnitude perturbations) in contrast to adversarial examples (vulnerability for small noises).

EAAI Journal 2025 Journal Article

Diameter-adjustable mandrel for thin-wall tube bending and its domain knowledge-integrated optimization design framework

  • Zili Wang
  • Jie Li
  • Xiaojian Liu
  • Shuyou Zhang
  • Yaochen Lin
  • Jianrong Tan

In response to the growing demand for small-batch bending tube production, traditional bending dies require separate customization for each tube size, resulting in extended design cycles and high costs. To meet bending requirements for tubes of different diameters using a single mandrel, a novel adjustable diameter mechanism (DAM) and its optimization design method are proposed. Initially, the DAM based on a planetary bevel gear-screw transmission set is developed for bending tubes of varying diameters. Subsequently, a domain knowledge-integrated optimization design framework is introduced. To reduce the cost of acquiring training samples for training surrogate models, a monotonicity-constrained neural network based on cascade boosting architecture (CB-MCNN) is introduced that enhances prediction accuracy while maintaining monotonicity. To improve the optimization speed and quality of Evolutionary Algorithms (EAs), a domain knowledge-guided EA (DK-EA) method is proposed, incorporating domain knowledge into the population initialization phase. The results indicate that: (1) CB-MCNN outperforms traditional methods and shows excellent performance on small-sample datasets. (2) DK-EA accelerates optimization processes and produces better outcomes. As a result, the domain knowledge-integrated optimization design framework enables the DAM to achieve a wider diameter variation range and enhanced reliability. The optimized DAM demonstrates the capability to bend tubes with diameters of 46–60 mm.

EAAI Journal 2025 Journal Article

Elements discriminative non-negative matrix factorization for data clustering

  • Jie Li
  • Xuzhu Shen
  • Chaoqian Li
  • Yaotang Li

Semi-supervised non-negative matrix factorization (NMF) is widely used in data clustering because it can generate more discriminative representations for high-dimensional data by leveraging partial label information. To advance research in this field, we propose a novel method, Element Discriminative NMF (EDNMF), which incorporates discrimination constraints based on the element ratio and element difference of the new representations of labeled data points. EDNMF is implemented in two variants, each with an iterative algorithm for solving the optimization problem. We further analyze the computational complexity and convergence of these algorithms. A key advantage of EDNMF is that its learned representations can serve directly as a clustering assignment matrix, thereby simplifying the clustering process. Extensive experiments on eight real-world datasets demonstrate that EDNMF consistently outperforms baseline methods, confirming its effectiveness in improving clustering performance. The code is available at https: //github. com/ljisxz/EDNMF.

EAAI Journal 2025 Journal Article

Ensemble fractional fuzzy dispersion entropy: A low bias approach for data analysis

  • Chengjiang Zhou
  • Longkun He
  • Xuanyu Liao
  • Jie Li
  • Hao Song
  • Yang Yang

Existing entropy methods exhibit strengths and limitations in capturing the irregularity and uncertainty of nonlinear systems. Dispersion entropy (DE), fuzzy dispersion entropy (FuDE), and ensemble dispersion entropy (EnsDE) are particularly vulnerable to parameters, noise, and outliers due to deficiencies in their mapping strategies and fuzzy classification, the inherent weaknesses of the Shannon entropy framework and suboptimal implementations lead to reduced accuracy in entropy estimation. Thus, we propose the ensemble fractional fuzzy dispersion entropy (EnsFrFuDE) and its two novel variants, E n s F r F u D E n p d f and E n s F r F u D E m e a n, designed to deliver more robust, accurate, and low-bias dispersion patterns for entropy estimation. The establishment of a generalized Gaussian distribution mapping, coupled with the meticulous refinement of five distinct mapping methodologies and the integration of three fuzzy functions, synergistically mitigates the sensitivity of entropy methods to noise and outliers. The fractional-order entropy framework, combined with n p d f process and overall m e a n, captures more nuanced hierarchical structures and multiscale features. Analysis of three groups of synthetic signals and real mechanical vibration datasets reveals that proposed methods exhibit a lower coefficient of variation for synthetic signals and greater variability across different fault patterns. Notably, our approach is effective not only for short signals and small sample sizes but also significantly outperforms 16 existing state-of-the-art entropy methods in noise robustness and estimation accuracy.

YNIMG Journal 2025 Journal Article

Express your feelings in words when you get sick of others’ pain: Effects of word labeling on pain empathy in phase-amplitude coupling

  • Jiahe Sun
  • Wenguang He
  • Hongfei Ji
  • Jie Li
  • Jie Zhuang

Revealing the neurophysiological mechanisms underlying pain empathy has profound implications for the understanding of social emotions and pain regulation. This study aimed to elucidate the mechanisms by which pain empathy modulates individual experiences, focusing on two primary questions: When individuals empathize with pain, do they attend more to the emotional content associated with pain or the nociceptive experience itself? How does lexical labeling modulate negative affect during pain empathy? In this EEG study, we employed a labeling paradigm divided into pain labeling, affect labeling and tool labeling to distinguish emotional content from the nociceptive experience of pain stimuli. We collected data from 39 participants and analyzed their EEG components and frequency-specific brain activity across the experimental conditions. Cross-frequency coupling analyses were conducted to uncover the mechanisms by which pain empathy modulates emotional responses. Our findings revealed that compared to pain labeling and tool labeling, affect labeling more effectively reduces the negative affect associated with pain empathy. This was evidenced by the decreased amplitude of the P300 component and lower theta-band activity within the prefrontal cortex, predominantly during the later stages of labeling. Additionally, affect labeling was associated with enhanced theta-gamma phase-amplitude coupling, with theta phase modulation in the posterior parietal and sensorimotor cortices influencing prefrontal gamma-band activity. These results suggest that, during pain empathy, individuals allocate greater attentional and cognitive resources to the emotional aspects of pain. Thus, implicit regulation of pain empathy involves coordinated interactions across multiple brain regions.

JBHI Journal 2025 Journal Article

From Micro to Meso: A Data-Driven Mesoscopic Region Division Method Based on Functional Connectivity for EEG-Based Driver Fatigue Detection

  • Lexing Zhong
  • Mingcheng Xu
  • Jie Li
  • Zhongfei Bai
  • Hongfei Ji
  • Lingyu Liu
  • Lingjing Jin

The integration of EEG signals and deep learning methods is emerging as an effective approach for brain fatigue detection, particularly utilizing Graph Neural Networks(GNNs) that excel in capturing complex electrode relationships. A significant challenge within GNNs is the construction of an effective adjacency matrix that enhances spatial information learning. Concurrently, electrode aggregation in EEG has emerged as a pivotal area of research. However, conventional partitioning methods depend on task-specific prior knowledge, limiting their generalizability across diverse tasks. To Address this issue, we propose a novel mesoscopic region division approach for EEG-based driver fatigue detection, leveraging inherent data characteristics and functional connectivity-based GNN. This method adopts a two-stage approach: initially, micro-electrodes exhibiting similar functional connectivity relationships are grouped as “mesoscopic region”; subsequently, all micro-electrodes in the same group are aggregated into virtual meso-electrodes, and the fatigue state classification is subsequently based on the functional connectivity between them. Applied to a public driver fatigue detection dataset, our approach surpasses existing state-of-the-art methods in performance. Additionally, interpretive analysis provides micro and mesoscopic insights into brain regions and neuronal connections associated with alert and fatigued states.

NeurIPS Conference 2025 Conference Paper

GD$^2$: Robust Graph Learning under Label Noise via Dual-View Prediction Discrepancy

  • Kailai Li
  • Jiong Lou
  • Jiawei Sun
  • Honghong Zeng
  • Wen Li
  • Chentao Wu
  • Yuan Luo
  • Wei Zhao

Graph Neural Networks (GNNs) achieve strong performance in node classification tasks but exhibit substantial performance degradation under label noise. Despite recent advances in noise-robust learning, a principled approach that exploits the node-neighbor interdependencies inherent in graph data for label noise detection remains underexplored. To address this gap, we propose GD$^2$, a noise-aware \underline{G}raph learning framework that detects label noise by leveraging \underline{D}ual-view prediction \underline{D}iscrepancies. The framework contrasts the \textit{ego-view}, constructed from node-specific features, with the \textit{structure-view}, derived through the aggregation of neighboring representations. The resulting discrepancy captures disruptions in semantic coherence between individual node representations and the structural context, enabling effective identification of mislabeled nodes. Building upon this insight, we further introduce a view-specific training strategy that enhances noise detection by amplifying prediction divergence through differentiated view-specific supervision. Extensive experiments on multiple datasets and noise settings demonstrate that \name~achieves superior performance over state-of-the-art baselines.

EAAI Journal 2025 Journal Article

Graph regularized One-hot-constrained nonnegative matrix factorization for data representation

  • Jie Li
  • Chaoqian Li

Non-negative matrix factorization (NMF) effectively reduces the dimensionality of high-dimensional data. The low-dimensional representations obtained by NMF are widely used in computer vision, information retrieval, and pattern recognition. Semi-supervised NMF can obtain more discriminative low-dimensional representation by incorporating partial label information. In this paper, we propose a semi-supervised NMF method called Graph regularized One-hot Constrained NMF(GOCNMF), by setting the low-dimensional representations of the labeled data points as One-hot vectors in the decomposition. The setting guarantees the clustering assignments of labeled data points are consistent with their ground truth classes. In clustering tasks, the low-dimensional representations of the unlabeled data points are guided by the graph regularization to approximate the One-hot vectors in new space, so that the low-dimensional representations generated by GOCNMF can directly be used as a clustering assignment matrix. In addition, the setting can also be considered as an initialization strategy for the optimization algorithm of GOCNMF, which can accelerate the convergence speed of the algorithm and reduce its computational complexity. The clustering experiments on six real datasets show that our proposed GOCNMF outperforms the comparison methods overall, which validates the effectiveness of our method. The code for reproducing our results can be obtained at: https: //github. com/ljisxz/GOCNMF.

NeurIPS Conference 2025 Conference Paper

GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification

  • Qiao Li
  • Jie Li
  • Yukang Zhang
  • Lei Tan
  • Jing Chen
  • Jiayi Ji

Aerial-Ground person re-identification (AG-ReID) is an emerging yet challenging task that aims to match pedestrian images captured from drastically different viewpoints, typically from unmanned aerial vehicles (UAVs) and ground-based surveillance cameras. The task poses significant challenges due to extreme viewpoint discrepancies, occlusions, and domain gaps between aerial and ground imagery. While prior works have made progress by learning cross-view representations, they remain limited in handling severe pose variations and spatial misalignment. To address these issues, we propose a Geometric and Semantic Alignment Network (GSAlign) tailored for AG-ReID. GSAlign introduces two key components to jointly tackle geometric distortion and semantic misalignment in aerial-ground matching: a Learnable Thin Plate Spline (LTPS) Transformation Module and a Dynamic Alignment Module (DAM). The LTPS module adaptively warps pedestrian features based on a set of learned keypoints, effectively compensating for geometric variations caused by extreme viewpoint changes. In parallel, the DAM estimates visibility-aware representation masks that highlight visible body regions at the semantic level, thereby alleviating the negative impact of occlusions and partial observations in cross-view correspondence. Extensive experiments on the challenging CARGO benchmark demonstrate the effectiveness of GSAlign, achieving significant improvements of +18. 8\% in mAP and +16. 8\% in Rank-1 accuracy over previous state-of-the-art methods.

EAAI Journal 2025 Journal Article

Improving model calibration in bone marrow cell classification through mixup and center loss fusion

  • Shuming Cheng
  • Qinghang Lu
  • Qianhang Guo
  • Yunqi Lin
  • Mingxin Li
  • Xingyu Zhao
  • Liang Guo
  • Jiaming Li

The analysis of bone marrow cell morphology is essential for the accurate diagnosis of hematological disorders. Traditional manual classification methods are time-consuming and labor-intensive. Although current automatic deep learning techniques mitigate these issues, they may still present significant risks in critical medical diagnostics due to overconfidence in predictions. To tackle these challenges, this paper proposes a novel calibration method called MixCL (Mix-Center Loss). MixCL combines the simple and effective data augmentation method Mixup with deep metric learning Center Loss, achieved through the design of a new loss function. By utilizing Mixup to generate mixing centers that enrich the feature sampling in the feature space, and leveraging the clustering effect of Center Loss to enhance the grouping of similar samples, MixCL combines the strengths of both methods. The effectiveness of MixCL is validated using three real bone marrow cell image datasets, demonstrating significant reductions in Expected Calibration Error (ECE) and Overconfidence Error (OE) for in-distribution samples. For example, in Shifted Windows Transformer model, ECE and OE metrics decreased across all datasets, with reductions averaging 1. 72% in ECE and 2. 10% in OE. The confidence Kernel Density Estimation (KDE) plot reveals that models using MixCL more effectively manage uncertainty in out-of-distribution samples, ensuring better differentiation between in-distribution and out-of-distribution samples. Thus, the proposed method effectively improves the calibration performance of the model while exhibiting better generalization performance, significantly improved when compared with current advanced bone marrow cell classification methods. Moreover, it has potential applications in various image classification fields, providing reliable confidence estimates.

NeurIPS Conference 2025 Conference Paper

JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language Models

  • Jiaxin Song
  • Yixu Wang
  • Jie Li
  • Xuan Tong
  • rui yu
  • Yan Teng
  • Xingjun Ma
  • Yingchun Wang

Vision-Language Models (VLMs) exhibit impressive performance, yet the integration of powerful vision encoders has significantly broadened their attack surface, rendering them increasingly susceptible to jailbreak attacks. However, lacking well-defined attack objectives, existing jailbreak methods often struggle with gradient-based strategies prone to local optima and lacking precise directional guidance, and typically decouple visual and textual modalities, thereby limiting their effectiveness by neglecting crucial cross-modal interactions. Inspired by the Eliciting Latent Knowledge (ELK) framework, we posit that VLMs encode safety-relevant information within their internal fusion-layer representations, revealing an implicit safety decision boundary in the latent space. This motivates exploiting boundary to steer model behavior. Accordingly, we propose \textbf{JailBound}, a novel latent space jailbreak framework comprising two stages: (1) \textbf{Safety Boundary Probing}, which addresses the guidance issue by approximating decision boundary within fusion layer's latent space, thereby identifying optimal perturbation directions towards the target region; and (2) \textbf{Safety Boundary Crossing}, which overcomes the limitations of decoupled approaches by jointly optimizing adversarial perturbations across both image and text inputs. This latter stage employs an innovative mechanism to steer the model's internal state towards policy-violating outputs while maintaining cross-modal semantic consistency. Extensive experiments on six diverse VLMs demonstrate JailBound's efficacy, achieves 94. 32\% white-box and 67. 28\% black-box attack success averagely, which are 6. 17\% and 21. 13\% higher than SOTA methods, respectively. Our findings expose a overlooked safety risk in VLMs and highlight the urgent need for more robust defenses. \textcolor{red}{Warning: This paper contains potentially sensitive, harmful and offensive content. }

EAAI Journal 2025 Journal Article

Knowledge alignment based on category-attributes for blood cell classification

  • Qianhang Guo
  • Xingyu Zhao
  • Yaodong Liang
  • Xiexin Liao
  • Jie Li
  • Jianping Guo
  • Qingmao Zhang
  • Kunyuan Xu

Blood cell image classification is crucial for diagnosing blood diseases. In this field, deep learning models have been extensively applied due to their high accuracy and efficiency. While the interpretability within models inherently possess expanded deployment potential, numerous interpretable approaches remain obscure and difficult for human comprehension. More detrimentally, interpretable models requiring domain-knowledge grounding face critical deployment barriers, as their reliance on costly expert-annotated datasets fundamentally restricts scalability across application domains. To address this issue, we introduce a novel knowledge alignment interpretability mechanism and the implementation, Knowledge-Aligned Attribute Network (KAAN), which leverages category-attributes alignment and cytomorphology. The concrete implementation, KAAN approach allows models of arbitrary architectures to learn cytomorphological knowledge through category-attribute labels, identify the true single-value attributes of images from multi-value category-attribute labels, and leverage these attributes for effective classification. Users can leverage reality-aligned attribute outputs to assess the credibility of a model’s classification results, while freely manipulating these attributes to fulfill practical engineering requirements such as false sample correction or manual hematological cell morphology training program. Meanwhile, knowledge alignment implementation could preserve the model’s original performance. More critically, the category-attribute labels are directly derived from existing knowledge texts, significantly reducing annotation costs during cross-domain deployment of KAAN approach. Comprehensive quantitative and qualitative experiments demonstrate that KAAN outperforms current state-of-the-art counterparts in both attribute recognition and classification accuracy, while exhibiting superior knowledge alignment characteristics. This novel interpretability mechanism not only demonstrates strong practical value but also fills the gap in interpretable models within the field of blood cell classification.

IJCAI Conference 2025 Conference Paper

Leveraging Peer-Informed Label Consistency for Robust Graph Neural Networks with Noisy Labels

  • Kailai Li
  • Jiawei Sun
  • Jiong Lou
  • Zhanbo Feng
  • Hefeng Zhou
  • Chentao Wu
  • Guangtao Xue
  • Wei Zhao

Graph Neural Networks (GNNs) excel in many applications but struggle when trained with noisy labels, especially as noise can propagate through the graph structure. Despite recent progress in developing robust GNNs, few methods exploit the intrinsic properties of graph data to filter out noise. In this paper, we introduce ProCon, a novel framework that identifies mislabeled nodes by measuring label consistency among semantically similar peers, which are determined by feature similarity and graph adjacency. Mislabeled nodes typically exhibit lower consistency with these peers, a signal we measure using pseudo-labels derived from representational prototypes. A Gaussian Mixture Model is fitted to the consistency distribution to identify clean samples, which refine prototype quality in an iterative feedback loop. Experiments on multiple datasets demonstrate that ProCon significantly outperforms state-of-the-art methods, effectively mitigating label noise and enhancing GNN robustness.

JBHI Journal 2025 Journal Article

M-NET: Transforming Single Nucleotide Variations Into Patient Feature Images for the Prediction of Prostate Cancer Metastasis and Identification of Significant Pathways

  • Li Zhou
  • Jie Li
  • Weilong Tan

High-performance prediction of prostate cancer metastasis based on single nucleotide variations remains a challenge. Therefore, we developed a novel biologically informed deep learning framework, named M-NET, for the prediction of prostate cancer metastasis. Within the framework, we transformed single nucleotide variations into patient feature images that are optimal for fitting convolutional neural networks. Moreover, we identified significant pathways associated with the metastatic status. The experimental results showed that M-NET significantly outperformed other comparison methods based on single nucleotide variations, achieving improvements in accuracy, precision, recall, F1-score, area under the receiver operating characteristics curve, and area under the precision-recall curve by 6. 3%, 8. 4%, 5. 1%, 0. 070, 0. 041, and 0. 026, respectively. Furthermore, M-NET identified some important pathways associated with the metastatic status, such as signaling by the hedgehog pathway. In summary, compared with other comparative methods, M-NET exhibited a better performance in the prediction of prostate cancer metastasis.

NeurIPS Conference 2025 Conference Paper

MDReID: Modality-Decoupled Learning for Any-to-Any Multi-Modal Object Re-Identification

  • Yingying Feng
  • Jie Li
  • Jie Hu
  • Yukang Zhang
  • Lei Tan
  • Jiayi Ji

The challenge of inconsistent modalities in real-world applications presents significant obstacles to effective object re-identification (ReID). However, most existing approaches assume modality-matched conditions, significantly limiting their effectiveness in modality-mismatched scenarios. To overcome this limitation and achieve a more flexible ReID, we introduce MDReID to allow any-to-any image-level ReID systems. MDReID is inspired by the widely recognized perspective that modality information comprises both modality-shared features, predictable across modalities, and unpredictable modality-specific features, which are inherently modality-dependent and consist of two key components: the Modality Decoupling Module (MDM) and Modality-aware Metric Learning (MML). Specifically, MDM explicitly decomposes modality features into modality-shared and modality-specific representations, enabling effective retrieval in both modality-aligned and mismatched scenarios. MML, a tailored metric learning strategy, further enhances feature discrimination and decoupling by exploiting distributional relationships between shared and specific modality features. Extensive experiments conducted on three challenging multi-modality ReID benchmarks (RGBNT201, RGBNT100, MSVR310) consistently demonstrate the superiority of MDL. MDReID achieves significant mAP improvements of 9. 8\%, 3. 0\%, and 11. 5\% in modality-matched scenarios, and average gains of 3. 4\%, 11. 8\%, and 10. 9\% in modality-mismatched scenarios, respectively.

IJCAI Conference 2025 Conference Paper

Multi-modal Anchor Gated Transformer with Knowledge Distillation for Emotion Recognition in Conversation

  • Jie Li
  • Shifei Ding
  • Lili Guo
  • Xuan Li

Emotion Recognition in Conversation (ERC) aims to detect the emotions of individual utterances within a conversation. Generating efficient and modality-specific representations for each utterance remains a significant challenge. Previous studies have proposed various models to integrate features extracted using different modality-specific encoders. However, they neglect the varying contributions of modalities to this task and introduce high complexity by aligning modalities at the frame level. To address these challenges, we propose the Multi-modal Anchor Gated Transformer with Knowledge Distillation (MAGTKD) for the ERC task. Specifically, prompt learning is employed to enhance textual modality representations, while knowledge distillation is utilized to strengthen representations of weaker modalities. Furthermore, we introduce a multi-modal anchor gated transformer to effectively integrate utterance-level representations across modalities. Extensive experiments on the IEMOCAP and MELD datasets demonstrate the effectiveness of knowledge distillation in enhancing modality representations and achieve state-of-the-art performance in emotion recognition. Our code is available at: https: //github. com/JieLi-dd/MAGTKD.

ICML Conference 2025 Conference Paper

Multi-Modal Object Re-identification via Sparse Mixture-of-Experts

  • Yingying Feng
  • Jie Li
  • Chi Xie
  • Lei Tan
  • Jiayi Ji

We present MFRNet, a novel network for multi-modal object re-identification that integrates multi-modal data features to effectively retrieve specific objects across different modalities. Current methods suffer from two principal limitations: (1) insufficient interaction between pixel-level semantic features across modalities, and (2) difficulty in balancing modality-shared and modality-specific features within a unified architecture. To address these challenges, our network introduces two core components. First, the Feature Fusion Module (FFM) enables fine-grained pixel-level feature generation and flexible cross-modal interaction. Second, the Feature Representation Module (FRM) efficiently extracts and combines modality-specific and modality-shared features, achieving strong discriminative ability with minimal parameter overhead. Extensive experiments on three challenging public datasets (RGBNT201, RGBNT100, and MSVR310) demonstrate the superiority of our approach in terms of both accuracy and efficiency, with 8. 4% mAP and 6. 9% accuracy improved in RGBNT201 with negligible additional parameters.

JBHI Journal 2025 Journal Article

Multi-Modality Regional Alignment Network for Covid X-Ray Survival Prediction and Report Generation

  • Zhusi Zhong
  • Jie Li
  • John Sollee
  • Scott Collins
  • Harrison Bai
  • Paul Zhang
  • Terrance Healey
  • Michael Atalay

In response to the worldwide COVID-19 pandemic, advanced automated technologies have emerged as valuable tools to aid healthcare professionals in managing an increased workload by improving radiology report generation and prognostic analysis. This study proposes a Multi-modality Regional Alignment Network (MRANet), an explainable model for radiology report generation and survival prediction that focuses on high-risk regions. By learning spatial correlation in the detector, MRANet visually grounds region-specific descriptions, providing robust anatomical regions with a completion strategy. The visual features of each region are embedded using a novel survival attention mechanism, offering spatially and risk-aware features for sentence encoding while maintaining global coherence across tasks. A cross-domain LLMs-Alignment is employed to enhance the image-to-text transfer process, resulting in sentences rich with clinical detail and improved explainability for radiologists. Multi-center experiments validate the overall performance and each module's composition within the model, encouraging further advancements in radiology report generation research emphasizing clinical interpretation and trustworthiness in AI models applied to medical studies.

IROS Conference 2025 Conference Paper

Multi-UAV Deployment in Obstacle-Cluttered Environments with LOS Connectivity

  • Yuda Chen
  • Shuaikang Wang
  • Jie Li
  • Meng Guo

A reliable communication network is essential for multiple UAVs operating within obstacle-cluttered environments, where limited communication due to obstructions often occurs. A common solution is to deploy intermediate UAVs to relay information via a multi-hop network, which introduces two challenges: (i) how to design the structure of multi-hop networks; and (ii) how to maintain connectivity during collaborative motion. To this end, this work first proposes an efficient constrained search method based on the minimum-edge RRT ⋆ algorithm, to find a spanning-tree topology that requires a less number of UAVs for the deployment task. Then, to achieve this deployment, a distributed model predictive control strategy is proposed for the online motion coordination. It explicitly incorporates not only the inter-UAV and UAV-obstacle distance constraints, but also the line-of-sight (LOS) connectivity constraint. These constraints are well-known to be nonlinear and often tackled by various approximations. In contrast, this work provides a theoretical guarantee that all agent trajectories are ensured to be collision-free with a team-wise LOS connectivity at all time. Numerous simulations are performed in 3D valley-like environments, while hardware experiments validate its dynamic adaptation when the deployment position changes online.

JBHI Journal 2025 Journal Article

RTGN: Robust Traditional Chinese Medicine Graph Networks for Patient Similarity Learning

  • Junjie Long
  • Jinghao Niu
  • Heping Wang
  • Jiaxi Liu
  • Jie Li
  • Wensheng Zhang

Traditional Chinese Medicine (TCM) boasts a long history and a unique diagnostic and therapeutic paradigm. Integrating TCM with Western medicine and modern medical devices has yielded numerous successful cases in recent years. TCM treatment has developed a special knowledge framework focusing on precise differentiation based on multidimensional information such as the patient's diseases, symptoms, and syndromes. This offers significant opportunities for AI research in similar patient scenarios within TCM contexts. However, traditional medicine's reliance on the physiological sensory judgment of human physicians to gather clinical information might lead to non-standardized descriptions and disturbances in patient assessments. Additionally, how to integrate TCM's fine-grained differentiation knowledge to design a patient similarity measure remains an open question. To address this, we first constructed a real-world dataset of TCM gastrointestinal malignancies (TCMGI) based on real cases in the Guang'anmen Hospital, China Academy of Chinese Medical Sciences. It contains 406 types of multidimensional information from 719 patients, organized in a graph structure. Second, we develop a novel deep learning framework, Robust Traditional Chinese Medicine Graph Networks (RTGN), which employs a Siamese network architecture with self-attention and self-supervision strategies to enhance robustness in patient retrieval. Lastly, we design a patient similarity metric integrating TCM and Western medicine approaches, demonstrating superior performance in depicting fine-grained patient similarities. Experimental results show our method outperforms existing best practices in patient retrieval accuracy. Moreover, the proposed similarity metric exhibits excellent performance in clustering tasks at various granularity levels, possibly supporting precision TCM patient retrieval and downstream tasks, such as prescription generation.

JBHI Journal 2025 Journal Article

Sequential sEMG Recognition With Knowledge Transfer and Dynamic Graph Network Based on Spatio-Temporal Feature Extraction Network

  • Zhilin Li
  • Xianghe Chen
  • Jie Li
  • Zhongfei Bai
  • Hongfei Ji
  • Lingyu Liu
  • Lingjing Jin

Surface electromyography (sEMG) signals are electrical signals released by muscles during movement, which can directly reflect the muscle conditions during various actions. When a series of continuous static actions are connected along the temporal axis, a sequential action is formed, which is more aligned with people's intuitive understanding of real-life movements. The signals acquired during sequential actions are known as sequential sEMG signals, including an additional dimension of sequence, embodying richer features compared to static sEMG signals. However, existing methods show inadequate utilization of the signals' sequential characteristics. Addressing these gaps, this paper introduces the Spatio-Temporal Feature Extraction Network (STFEN), which includes a Sequential Feature Analysis Module based on static-sequential knowledge transfer, and a Spatial Feature Analysis Module based on dynamic graph networks to analyze the internal relationships between the leads. The effectiveness of STFEN is tested on both modified publicly available datasets and on our acquired Arabic Digit Sequential Electromyography (ADSE) dataset. The results show that STFEN outperforms existing models in recognizing sequential sEMG signals. Experiments have confirmed the reliability and wide applicability of STFEN in analyzing complex muscle activities. Furthermore, this work also suggests STFEN's potential benefits in rehabilitation medicine, particularly for stroke recovery, and shows promising future applications.

NeurIPS Conference 2025 Conference Paper

Sparse Meets Dense: Unified Generative Recommendations with Cascaded Sparse-Dense Representations

  • Yuhao Yang
  • ZhI JI
  • Zhaopeng Li
  • Yi Li
  • Zhonglin Mo
  • Yue Ding
  • Kai Chen
  • Zijian Zhang

Generative models have recently gained attention in recommendation systems by directly predicting item identifiers from user interaction sequences. However, existing methods suffer from significant information loss due to the separation of stages such as quantization and sequence modeling, hindering their ability to achieve the modeling precision and accuracy of sequential dense retrieval techniques. Integrating generative and dense retrieval methods remains a critical challenge. To address this, we introduce the Cascaded Organized Bi-Represented generAtive retrieval (COBRA) framework, which innovatively integrates sparse semantic IDs and dense vectors through a cascading process. Our method alternates between generating these representations by first generating sparse IDs, which serve as conditions to aid in the generation of dense vectors. End-to-end training enables dynamic refinement of dense representations, capturing both semantic insights and collaborative signals from user-item interactions. During inference, COBRA employs a coarse-to-fine strategy, starting with sparse ID generation and refining them into dense vectors via the generative model. We further propose BeamFusion, an innovative approach combining beam search with nearest neighbor scores to enhance inference flexibility and recommendation diversity. Extensive experiments on public datasets and offline tests validate our method's robustness. Online A/B tests on a real-world advertising platform with over 200 million daily users demonstrate substantial improvements in key metrics, highlighting COBRA's practical advantages.

EAAI Journal 2025 Journal Article

WFC-BSN: Wavelet fusion-based conditional blind-spot network for self-supervised forward sonar denoising

  • Ziqi Xia
  • Jie Li
  • Wenpei Jiao
  • Chunyan Zhang
  • Jianlei Zhang
  • Guangming Xie

Forward-looking sonar is a critical technology for underwater information acquisition, but complex environment and equipment limitation introduce significant noise, degrading image quality and hindering AI-based interpretation. Current self-supervised denoising methods, primarily designed for optical images, struggle to address inherent structural noise and restore local details of imaging targets. Additionally, these methods lack ground truth for evaluation, complicating the application of reference-based metrics in sonar denoising tasks. To address these issues, this paper proposes a wavelet fusion-based conditional blind-spot network (WFC-BSN) for sonar image denoising without needing ground truth data. The core of this network is the wavelet transform fusion block (WTFB), which decomposes feature maps into multiple frequency-domain components, merging summed and differential features. This enhances the receptive field and boosts the network’s ability to recover fine details. Moreover, WFC-BSN introduces Blind-Self Displacement Loss, comprising Blind-Spot Displacement Loss and Self-Similar Displacement Loss. The former enhances noise robustness through data perturbations, and the latter improves structural reconstruction by retaining original pixel information. We also present a novel method for generating sonar ground truth by extracting noise-free feature from continuous frames. This method utilizes forward-looking sonar data collected from real ocean environment to construct a clean-noise paired dataset, enabling reference-based denoising performance evaluation. Experimental results demonstrate that whether for high- or low-frequency sonar images, WFC-BSN surpasses both traditional non-learning and existing self-supervised denoising techniques in effectiveness and detail preservation. Additionally, denoised images show improved accuracy in detection tasks, underscoring the method’s efficacy and practical value in underwater applications. Code and Dataset can be found at https: //github. com/xzq-2000/WFC-BSN.

EAAI Journal 2024 Journal Article

A novel bone marrow cell recognition method based on multi-scale information and reject option

  • Haisen He
  • Zilan Li
  • Yunqi Lin
  • Tongyi Wei
  • Qianghang Guo
  • Qinghang Lu
  • Liang Guo
  • Qingmao Zhang

The morphology of bone marrow cells is crucial for the diagnostics of blood disorders. However, the traditional approach of cell sorting and counting under a microscope is time-consuming and error-prone. To assist experts in diagnosis, there is considerable interest at present in using neural networks to develop automatic recognition algorithms. However, the results obtained from these approaches indicate that the achieved accuracy is not entirely convincing, and there remains a risk of misdiagnosis. In fact, a reliable collaborative recognition algorithm not only needs to have high recognition accuracy but also must be able to confidently reject low-confidence classification results and leave them for human experts to handle. Therefore, this paper proposes a method that combines high accuracy with a rejection recognition option for bone marrow cell recognition. This method incorporates Crossformer (Cross) and Class Centroid Learning (CCL), which integrate multi-scale information from cell images, enhance image feature differentiation, and effectively reject low-confidence predictions. In the experiment, the recognition accuracy of Cross-CCL reaches 94. 41%, surpassing the performance of the current mainstream and advanced methods for bone marrow cell recognition. Simultaneously, CCL serves as a dependable collaborative algorithm by effectively rejecting 21. 07% of cell images that represent low-confidence prediction results. Moreover, the accuracy of identification for the remaining portions exceeds 98%. This rejection mechanism not only enhances the efficiency of expert resource allocation but also ensures a higher level of confidence in the identified cell images.

ICRA Conference 2024 Conference Paper

Augmenting Lane Perception and Topology Understanding with Standard Definition Navigation Maps

  • Katie Luo
  • Xinshuo Weng
  • Yan Wang 0051
  • Shuang Wu
  • Jie Li
  • Kilian Q. Weinberger
  • Yue Wang 0041
  • Marco Pavone 0001

Autonomous driving has traditionally relied heavily on costly and labor-intensive High Definition (HD) maps, hindering scalability. In contrast, Standard Definition (SD) maps are more affordable and have worldwide coverage, offering a scalable alternative. In this work, we systematically explore the effect of SD maps for real-time lane-topology understanding. We propose a novel framework to integrate SD maps into online map prediction and propose a Transformer-based encoder, SD Map Encoder Representations from transFormers, to leverage priors in SD maps for the lane-topology prediction task. This enhancement consistently and significantly boosts (by up to 60%) lane detection and topology prediction on current state-of-the-art online map prediction methods without bells and whistles and can be immediately incorporated into any Transformer-based lane-topology method. Code is available at https://github.com/NVlabs/SMERF.

JBHI Journal 2024 Journal Article

De-Biased Disentanglement Learning for Pulmonary Embolism Survival Prediction on Multimodal Data

  • Zhusi Zhong
  • Jie Li
  • Shreyas Kulkarni
  • Helen Zhang
  • Fayez H. Fayad
  • Yang Li
  • Scott Collins
  • Harrison Bai

Health disparities among marginalized populations with lower socioeconomic status significantly impact the fairness and effectiveness of healthcare delivery. The increasing integration of artificial intelligence (AI) into healthcare presents an opportunity to address these inequalities, provided that AI models are free from bias. This paper aims to address the bias challenges by population disparities within healthcare systems, existing in the presentation of and development of algorithms, leading to inequitable medical implementation for conditions such as pulmonary embolism (PE) prognosis. In this study, we explore the diverse bias in healthcare systems, which highlights the demand for a holistic framework to reducing bias by complementary aggregation. By leveraging de-biasing deep survival prediction models, we propose a framework that disentangles identifiable information from images, text reports, and clinical variables to mitigate potential biases within multimodal datasets. Our study offers several advantages over traditional clinical-based survival prediction methods, including richer survival-related characteristics and bias-complementary predicted results. By improving the robustness of survival analysis through this framework, we aim to benefit patients, clinicians, and researchers by enhancing fairness and accuracy in healthcare AI systems.

IROS Conference 2024 Conference Paper

Design and implementation of a novel wheel-based cable inspection robot

  • Mengqi Hou
  • Jie Li
  • Fengyu Xu
  • Lezhi Hu

Regular maintenance and inspection of cables are essential for cable-stayed bridges and suspension bridges, as cables serve as the core components. In order to enable automated detection of cables, this paper proposes a novel wheeled cable inspection robot. The robot utilizes a bilateral wheel structure and is composed of four independent suspension mechanisms. By collaborating with a lifting mechanism, the robot achieves functions such as adhesion, climbing, and obstacle overcoming. The robot is powered by a lithium polymer battery and operated via wireless control by ground personnel. This paper provides a detailed exposition on the structure design and control system of robots, and conducts a mechanical analysis of the suspension mechanism of robots. The maximum obstacle-negotiation height of the robot is calculated, and a mechanical model for cable climbing is established. During prototype testing, the robot demonstrated a mass of 6. 7kg, a maximum payload capacity of 6kg, a maximum obstacle height of 10mm, and a fastest climbing speed of 14m/min. These specifications meet the requirements of practical inspections.

IJCAI Conference 2024 Conference Paper

Multi-Granularity Graph-Convolution-Based Method for Weakly Supervised Person Search

  • Haichun Tai
  • De Cheng
  • Jie Li
  • Nannan Wang
  • Xinbo Gao

One-step Weakly Supervised Person Search (WSPS) jointly performs pedestrian detection and person Re-IDentification (ReID) only with bounding box annotations, which makes the traditional person ReID problem more suitable and efficient for real-world applications. However, this task is very challenging due to the following reasons: 1) large feature gap between person ReID and general object detection tasks when learning shared representations; 2) difficult pseudo identity estimation for each person image with unrefined raw detection and dramatic scale changes. To address above issues, we propose a multi-granularity graph convolution framework to jointly optimize the aligned task features, as well as to assist the pseudo label estimation. Specifically, the multi-granularity feature alignment module (MFA) in the designed two-branch framework, employs cluster-level bi-directional interaction of various granularity information to narrow down the large feature gap. Further, upon the MFA module, we introduce the multi-granularity graph-convolution-based pseudo-label estimation module, to enhance feature representations for distinguishing diverse identities. Extensive experimental results demonstrate the effectiveness of the proposed method, and show superior performances to state-of-the art methods by a large margin on CUHK-SYSU and PRW datasets.

AAAI Conference 2024 Conference Paper

Multi-Scene Generalized Trajectory Global Graph Solver with Composite Nodes for Multiple Object Tracking

  • Yan Gao
  • Haojun Xu
  • Jie Li
  • Nannan Wang
  • Xinbo Gao

The global multi-object tracking (MOT) system can consider interaction, occlusion, and other ``visual blur'' scenarios to ensure effective object tracking in long videos. Among them, graph-based tracking-by-detection paradigms achieve surprising performance. However, their fully-connected nature poses storage space requirements that challenge algorithm handling long videos. Currently, commonly used methods are still generated trajectories by building one-forward associations across frames. Such matches produced under the guidance of first-order similarity information may not be optimal from a longer-time perspective. Moreover, they often lack an end-to-end scheme for correcting mismatches. This paper proposes the Composite Node Message Passing Network (CoNo-Link), a multi-scene generalized framework for modeling ultra-long frames information for association. CoNo-Link's solution is a low-storage overhead method for building constrained connected graphs. In addition to the previous method of treating objects as nodes, the network innovatively treats object trajectories as nodes for information interaction, improving the graph neural network's feature representation capability. Specifically, we formulate the graph-building problem as a top-k selection task for some reliable objects or trajectories. Our model can learn better predictions on longer-time scales by adding composite nodes. As a result, our method outperforms the state-of-the-art in several commonly used datasets.

YNIMG Journal 2024 Journal Article

SiMix: A domain generalization method for cross-site brain MRI harmonization via site mixing

  • Chundan Xu
  • Jie Li
  • Yakui Wang
  • Lixue Wang
  • Yizhe Wang
  • Xiaofeng Zhang
  • Weiqi Liu
  • Jingang Chen

Brain magnetic resonance imaging (MRI) is widely used in clinical practice for disease diagnosis. However, MRI scans acquired at different sites can have different appearances due to the difference in the hardware, pulse sequence, and imaging parameter. It is important to reduce or eliminate such cross-site variations with brain MRI harmonization so that downstream image processing and analysis is performed consistently. Previous works on the harmonization problem require the data acquired from the sites of interest for model training. But in real-world scenarios there can be test data from a new site of interest after the model is trained, and training data from the new site is unavailable when the model is trained. In this case, previous methods cannot optimally handle the test data from the new unseen site. To address the problem, in this work we explore domain generalization for brain MRI harmonization and propose Site Mix (SiMix). We assume that images of travelling subjects are acquired at a few existing sites for model training. To allow the training data to better represent the test data from unseen sites, we first propose to mix the training images belonging to different sites stochastically, which substantially increases the diversity of the training data while preserving the authenticity of the mixed training images. Second, at test time, when a test image from an unseen site is given, we propose a multiview strategy that perturbs the test image with preserved authenticity and ensembles the harmonization results of the perturbed images for improved harmonization quality. To validate SiMix, we performed experiments on the publicly available SRPBS dataset and MUSHAC dataset that comprised brain MRI acquired at nine and two different sites, respectively. The results indicate that SiMix improves brain MRI harmonization for unseen sites, and it is also beneficial to the harmonization of existing sites.

IROS Conference 2023 Conference Paper

A Distributed Scheduling Method for Networked UAV Swarm based on Computing for Communication

  • Runfeng Chen
  • Jie Li
  • Yiting Chen
  • Yuchong Huang

UAV swarms have attracted much attention for post-disaster search and rescue, pollution monitoring and trace-ability, etc. , where distributed scheduling is required to arrange careful tasks and time quickly. The market-based methods are widely favored but they rely on the environmentally influenced communication network to complete negotiation, while the on-board computing of UAV is robust and redundant. This paper proposes a distributed scheduling method for networked UAV swarm based on computing for communication, which trades a modest increase in computing for a significant decrease in communication. First, by analyzing the task removal strategies of two representative methods, the consensus-based bundle algorithm (CBBA) and performance impact (PI) algorithm, a new removal strategy is proposed, which expands the explo-ration of the bundle and can potentially reduce communication rounds. Second, the proposed task-related optimization method can extract task conflict nodes from the native communication protocol, and use the sampling and estimation strategies to resolve task conflicts in advance. Third, historical bids are cleverly used to infer others' locations, which is necessary for task-related optimization. Fourth, to verify the algorithm in real communication, a hardware-in-the-loop (HIL) ad-hoc network simulation system is constructed, which uses real network protocols and simulated channel transmissions. Finally, the HIL Monte Carlo simulation results show that, compared with CBBA and PI, the proposed method can significantly reduce the number of communication rounds and the total scheduling time, without increasing the communication protocol overhead and loss of optimization.

JAAMAS Journal 2023 Journal Article

A performance-impact based multi-task distributed scheduling algorithm with task removal inference and deadlock avoidance

  • Jie Li
  • Runfeng Chen
  • Xiangke Wang

Abstract Multi-task distributed scheduling (MTDS) remains a challenging problem for multi-agent systems used for uncertain and dynamic real-world tasks such as search-and-rescue. The Performance Impact (PI) algorithm is an excellent solution for MTDS, but it suffers from the problem of non-convergence that it may fall into an infinite cycle of exchanging the same task. In this paper, we improve the PI algorithm through the integration of a task removal inference strategy and a deadlock avoidance mechanism. Specifically, the task removal inference strategy results in better exploration performance than the original PI, improving the suboptimal solutions caused by the heuristics for local task selection as done in PI. In addition, we design a deadlock avoidance mechanism that limits the number of times of removing the same task and isolating consecutive inclusions of the same task. Therefore, it guarantees the convergence of the MTDS algorithm. We demonstrate the advantage of the proposed algorithm over the original PI algorithm through Monte Carlo simulation of the search-and-rescue task. The results show that the proposed algorithm can obtain a lower average time cost and the highest total allocation number.

NeurIPS Conference 2023 Conference Paper

Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise

  • Arpit Bansal
  • Eitan Borgnia
  • Hong-Min Chu
  • Jie Li
  • Hamid Kazemi
  • Furong Huang
  • Micah Goldblum
  • Jonas Geiping

Standard diffusion models involve an image transform -- adding Gaussian noise -- and an image restoration operator that inverts this degradation. We observe that the generative behavior of diffusion models is not strongly dependent on the choice of image degradation, and in fact, an entire family of generative models can be constructed by varying this choice. Even when using completely deterministic degradations (e. g. , blur, masking, and more), the training and test-time update rules that underlie diffusion models can be easily generalized to create generative models. The success of these fully deterministic models calls into question the community's understanding of diffusion models, which relies on noise in either gradient Langevin dynamics or variational inference and paves the way for generalized diffusion models that invert arbitrary processes.

AAAI Conference 2023 Conference Paper

Incremental Image De-raining via Associative Memory

  • Yi Gu
  • Chao Wang
  • Jie Li

While deep learning models have achieved the state-of-the-art performance on single-image rain removal, most methods only consider learning fixed mapping rules on the single synthetic dataset for lifetime. This limits the real-life application as iterative optimization may change mapping rules and training samples. However, when models learn a sequence of datasets in multiple incremental steps, they are susceptible to catastrophic forgetting that adapts to new incremental episodes while failing to preserve previously acquired mapping rules. In this paper, we argue the importance of sample diversity in the episodes on the iterative optimization, and propose a novel memory management method, Associative Memory, to achieve incremental image de-raining. It bridges connections between current and past episodes for feature reconstruction by sampling domain mappings of past learning steps, and guides the learning to trace the current pathway back to the historical environment without storing extra data. Experiments demonstrate that our method can achieve better performance than existing approaches on both inhomogeneous and incremental datasets within the spectrum of highly compact systems.

AAAI Conference 2023 Conference Paper

MRCN: A Novel Modality Restitution and Compensation Network for Visible-Infrared Person Re-identification

  • Yukang Zhang
  • Yan Yan
  • Jie Li
  • Hanzi Wang

Visible-infrared person re-identification (VI-ReID), which aims to search identities across different spectra, is a challenging task due to large cross-modality discrepancy between visible and infrared images. The key to reduce the discrepancy is to filter out identity-irrelevant interference and effectively learn modality-invariant person representations. In this paper, we propose a novel Modality Restitution and Compensation Network (MRCN) to narrow the gap between the two modalities. Specifically, we first reduce the modality discrepancy by using two Instance Normalization (IN) layers. Next, to reduce the influence of IN layers on removing discriminative information and to reduce modality differences, we propose a Modality Restitution Module (MRM) and a Modality Compensation Module (MCM) to respectively distill modality-irrelevant and modality-relevant features from the removed information. Then, the modality-irrelevant features are used to restitute to the normalized visible and infrared features, while the modality-relevant features are used to compensate for the features of the other modality. Furthermore, to better disentangle the modality-relevant features and the modality-irrelevant features, we propose a novel Center-Quadruplet Causal (CQC) loss to encourage the network to effectively learn the modality-relevant features and the modality-irrelevant features. Extensive experiments are conducted to validate the superiority of our method on the challenging SYSU-MM01 and RegDB datasets. More remarkably, our method achieves 95.1% in terms of Rank-1 and 89.2% in terms of mAP on the RegDB dataset.

JBHI Journal 2023 Journal Article

Predicting Drug-Disease Associations Through Similarity Network Fusion and Multi-View Feature Projection Representation

  • Shiming Wang
  • Jie Li
  • Dong Wang
  • Dechen Xu
  • Jiahuan Jin
  • Yadong Wang

Predicting drug-disease associations (DDAs) through computational methods has become a prevalent trend in drug development because of their high efficiency and low cost. Existing methods usually focus on constructing heterogeneous networks by collecting multiple data resources to improve prediction ability. However, potential association possibilities of numerous unconfirmed drug-related or disease-related pairs are not sufficiently considered. In this article, we propose a novel computational model to predict new DDAs. First, a heterogeneous network is constructed, including four types of nodes (drugs, targets, cell lines, diseases) and three types of edges (associations, association scores, similarities). Second, an updating and merging-based similarity network fusion method, termed UM-SF, is presented to fuse various similarity networks with diverse weights. Finally, an intermediate layer-mediated multi-view feature projection representation method, termed IM-FP, is proposed to calculate the predicted DDA scores. This method uses multiple association scores to construct multi-view drug features, then projects them into disease space through the intermediate layer, where an intermediate layer similarity constraint is designed to learn the projection matrices. Results of comparative experiments reveal the effectiveness of our innovations. Comparisons with other state-of-the-art models by the 10-fold cross-validation experiment indicate our model's advantage on AUROC and AUPR metrics. Moreover, our proposed model successfully predicted 107 novel high-ranked DDAs.

YNIMG Journal 2023 Journal Article

RETRACTED: Transcriptional correlates of frequency-dependent brain functional activity associated with symptom severity in degenerative cervical myelopathy

  • Xing Guo
  • Jie Li
  • Qian Su
  • Jiajun Song
  • Cai Cheng
  • Xu Chu
  • Rui Zhao

BACKGROUND: Neuroimaging techniques provide insights into the brain abnormalities secondary to degenerative cervical myelopathy (DCM) and their association with neurological deficits. However, the neural correlates underlying the discrepancy between symptom severity and the degree of spinal cord compression, as well as the transcriptional correlates of these cortical abnormalities, remain unknown in DCM patients. METHODS: In this cross-sectional study, which collected resting-state functional MRI (rs-fMRI) images and the Japanese Orthopedic Association (JOA) score, enrolled 104 participants (54 patients and 50 healthy controls). The frequency-dependent amplitude of low-frequency fluctuation (ALFF) was obtained for all participants. We investigated the ALFF differences between mild-symptom DCM patients and severe-symptom DCM patients while carefully matching the degree of compression between these two groups via both univariate comparison and searchlight classification for three frequency bands (e.g., Slow-4, Slow-5, and Full-band). Additionally, we identified genes associated with symptom severity in DCM patients by linking the spatial patterns of gene expression of Allen Human Brain Atlas and brain functional differences between mild symptom and severe symptom groups. RESULTS: (1) We found that the frequency-specific brain activities within the sensorimotor network (SMN), visual network (VN), and default mode network (DMN) were associated with the varying degrees of functional impairment in DCM patients; (2) the frequency-specific brain activity within the SMN correlated with the functional recovery in patients with DCM; (3) a spatial correlation between the brain-wide expression of genes involved in neuronal migration and the brain functional activities associated with symptom severity was identified in DCM patients. CONCLUSION: In conclusion, our study bridges gaps between genes, cell classes, biological processes, and brain functional correlates of DCM. While our findings are correlational in nature, they suggest that the neural activities of sensorimotor cortices in DCM are associated with the severity of symptoms and might be associated with neuronal migration within the brain.

IROS Conference 2022 Conference Paper

Attention-Based Deep Driving Model for Autonomous Vehicles with Surround-View Cameras

  • Yang Zhao 0024
  • Jie Li
  • Rui Huang 0008
  • Boqi Li 0001
  • Ao Luo
  • Yaochen Li
  • Hong Cheng 0002

Experienced human drivers always make safe driving decisions by selectively observing the front, rear and side- view mirrors. Several end - to-end methods have been pro-posed to learn driving models with multi-view visual infor-mation. However, these benchmark methods lack semantic understanding of multi-view image contents, where human drivers usually reason these information for decision making with different visual region of interests. In this paper, we propose an attention-based deep learning method to learn a driving model with input of surround-view visual information and the route planner, in which a multi-view attention module is designed for obtaining region of interests from human drivers. We evaluate our model on the Drive360 dataset with comparison of benchmarking deep driving models. Results demonstrate that our model achieves a competitive accuracy in both steering angle and speed prediction than benchmarking methods. Code is available at https://githuh.com/jet-uestc/MVA-Net.

AAAI Conference 2022 Conference Paper

Fully Attentional Network for Semantic Segmentation

  • Qi Song
  • Jie Li
  • Chenghong Li
  • Hao Guo
  • Rui Huang

Recent non-local self-attention methods have proven to be effective in capturing long-range dependencies for semantic segmentation. These methods usually form a similarity map of RC×C (by compressing spatial dimensions) or RHW ×HW (by compressing channels) to describe the feature relations along either channel or spatial dimensions, where C is the number of channels, H and W are the spatial dimensions of the input feature map. However, such practices tend to condense feature dependencies along the other dimensions, hence causing attention missing, which might lead to inferior results for small/thin categories or inconsistent segmentation inside large objects. To address this problem, we propose a new approach, namely Fully Attentional Network (FLANet), to encode both spatial and channel attentions in a single similarity map while maintaining high computational efficiency. Specifically, for each channel map, our FLANet can harvest feature responses from all other channel maps, and the associated spatial positions as well, through a novel fully attentional module. Our new method has achieved state-of-the-art performance on three challenging semantic segmentation datasets, i. e. , 83. 6%, 46. 99%, and 88. 5% on the Cityscapes test set, the ADE20K validation set, and the PASCAL VOC test set, respectively.

AAAI Conference 2022 Conference Paper

Input-Specific Robustness Certification for Randomized Smoothing

  • Ruoxin Chen
  • Jie Li
  • Junchi Yan
  • Ping Li
  • Bin Sheng

Although randomized smoothing has demonstrated high certified robustness and superior scalability to other certified defenses, the high computational overhead of the robustness certification bottlenecks the practical applicability, as it depends heavily on the large sample approximation for estimating the confidence interval. In existing works, the sample size for the confidence interval is universally set and agnostic to the input for prediction. This Input-Agnostic Sampling (IAS) scheme may yield a poor Average Certified Radius (ACR)-runtime trade-off which calls for improvement. In this paper, we propose Input-Specific Sampling (ISS) acceleration to achieve the cost-effectiveness for robustness certification, in an adaptive way of reducing the sampling size based on the input characteristic. Furthermore, our method universally controls the certified radius decline from the ISS sample size reduction. The empirical results on CIFAR-10 and ImageNet show that ISS can speed up the certification by more than three times at a limited cost of 0. 05 certified radius. Meanwhile, ISS surpasses IAS on the average certified radius across the extensive hyperparameter settings. Specifically, ISS achieves ACR=0. 958 on ImageNet in 250 minutes, compared to ACR=0. 917 by IAS under the same condition. We release our code in https: //github. com/roy-ch/Input-Specific-Certification.

YNIMG Journal 2022 Journal Article

Intracranial brain-computer interface spelling using localized visual motion response

  • Dingkun Liu
  • Xin Xu
  • Dongyang Li
  • Jie Li
  • Xinguang Yu
  • Zhipei Ling
  • Bo Hong

Intracranial brain-computer interfaces (BCIs) can assist severely disabled persons in text communication and environmental control with high precision and speed. Nevertheless, sustainable BCI implants require minimal invasiveness. One of the implantation strategies is to adopt localized and robust cortical activities to drive BCI communication and to make a precise presurgical planning. The visual motion response is a good candidate for inclusion in this strategy because of its focal activity over the middle temporal visual area (MT). Here, we developed an intracranial BCI for spelling, utilizing only three electrodes over the MT area. The best recording electrodes were decided by preoperative functional magnetic resonance imaging (MRI) localization of the MT, and local neural activities were further enhanced by differential rereferencing of these electrodes. The BCI spelling system was validated both offline and online by five epilepsy patients, achieving the fastest speed of 62 bits/min, i.e., 12 characters/min. Moreover, the response patterns of dual-directional visual motion stimuli provided an additional dimension of BCI target encoding and paved the way for a higher information transfer rate of intracranial BCI spelling.

AAAI Conference 2022 Conference Paper

Learning to Learn Transferable Attack

  • Shuman Fang
  • Jie Li
  • Xianming Lin
  • Rongrong Ji

Transfer adversarial attack is a non-trivial black-box adversarial attack that aims to craft adversarial perturbations on the surrogate model and then apply such perturbations to the victim model. However, the transferability of perturbations from existing methods is still limited, since the adversarial perturbations are easily overfitting with a single surrogate model and specific data pattern. In this paper, we propose a Learning to Learn Transferable Attack (LLTA) method, which makes the adversarial perturbations more generalized via learning from both data and model augmentation. For data augmentation, we adopt simple random resizing and padding. For model augmentation, we randomly alter the back propagation instead of the forward propagation to eliminate the effect on the model prediction. By treating the attack of both specific data and a modified model as a task, we expect the adversarial perturbations to adopt enough tasks for generalization. To this end, the meta-learning algorithm is further introduced during the iteration of perturbation generation. Empirical results on the widely-used dataset demonstrate the effectiveness of our attack method with a 12. 85% higher success rate of transfer attack comparing with the state-of-the-art methods. We also evaluate our method on the real-world online system, i. e. , Google Cloud Vision API, to further show the practical potentials of our method.

TIST Journal 2022 Journal Article

Utility-aware and Privacy-preserving Trajectory Synthesis Model that Resists Social Relationship Privacy Attacks

  • Zhirun Zheng
  • Zhetao Li
  • Jie Li
  • Hongbo Jiang
  • Tong Li
  • Bin Guo

For academic research and business intelligence, trajectory data has been widely collected and analyzed. Releasing trajectory data to a third party may lead to serious privacy leakage, which has spawned considerable researches on trajectory privacy protection technology. However, existing work suffers from several shortcomings. They either focus on point-based location privacy, ignoring the spatio-temporal correlations among locations within a trajectory, or they protect the privacy of each user separately without considering privacy leakage of the social relationship between trajectories of different users. Besides, they fail to balance privacy protection and data utility. Motivated by these limitations, in this article, we propose S 3 T -Trajectory, which is a utility-aware and privacy-preserving trajectory synthesis model that Resists social relationship privacy attacks. Specifically, we first develop a time-dependent Markov chain based on an adaptive spatio-temporal discrete grid to efficiently and accurately capture human mobility behavior. Then, we propose three mobility feature metrics from spatio-temporal, semantic, and social dimensions. On the basis of the metrics, we construct a bi-level optimization problem to accomplish the utility-aware and privacy-preserving trajectory synthesizing. The upper-level objective guarantees data utility and the lower-level optimization problems (or upper-level constraints) provides two-layer privacy protection for S 3 T -Trajectory, i.e., resisting location inference attacks and social relationship privacy attacks. We conduct extensive experiments on large-scale real-world datasets loc-Gowalla and loc-Brightkite. The experimental results demonstrate the effectiveness and robustness of S 3 T Trajectory. Compared with the baseline models, S 3 T Trajectory achieves between 7.8% and 23.8% performance improvement in resisting social relationship privacy attacks and achieves at least 5.19% improvement regarding data utility.

IJCAI Conference 2021 Conference Paper

IMENet: Joint 3D Semantic Scene Completion and 2D Semantic Segmentation through Iterative Mutual Enhancement

  • Jie Li
  • Laiyan Ding
  • Rui Huang

3D semantic scene completion and 2D semantic segmentation are two tightly correlated tasks that are both essential for indoor scene understanding, because they predict the same semantic classes, using positively correlated high-level features. Current methods use 2D features extracted from early-fused RGB-D images for 2D segmentation to improve 3D scene completion. We argue that this sequential scheme does not ensure these two tasks fully benefit each other, and present an Iterative Mutual Enhancement Network (IMENet) to solve them jointly, which interactively refines the two tasks at the late prediction stage. Specifically, two refinement modules are developed under a unified framework for the two tasks. The first is a 2D Deformable Context Pyramid (DCP) module, which receives the projection from the current 3D predictions to refine the 2D predictions. In turn, a 3D Deformable Depth Attention (DDA) module is proposed to leverage the reprojected results from 2D predictions to update the coarse 3D predictions. This iterative fusion happens to the stable high-level features of both tasks at a late stage. Extensive experiments on NYU and NYUCAD datasets verify the effectiveness of the proposed iterative late fusion scheme, and our approach outperforms the state of the art on both 3D semantic scene completion and 2D semantic segmentation.

JBHI Journal 2021 Journal Article

ML-Net: Multi-Channel Lightweight Network for Detecting Myocardial Infarction

  • Yangjie Cao
  • Tingting Wei
  • Bo Zhang
  • Nan Lin
  • Joel J. P. C. Rodrigues
  • Jie Li
  • Di Zhang

Due to the complexity of myocardial infarction (MI) waveform, most traditional automatic diagnosis models rarely detect it, while those able to detect MI often require high computing and storage capacity, rendering them unsuitable for portable devices. Therefore, in order for convenient real-time MI detection, it is essential to design lightweight models suitable for resource-limited portable devices. This paper proposes a novel multi-channel lightweight model (ML-Net), that provides a new solution for portable detection devices with limited resources. In ML-Net, each electrocardiogram (ECG) lead is assigned an independent channel, ensuring data independence and preserve the ECG characteristics of different angles represented by different leads. Moreover, convolution kernels of heterogeneous sizes are utilized to achieve accurate classification with only a small amount of lead data. Extensive experiments over actual ECG data from the PTB diagnostic database are conducted to evaluate ML-Net. The results show that ML-Net outperforms comparable schemes in diagnosing MI, and it requires lower computational cost and less memory, so that portable devices can be more widely used in the field of Internet of Medical Things(IoMT).

YNIMG Journal 2021 Journal Article

Reconstructing seen image from brain activity by visually-guided cognitive representation and adversarial learning

  • Ziqi Ren
  • Jie Li
  • Xuetong Xue
  • Xin Li
  • Fan Yang
  • Zhicheng Jiao
  • Xinbo Gao

Reconstructing perceived stimulus (image) only from human brain activity measured with functional Magnetic Resonance Imaging (fMRI) is a significant task in brain decoding. However, the inconsistent distribution and representation between fMRI signals and visual images cause great 'domain gap'. Moreover, the limited fMRI data instances generally suffer from the issues of low signal noise ratio (SNR), extremely high dimensionality, and limited spatial resolution. Existing methods are often affected by these issues so that a satisfactory reconstruction is still an open problem. In this paper, we show that it is possible to obtain a promising solution by learning visually-guided latent cognitive representations from the fMRI signals, and inversely decoding them to the image stimuli. The resulting framework is called Dual-Variational Autoencoder/ Generative Adversarial Network (D-Vae/Gan), which combines the advantages of adversarial representation learning with knowledge distillation. In addition, we introduce a novel three-stage learning strategy which enables the (cognitive) encoder to gradually distill useful knowledge from the paired (visual) encoder during the learning process. Extensive experimental results on both artificial and natural images have demonstrated that our method could achieve surprisingly good results and outperform the available alternatives.

AAAI Conference 2021 Conference Paper

Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene Completion

  • Xu Yan
  • Jiantao Gao
  • Jie Li
  • Ruimao Zhang
  • Zhen Li
  • Rui Huang
  • Shuguang Cui

LiDAR point cloud analysis is a core task for 3D computer vision, especially for autonomous driving. However, due to the severe sparsity and noise interference in the single sweep Li- DAR point cloud, the accurate semantic segmentation is nontrivial to achieve. In this paper, we propose a novel sparse Li- DAR point cloud semantic segmentation framework assisted by learned contextual shape priors. In practice, an initial semantic segmentation (SS) of a single sweep point cloud can be achieved by any appealing network and then flows into the semantic scene completion (SSC) module as the input. By merging multiple frames in the LiDAR sequence as supervision, the optimized SSC module has learned the contextual shape priors from sequential LiDAR data, completing the sparse single sweep point cloud to the dense one. Thus, it inherently improves SS optimization through fully end-toend training. Besides, a Point-Voxel Interaction (PVI) module is proposed to further enhance the knowledge fusion between SS and SSC tasks, i. e. , promoting the interaction of incomplete local geometry of point cloud and complete voxelwise global structure. Furthermore, the auxiliary SSC and PVI modules can be discarded during inference without extra burden for SS. Extensive experiments confirm that our JS3C- Net achieves superior performance on both SemanticKITTI and SemanticPOSS benchmarks, i. e. , 4% and 3% improvement correspondingly.

AAAI Conference 2020 Conference Paper

Facial Attribute Capsules for Noise Face Super Resolution

  • Jingwei Xin
  • Nannan Wang
  • Xinrui Jiang
  • Jie Li
  • Xinbo Gao
  • Zhifeng Li

Existing face super-resolution (SR) methods mainly assume the input image to be noise-free. Their performance degrades drastically when applied to real-world scenarios where the input image is always contaminated by noise. In this paper, we propose a Facial Attribute Capsules Network (FACN) to deal with the problem of high-scale super-resolution of noisy face image. Capsule is a group of neurons whose activity vector models different properties of the same entity. Inspired by the concept of capsule, we propose an integrated representation model of facial information, which named Facial Attribute Capsule (FAC). In the SR processing, we first generated a group of FACs from the input LR face, and then reconstructed the HR face from this group of FACs. Aiming to effectively improve the robustness of FAC to noise, we generate FAC in semantic, probabilistic and facial attributes manners by means of integrated learning strategy. Each FAC can be divided into two sub-capsules: Semantic Capsule (SC) and Probabilistic Capsule (PC). Them describe an explicit facial attribute in detail from two aspects of semantic representation and probability distribution. The group of FACs model an image as a combination of facial attribute information in the semantic space and probabilistic space by an attributedisentangling way. The diverse FACs could better combine the face prior information to generate the face images with fine-grained semantic attributes. Extensive benchmark experiments show that our method achieves superior hallucination results and outperforms state-of-the-art for very low resolution (LR) noise face image super resolution.

JBHI Journal 2020 Journal Article

Progressive Sub-Band Residual-Learning Network for MR Image Super Resolution

  • Xuetong Xue
  • Ying Wang
  • Jie Li
  • Zhicheng Jiao
  • Ziqi Ren
  • Xinbo Gao

High-resolution (HR) magnetic resonance images (MRI) provide more detailed information for clinical application. However, HR MRI is less available because of the longer scan time and lower signal-to-noise ratio. Spatial resolution is one of the key parameters of MRI. The image post-processing technique super-resolution (SR) is an alternative approach to improve the spatial resolution of MR images. Inspired by advanced deep learning based SR methods, we propose an MRI SR model named progressive sub-band residual learning SR network (PSR-SRN). The proposed model contains two parallel progressive learning streams, where one stream learns on missed high-frequency residuals by sub-band residual learning unit (ISRL) and the other focuses on reconstructing refined MR image. These two streams complement each other and enable to learn complex mappings between “Low-” and “High-” resolution MR images. Besides, we introduce brain-like mechanisms (in-depth supervision and local feedback mechanism) and progressive sub-band learning strategy to emphasize variant textures of MRI. Compared with traditional and deep learning MRI SR methods, our PSR-SRN model shows superior performance.

AAAI Conference 2020 Conference Paper

Video Face Super-Resolution with Motion-Adaptive Feedback Cell

  • Jingwei Xin
  • Nannan Wang
  • Jie Li
  • Xinbo Gao
  • Zhifeng Li

Video super-resolution (VSR) methods have recently achieved a remarkable success due to the development of deep convolutional neural networks (CNN). Current state-of-the-art CNN methods usually treat the VSR problem as a large number of separate multi-frame super-resolution tasks, at which a batch of low resolution (LR) frames is utilized to generate a single high resolution (HR) frame, and running a slide window to select LR frames over the entire video would obtain a series of HR frames. However, duo to the complex temporal dependency between frames, with the number of LR input frames increase, the performance of the reconstructed HR frames become worse. The reason is in that these methods lack the ability to model complex temporal dependencies and hard to give an accurate motion estimation and compensation for VSR process. Which makes the performance degrade drastically when the motion in frames is complex. In this paper, we propose a Motion-Adaptive Feedback Cell (MAFC), a simple but effective block, which can efficiently capture the motion compensation and feed it back to the network in an adaptive way. Our approach efficiently utilizes the information of the inter-frame motion, the dependence of the network on motion estimation and compensation method can be avoid. In addition, benefiting from the excellent nature of MAFC, the network can achieve better performance in the case of extremely complex motion scenarios. Extensive evaluations and comparisons validate the strengths of our approach, and the experimental results demonstrated that the proposed framework is outperform the state-of-the-art methods.

IJCAI Conference 2019 Conference Paper

Face Photo-Sketch Synthesis via Knowledge Transfer

  • Mingrui Zhu
  • Nannan Wang
  • Xinbo Gao
  • Jie Li
  • Zhifeng Li

Despite deep neural networks have demonstrated strong power in face photo-sketch synthesis task, their performance, however, are still limited by the lack of training data (photo-sketch pairs). Knowledge Transfer (KT), which aims at training a smaller and fast student network with the information learned from a larger and accurate teacher network, has attracted much attention recently due to its superior performance in the acceleration and compression of deep neural networks. This work has brought us great inspiration that we can train a relatively small student network on very few training data by transferring knowledge from a larger teacher model trained on enough training data for other tasks. Therefore, we propose a novel knowledge transfer framework to synthesize face photos from face sketches or synthesize face sketches from face photos. Particularly, we utilize two teacher networks trained on large amount of data in related task to learn the knowledge of face photos and face sketches separately and transfer them to two student networks simultaneously. In addition, the two student networks, one for photo? sketch task and the other for sketch? photo task, can transfer their knowledge mutually. With the proposed method, we can train our model which has superior performance using a small set of photo-sketch pairs. We validate the effectiveness of our method across several datasets. Quantitative and qualitative evaluations illustrate that our model outperforms other state-of-the-art methods in generating face sketches (or photos) with high visual quality and recognition ability.

IJCAI Conference 2019 Conference Paper

Group Reconstruction and Max-Pooling Residual Capsule Network

  • Xinpeng Ding
  • Nannan Wang
  • Xinbo Gao
  • Jie Li
  • Xiaoyu Wang

In capsule networks, the mapping of low-level capsules to high-level capsules is achieved by a routing-by-agreement algorithm. Since the capsule is made up of collections of neurons and the routing mechanism involves all the capsules instead of simply discarding some of the neurons like Max-Pooling, the capsule network has stronger representation ability than the traditional neural network. However, considering too much low-level capsules' information will cause its corresponding upper layer capsules to be interfered by other irrelevant information or noise capsules. Therefore, the original capsule network does not perform well on complex data structure. What's worse, computational complexity becomes a bottleneck in dealing with large data networks. In order to solve these shortcomings, this paper proposes a group reconstruction and max-pooling residual capsule network (GRMR-CapsNet). We build a block in which all capsules are divided into different groups and perform group reconstruction routing algorithm to obtain the corresponding high-level capsules. Between the lower and higher layers, Capsule Max-Pooling is adopted to prevent overfitting. We conduct experiments on CIFAR-10/100 and SVHN datasets and the results show that our method can perform better against state-of-the-arts.

AAAI Conference 2019 Conference Paper

HSME: Hypersphere Manifold Embedding for Visible Thermal Person Re-Identification

  • Yi Hao
  • Nannan Wang
  • Jie Li
  • Xinbo Gao

Person Re-identification(re-ID) has great potential to contribute to video surveillance that automatically searches and identifies people across different cameras. Heterogeneous person re-identification between thermal(infrared) and visible images is essentially a cross-modality problem and important for night-time surveillance application. Current methods usually train a model by combining classification and metric learning algorithms to obtain discriminative and robust feature representations. However, the combined loss function ignored the correlation between classification subspace and feature embedding subspace. In this paper, we use Sphere Softmax to learn a hypersphere manifold embedding and constrain the intra-modality variations and cross-modality variations on this hypersphere. We propose an end-to-end dualstream hypersphere manifold embedding network(HSMEnet) with both classification and identification constraint. Meanwhile, we design a two-stage training scheme to acquire decorrelated features, we refer the HSME with decorrelation as D-HSME. We conduct experiments on two crossmodality person re-identification datasets. Experimental results demonstrate that our method outperforms the state-of-the-art methods on two datasets. On RegDB dataset, rank-1 accuracy is improved from 33.47% to 50.85%, and mAP is improved from 31.83% to 47.00%.

AAAI Conference 2019 Conference Paper

Learning Neural Bag-of-Matrix-Summarization with Riemannian Network

  • Hong Liu
  • Jie Li
  • Yongjian Wu
  • Rongrong Ji

Symmetric positive defined (SPD) matrix has attracted increasing research focus in image/video analysis, which merits in capturing the Riemannian geometry in its structured 2D feature representation. However, computation in the vector space on SPD matrices cannot capture the geometric properties, which corrupts the classification performance. To this end, Riemannian based deep network has become a promising solution for SPD matrix classification, because of its excellence in performing non-linear learning over SPD matrix. Besides, Riemannian metric learning typically adopts a kNN classifier that cannot be extended to large-scale datasets, which limits its application in many time-efficient scenarios. In this paper, we propose a Bag-of-Matrix-Summarization (BoMS) method to be combined with Riemannian network, which handles the above issues towards highly efficient and scalable SPD feature representation. Our key innovation lies in the idea of summarizing data in a Riemannian geometric space instead of the vector space. First, the whole training set is compressed with a small number of matrix features to ensure high scalability. Second, given such a compressed set, a constant-length vector representation is extracted by efficiently measuring the distribution variations between the summarized data and the latent feature of the Riemannian network. Finally, the proposed BoMS descriptor is integrated into the Riemannian network, upon which the whole framework is end-to-end trained via matrix back-propagation. Experiments on four different classification tasks demonstrate the superior performance of the proposed method over the state-of-the-art methods.

IJCAI Conference 2019 Conference Paper

Multi-Margin based Decorrelation Learning for Heterogeneous Face Recognition

  • Bing Cao
  • Nannan Wang
  • Xinbo Gao
  • Jie Li
  • Zhifeng Li

Heterogeneous face recognition (HFR) refers to matching face images acquired from different domains with wide applications in security scenarios. However, HFR is still a challenging problem due to the significant cross-domain discrepancy and the lacking of sufficient training data in different domains. This paper presents a deep neural network approach namely Multi-Margin based Decorrelation Learning (MMDL) to extract decorrelation representations in a hyperspherical space for cross-domain face images. The proposed framework can be divided into two components: heterogeneous representation network and decorrelation representation learning. First, we employ a large scale of accessible visual face images to train heterogeneous representation network. The decorrelation layer projects the output of the first component into decorrelation latent subspace and obtain decorrelation representation. In addition, we design a multi-margin loss (MML), which consists of tetradmargin loss (TML) and heterogeneous angular margin loss (HAML), to constrain the proposed framework. Experimental results on two challenging heterogeneous face databases show that our approach achieves superior performance on both verification and recognition tasks, comparing with state-of-the-art methods.

AAAI Conference 2019 Conference Paper

Residual Attribute Attention Network for Face Image Super-Resolution

  • Jingwei Xin
  • Nannan Wang
  • Xinbo Gao
  • Jie Li

Facial prior knowledge based methods recently achieved great success on the task of face image super-resolution (SR). The combination of different type of facial knowledge could be leveraged for better super-resolving face images, e.g., facial attribute information with texture and shape information. In this paper, we present a novel deep end-to-end network for face super resolution, named Residual Attribute Attention Network (RAAN), which realizes the efficient feature fusion of various types of facial information. Specifically, we construct a multi-block cascaded structure network with dense connection. Each block has three branches: Texture Prediction Network (TPN), Shape Generation Network (SGN) and Attribute Analysis Network (AAN). We divide the task of face image reconstruction into three steps: extracting the pixel level representation information from the input very low resolution (LR) image via TPN and SGN, extracting the semantic level representation information by AAN from the input, and finally combining the pixel level and semantic level information to recover the high resolution (HR) image. Experiments on benchmark database illustrate that RAAN significantly outperforms state-of-the-arts for very low-resolution face SR problem, both quantitatively and qualitatively.

IROS Conference 2019 Conference Paper

ViLiVO: Virtual LiDAR-Visual Odometry for an Autonomous Vehicle with a Multi-Camera System

  • Zhenzhen Xiang
  • Jingrui Yu
  • Jie Li
  • Jianbo Su

In this paper, we present a multi-camera visual odometry (VO) system for an autonomous vehicle. Our system mainly consists of a virtual LiDAR and a pose tracker. We use a perspective transformation method to synthesize a surroundview image from undistorted fisheye camera images. With a semantic segmentation model, the free space can be extracted. The scans of the virtual LiDAR are generated by discretizing the contours of the free space. As for the pose tracker, we propose a visual odometry system fusing both the feature matching and the virtual LiDAR scan matching results. Only those feature points located in the free space area are utilized to ensure the 2D-2D matching for pose estimation. Furthermore, bundle adjustment (BA) is performed to minimize the feature points reprojection error and scan matching error. We apply our system to an autonomous vehicle equipped with four fisheye cameras. The testing scenarios include an outdoor parking lot as well as an indoor garage. Experimental results demonstrate that our system achieves a more robust and accurate performance comparing with a fisheye camera based monocular visual odometry system.

IJCAI Conference 2018 Conference Paper

Aspect Sentiment Classification with both Word-level and Clause-level Attention Networks

  • Jingjing Wang
  • Jie Li
  • Shoushan Li
  • Yangyang Kang
  • Min Zhang
  • Luo Si
  • Guodong Zhou

Aspect sentiment classification, a challenging task in sentiment analysis, has been attracting more and more attention in recent years. In this paper, we highlight the need for incorporating the importance degrees of both words and clauses inside a sentence and propose a hierarchical network with both word-level and clause-level attentions to aspect sentiment classification. Specifically, we first adopt sentence-level discourse segmentation to segment a sentence into several clauses. Then, we leverage multiple Bi-directional LSTM layers to encode all clauses and propose a word-level attention layer to capture the importance degrees of words in each clause. Third and finally, we leverage another Bi-directional LSTM layer to encode the outputs from the former layers and propose a clause-level attention layer to capture the importance degrees of all the clauses inside a sentence. Experimental results on the laptop and restaurant datasets from SemEval-2015 demonstrate the effectiveness of our proposed approach to aspect sentiment classification.

AAAI Conference 2018 Conference Paper

Asymmetric Joint Learning for Heterogeneous Face Recognition

  • Bing Cao
  • Nannan Wang
  • Xinbo Gao
  • Jie Li

Heterogeneous face recognition (HFR) refers to matching a probe face image taken from one modality to face images acquired from another modality. It plays an important role in security scenarios. However, HFR is still a challenging problem due to great discrepancies between cross-modality images. This paper proposes an asymmetric joint learning (AJL) approach to handle this issue. The proposed method transforms the cross-modality differences mutually by incorporating the synthesized images into the learning process which provides more discriminative information. Although the aggregated data would augment the scale of intraclasses, it also reduces the diversity (i. e. discriminative information) for inter-classes. Then, we develop the AJL model to balance this dilemma. Finally, we could obtain the similarity score between two heterogeneous face images through the log-likelihood ratio. Extensive experiments on viewed sketch database, forensic sketch database and near infrared image database illustrate that the proposed AJL-HFR method achieve superior performance in comparison to state-of-theart methods.

IJCAI Conference 2018 Conference Paper

Deep Attribute Guided Representation for Heterogeneous Face Recognition

  • Decheng Liu
  • Nannan Wang
  • Chunlei Peng
  • Jie Li
  • Xinbo Gao

Heterogeneous face recognition (HFR) is a challenging problem in face recognition, subject to large texture and spatial structure differences of face images. Different from conventional face recognition in homogeneous environments, there exist many face images taken from different sources (including different sensors or different mechanisms) in reality. Motivated by human cognitive mechanism, we naturally utilize the explicit invariant semantic information (face attributes) to help address the gap of different modalities. Existing related face recognition methods mostly regard attributes as the high level feature integrated with other engineering features enhancing recognition performance, ignoring the inherent relationship between face attributes and identities. In this paper, we propose a novel deep attribute guided representation based heterogeneous face recognition method (DAG-HFR) without labeling attributes manually. Deep convolutional networks are employed to directly map face images in heterogeneous scenarios to a compact common space where distances mean similarities of pairs. An attribute guided triplet loss (AGTL) is designed to train an end-to-end HFR network which could effectively eliminate defects of incorrectly detected attributes. Extensive experiments on multiple heterogeneous scenarios (composite sketches, resident ID cards) demonstrate that the proposed method achieves superior performances compared with state-of-the-art methods.

IJCAI Conference 2018 Conference Paper

From Reality to Perception: Genre-Based Neural Image Style Transfer

  • Zhuoqi Ma
  • Nannan Wang
  • Xinbo Gao
  • Jie Li

We introduce a novel thought for integrating artists’ perceptions on the real world into neural image style transfer process. Conventional approaches commonly migrate color or texture patterns from style image to content image, but the underlying design aspect of the artist always get overlooked. We want to address the in-depth genre style, that how artists perceive the real world and express their perceptions in the artwork. We collect a set of Van Gogh’s paintings and cubist artworks, and their semantically corresponding real world photos. We present a novel genre style transfer framework modeled after the mechanism of actual artwork production. The target style representation is reconstructed based on the semantic correspondence between real world photo and painting, which enable the perception guidance in style transfer. The experimental results demonstrate that our method can capture the overall style of a genre or an artist. We hope that this work provides new insight for including artists’ perceptions into neural style transfer process, and helps people to understand the underlying characters of the artist or the genre.

IJCAI Conference 2017 Conference Paper

Deep Graphical Feature Learning for Face Sketch Synthesis

  • Mingrui Zhu
  • Nannan Wang
  • Xinbo Gao
  • Jie Li

The exemplar-based face sketch synthesis method generally contains two steps: neighbor selection and reconstruction weight representation. Pixel intensities are widely used as features by most of the existing exemplar-based methods, which lacks of representation ability and robustness to light variations and clutter backgrounds. We present a novel face sketch synthesis method combining generative exemplar-based method and discriminatively trained deep convolutional neural networks (dCNNs) via a deep graphical feature learning framework. Our method works in both two steps by using deep discriminative representations derived from dCNNs. Instead of using it directly, we boost its representation capability by a deep graphical feature learning framework. Finally, the optimal weights of deep representations and optimal reconstruction weights for face sketch synthesis can be obtained simultaneously. With the optimal reconstruction weights, we can synthesize high quality sketches which is robust against light variations and clutter backgrounds. Extensive experiments on public face sketch databases show that our method outperforms state-of-the-art methods, in terms of both synthesis quality and recognition ability.

TCS Journal 2015 Journal Article

Improved even order magic square construction algorithms and their applications in multi-user shared electronic accounts

  • Zhenhua Duan
  • Jin Liu
  • Jie Li
  • Cong Tian

This paper presents improved even order magic square construction algorithms, including both single even order magic square and double even order magic square construction algorithms. Further, in order to show how the algorithms work, two specific magic squares are constructed. Moreover, the correctness of the algorithms is proved, and the complexity analysis of the algorithms is given. Finally, the improved even order magic square construction algorithms are applied in secure communication and authentication areas for multi-user shared electronic account in detail.

v2026.09.13