Arrow Research search

Author name cluster

Zhe Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

47 papers
2 author rows

Possible papers

47

EAAI Journal 2026 Journal Article

A linear Diophantine fuzzy hybrid decision support system for sustainability evaluation of renewable energy resources

  • Zhe Liu
  • Zhipeng Chen
  • Sukumar Letchmunan
  • Yulong Huang
  • Tapan Senapati
  • Dragan Pamucar

As the global energy system increasingly transforms to renewable energy sources (RERs), such as solar, wind, hydro, biomass, and geothermal, there is an increasing need for intelligent and robust decision-making models that can assess sustainability in multiple dimensions, including environmental, economic, technical, and social. However, existing multi-criteria decision-making (MCDM) methods often struggle with modeling uncertainty and subjectivity. They typically lack the flexibility to integrate both objective data and expert judgment in a unified, interpretable framework. To address these limitations, this paper proposes an artificial intelligence-driven hybrid decision support system combined with linear Diophantine fuzzy sets (LDFSs) to model the uncertainty in RERs evaluation. We first introduce new entropy and divergence measures for LDFSs to accurately quantify uncertainty and information diversity, overcoming the shortcomings of existing measures that may yield inconsistent outcomes. Subsequently, we integrate entropy-divergence measures (EDM) for objective criteria weighting with step-wise weight assessment ratio analysis (SWARA) to incorporate subjective expert preferences, ensuring a balanced weighting strategy. Finally, we employ the operational competitiveness rating analysis (OCRA) method to rank alternatives, enhancing its interpretability in complex decision-making. The proposed model is validated through a case study on the evaluation of RERs. The model successfully ranked alternatives based on sustainability criteria, identifying wind energy as the optimal choice for sustainable energy deployment. Sensitivity and comparison analysis further confirm the robustness and effectiveness of the proposed model. This work provides a scientifically rigorous and practical artificial intelligence-based tool to support informed decision-making for sustainable RERs deployment in diverse and uncertain contexts.

EAAI Journal 2026 Journal Article

A semantic segmentation model for early-stage fire detection from aerial remote sensing

  • Zhe Liu
  • Yu Sun
  • Xiangyuan Jiang
  • Pei Duan
  • Ming Li

For forest fire disasters threatening to the ecological environment and human life safety, current research focuses solely on detecting either flame or smoke. This often leads to missed detection or false detection. In this paper, we propose a semantic segmentation model that aims to accurately segment flame and smoke simultaneously. A Compact Atrous Spatial Pyramid Pooling module is developed with the objective of capturing multi-scale contextual information efficiently, addressing the significant scale disparities between flame and smoke. Additionally, a Bottom-up Detail-informed Feature Fusion Module is proposed, which leverages shallow features to guide cross-layer feature fusion, thereby enhancing the detection accuracy of small targets. Lastly, a Foreground Emphasis Module is proposed to mitigate the issue of foreground sparsity that commonly exists in remote sensing images of early forest fires. This module utilizes foreground classification results to guide segmentation, making the model focus more on the identification of foreground. Experimental results suggest that our method markedly surpasses other methods in early-stage fire scenarios and achieves accurate disaster area segmentation in various scenarios such as urban fires. In addition, a processing speed of 41. 83 frames per second is attainable on TITAN Xp devices, which fully demonstrates its excellent segmentation performance and efficient real-time processing capability.

EAAI Journal 2026 Journal Article

Deep dynamic image prior for three-dimensional time-sequence pulmonary electrical impedance tomography

  • Hao Fang
  • Hao Yu
  • Sihao Teng
  • Tao Zhang
  • Siyi Yuan
  • Huaiwu He
  • Zhe Liu
  • Yunjie Yang

Unsupervised learning methods, such as Deep Image Prior (DIP), have shown great potential in engineering imaging due to their training-data-free nature and high generalization capability. However, their reliance on numerous network parameter iterations results in high computational costs, limiting their practical application, particularly in complex three-dimensional (3D) or time-sequence tomographic imaging tasks. To overcome these challenges, we propose Deep Dynamic Image Prior (D 2 IP), a novel framework for three-dimensional time-sequence imaging. D 2 IP introduces three key strategies — Unsupervised Parameter Warm-Start (UPWS), Temporal Parameter Propagation (TPP), and a customized lightweight reconstruction backbone, Three-dimensional Fast Residual U-Net (3D-FastResUNet) — to accelerate convergence, enforce temporal coherence, and improve computational efficiency. Experimental results on both simulated and clinical pulmonary datasets demonstrate that D 2 IP enables fast and accurate 3D time-sequence Electrical Impedance Tomography (tsEIT) reconstruction. Compared to the state-of-the-art Regularized Shallow Image Prior (R-SIP) baseline, D 2 IP delivers superior image quality — with a 24. 8% increase in average Mean Structural Similarity Index (MSSIM) and an 8. 1% reduction in Relative Error (ERR) — alongside significantly reduced computational time (7. 1× faster), demonstrating its promise for artificial intelligence (AI)-driven medical engineering applications, as exemplified by clinical dynamic pulmonary imaging.

AAAI Conference 2026 Conference Paper

Many Minds, One Path: LLM-Augmented Consensus Decision for Distributed Control in Multi-Agent Collaborative Stable Scenarios

  • Zhuohao Yu
  • Zhe Liu
  • Tao Ren
  • Chenxue Wang
  • Junjie Wang
  • Qing Wang

Distributed multi-agent systems are increasingly deployed in dynamic and high-stakes environments such as power grids, intelligent traffic systems, and collaborative robotics. In these systems, long-term stability, the ability to maintain coherent and safe system behavior over time, is critical but underexplored in existing research. This paper presents LLMASC, a framework designed to enhance long-term stability in multi-agent collaboration by combining semantic reasoning with decentralized control. LLMASC comprises three key components: a Semantic Perception Encoder that transforms heterogeneous agent observations into structured natural language; an LLM-Guided Consensus Decision module that enables strategic alignment through proposal exchange and voting; and a Policy Execution Controller that maps high-level plans to executable actions via reinforcement learning. We evaluate LLMASC across three representative simulation domains (Multi-Walker, Simulation of Urban Mobility and Power Grid Stabilization), spanning both physical and cyber-physical systems. Experiments show that LLMASC consistently outperforms the best baselines, improving stability rates by up to 44% and long-term success by 31%. Further analysis confirms its decision-making efficiency and robustness under varying agent populations and model choices.

EAAI Journal 2025 Journal Article

Auto feature weighted c -means type clustering methods for color image segmentation

  • Sijia Zhu
  • Zhe Liu
  • Sukumar Letchmunan
  • Haoye Qiu

To address the limitations of existing hard c -means (HCM) and fuzzy c -means (FCM) methods, we develop four novel clustering methods: vector-weighted alternative hard c -means (VWAHCM), matrix-weighted alternative hard c -means (MWAHCM), vector-weighted alternative fuzzy c -means (VWAFCM), and matrix-weighted alternative fuzzy c -means (MWAFCM). These methods enhance clustering performance by incorporating non-Euclidean norm metrics and vector-weighted and matrix-weighted schemes without adding extra parameters. Our methods modify the traditional weight constraint from a sum to a product of weights, thereby improving robustness and accuracy. Comprehensive experiments conduct on various real-world datasets and color image segmentation tasks demonstrate the superiority of the proposed methods over traditional HCM and FCM variants. The results show significant improvements in clustering Accuracy ( A C C ), Normalized mutual information ( N M I ), Rand index ( R I ), and Fowlkes–Mallows index ( F M ). Furthermore, the proposed methods exhibit fast convergence and robust performance, proving their effectiveness in practical applications.

JBHI Journal 2025 Journal Article

Bidirectional Prototype-Guided Consistency Constraint for Semi-Supervised Fetal Ultrasound Image Segmentation

  • Chongwen Lyu
  • Kai Han
  • Lu Liu
  • Jun Chen
  • Lele Ma
  • Zheng Pang
  • Zhe Liu

Fetal ultrasound (US) image segmentation plays an important role in fetal development assessment, maternal pregnancy management, and intrauterine surgery planning. However, obtaining large-scale, accurately annotated fetal US imaging data is time-consuming and labor-intensive, posing challenges to the application of deep learning in this field. To address this challenge, we propose a semi-supervised fetal US image segmentation method based on bidirectional prototype-guided consistency constraint (BiPCC). BiPCC utilizes the prototype to bridge labeled and unlabeled data and establishes interaction between them. Specifically, the model generates pseudo-labels using prototypes from labeled data and then utilizes these pseudo-labels to generate pseudo-prototypes for segmenting the labeled data inversely, thereby achieving bidirectional consistency. Additionally, uncertainty-based cross-supervision is incorporated to provide additional supervision signals, thereby enhancing the quality of pseudo-labels. Extensive experiments on two fetal US datasets demonstrate that BiPCC outperforms state-of-the-art methods for semi-supervised fetal US segmentation. Furthermore, experimental results on two additional medical segmentation datasets exhibit BiPCC's outstanding generalization capability for diverse medical image segmentation tasks. Our proposed method offers a novel insight for semi-supervised fetal US image segmentation and holds promise for further advancing the development of intelligent healthcare.

AAAI Conference 2025 Conference Paper

Differential Private Stochastic Optimization with Heavy-tailed Data: Towards Optimal Rates

  • Puning Zhao
  • Jiafei Wu
  • Zhe Liu
  • Chong Wang
  • Rongfei Fan
  • Qingming Li

We study convex optimization problems under differential privacy (DP). With heavy-tailed gradients, existing works achieve suboptimal rates. The main obstacle is that existing gradient estimators have suboptimal tail property, resulting in a superfluous factor of d in the union bound. In this paper, we explore algorithms achieving optimal rates of DP optimization with heavy-tailed gradients. Our first method is a simple clipping approach. Under bounded p-th order moments of gradients, with n samples, it achieves minimax optimal population risk with epsilon less than 1/d. We then propose an iterative updating method, which is more complex but achieves this rate for all epsilon smaller than 1. The results significantly improve over existing methods. Such improvement relies on a careful treatment of the tail behavior of gradient estimators. Our results match the minimax lower bound, indicating that the theoretical limit of stochastic convex optimization under DP is achievable.

AAAI Conference 2025 Conference Paper

DR-Encoder: Encode Low-rank Gradients with Random Prior for Large Language Models Differentially Privately

  • Huiwen Wu
  • Deyi Zhang
  • Xiaohan Li
  • Xiaogang Xu
  • Jiafei Wu
  • Zhe Liu

The emergence of the large language model (LLM) has shown its superiority in a wide range of disciplines, including language understanding and translation, relational logic reasoning, and even partial differential equations solving. The transformer is the pervasive backbone architecture for the foundation model construction. It is vital to research how to adjust the Transformer architecture to achieve an end-to-end privacy guarantee in LLM fine-tuning. This paper investigates three potential information leaks during a federated fine-tuning procedure for LLM (FedLLM). Based on the potential information leakage, we insert two-stage randomness into FedLLM to provide an end-to-end privacy guarantee solution. The first stage is to train a gradient auto-encoder with a Gaussian random prior based on the statistical information of the gradients generated by local clients. The second stage is fine-tuning the overall LLM with a differential privacy guarantee by adopting appropriate Gaussian noises. We show our proposed method's efficiency and accuracy gains with several foundation models and two popular evaluation benchmarks. Furthermore, we present a comprehensive privacy analysis with Gaussian Differential Privacy (GDP) and Renyi Differential Privacy (RDP).

EAAI Journal 2025 Journal Article

Enhancing neighborhood-based co-clustering contrastive learning for multi-entity recommendation

  • Juan Liao
  • Aman Jantan
  • Zhe Liu

To enhance recommendation performance, multi-behavior has become popular for addressing data sparsity in recommendation systems (RS). Meanwhile, due to its self-supervised nature which is not reliant on large amounts of labeled data, contrastive learning (CL) has emerged as a critical technique for solving the cold start-induced data sparsity problem. In the context of multi-behavior RS based on CL, existing methods mainly construct positive and negative sample pairs from implicit user behaviors. However, these methods often rely on black-box neural networks (a common artificial intelligence technique) to simulate the relationships between users and items, making the target nodes susceptible to unreliable neighborhood representations during high-order graph convolution. Specifically, (1) with respect to different user-relationship classifications, although users may belong to different clusters, they could also belong to multiple clusters simultaneously (for example, user u might belong to both sport-cluster and book-lover-cluster). For users associated with multiple clusters, the RS should provide clear guidance on how to effectively recommend items. (2) Because different clusters of users have different preferences, the feature extraction should also differ. If clustering feature extraction originates from unreliable neighbor nodes, it is prone to generating noisy data. To address these challenges, we propose the Enhancing Neighborhood-based Co-clustering Contrastive Learning (NCCL) model, which simultaneously leverages user-level and item-level CL from two perspectives: user-soft-clustering relationships (horizontal) and item-hard-clustering (vertical). Specifically, NCCL combines user-item and item-entity clusters. In the horizontal embedding, CL is performed based on user-centric user-item interactions, whereas in the vertical embedding, CL is conducted using item-centric item-entity aggregations. NCCL aims to incorporate multi-entity data into CL to mitigate data sparsity issues. The model extracts feature values of users and items from different granular perspectives, eliminating noise occasioned by neighboring clusters, which enhances the RS accuracy. Extensive experiments on four public datasets demonstrate that the proposed NCCL significantly outperforms the state-of-the-art techniques, with average improvements of 19. 38% in Hit Ratio (HR) and 38. 66% in Normalized Discounted Cumulative Gain (NDCG).

TCS Journal 2025 Journal Article

Finding fair and efficient allocations of indivisible chores

  • Zhe Liu
  • Wenguo Yang
  • Suixiang Gao

Fair resource allocation has found widespread application across various fields, such as economics and computer science, and has garnered significant attention. In this paper, we address the problem of allocating a set of indivisible chores among a group of agents. Our objective is to achieve an allocation that satisfies the fairness criterion of weighted equitable up to one item and the efficiency criterion of fractional Pareto-Optimal. Inspired by Fisher Market equilibrium [9], we construct a WEQ1 allocation among Market equilibrium through item transferring and price raising. Our algorithm provides a WEQ1+fPO allocation for additive valuations in pseudo-polynomial time. Moreover, we consider a special k-ary instance, i. e. , each agent has at most k different disutility values for chores when k is a constant. We show that a WEQ1+fPO allocation can be computed in polynomial time in such case.

AAAI Conference 2025 Conference Paper

FLAME: Learning to Navigate with Multimodal LLM in Urban Environments

  • Yunzhe Xu
  • Yiyuan Pan
  • Zhe Liu
  • Hesheng Wang

Large Language Models (LLMs) have demonstrated potential in Vision-and-Language Navigation (VLN) tasks, yet current applications face challenges. While LLMs excel in general conversation scenarios, they struggle with specialized navigation tasks, yielding suboptimal performance compared to specialized VLN models. We introduce FLAME (FLAMingo-Architected Embodied Agent), a novel Multimodal LLM-based agent and architecture designed for urban VLN tasks that efficiently handles multiple observations. Our approach implements a three-phase tuning technique for effective adaptation to navigation tasks, including single perception tuning for street view description, multiple perception tuning for route summarization, and end-to-end training on VLN datasets. The augmented datasets are synthesized automatically. Experimental results demonstrate FLAME's superiority over existing methods, surpassing state-of-the-art methods by a 7.3% increase in task completion on Touchdown dataset. This work showcases the potential of Multimodal LLMs (MLLMs) in complex navigation tasks, representing an advancement towards applications of MLLMs in the field of embodied intelligence.

EAAI Journal 2025 Journal Article

Fuzzy clustering-based dual-channel contrastive learning for multi-behavior recommendation

  • Juan Liao
  • Aman Jantan
  • Zhe Liu
  • Tapan Senapati
  • Gözde Ulutagay
  • Laith Abualigah
  • Omed Hassan Ahmed

In recommender systems, diverse user behaviors (e. g. , clicking, purchasing, sharing) provide valuable insights into user preferences. While multi-behavior recommendation models have shown promise, existing models often suffer from excessive complexity or fail to effectively capture relationships between behaviors. Two major challenges persist: (1) Most models focus primarily on user–item interactions, overlooking the dominant role of content data in real-world applications. Additionally, the high proportion of non-interacted items exacerbates the sparsity of target behavior data. (2) Current methods jointly model users and items but fail to explicitly distinguish their unique characteristics, leading to an incomplete understanding of diverse item behaviors. To address these issues, we propose Fuzzy Clustering-based Dual-Channel Contrastive Learning(a commonly used algorithm in artificial intelligence) (FCCL) model for multi-behavior recommendation. FCCL first employs a graph convolutional network to generate user and item embeddings independently, leveraging contrastive learning (CL) to capture explicit and implicit features. Subsequently, a dual-channel linear module based on fuzzy clustering is introduced to model both user interest diffusion and item provider influence. In the first layer, a user-level fuzzy clustering CL method is proposed to capture user–item similarities through a fused loss function. The second layer applies item-level hard clustering to characterize item-entity relationships, mitigating sparsity by identifying relevant items, including non-interacted ones. Finally, these tasks are integrated to enhance the quality of user and item embeddings, and a dual-channel optimization mechanism is established to optimize model parameters. Extensive experiments conducted on several public datasets demonstrate that FCCL significantly outperforms existing multi-behavior recommendation models in terms of effectiveness.

IROS Conference 2025 Conference Paper

GIPD: Global Intent Prediction and Decomposition of Cooperative Multi-Robot System in Non-Communication Environments

  • Yu Zhao
  • Zhe Liu
  • Haoyu Wei
  • Kai Wang
  • Haitao Wang
  • Duwen Zhai
  • Kefan Jin
  • Haibin Shao

In complex multi-robot application scenarios, particularly in dynamically adversarial, hazardous, or disaster environments, traditional cooperation paradigms face significant challenges due to unreliable or absent communication links. Achieving efficient cooperation in the absence of communication has become a key bottleneck limiting the performance of multirobot systems. In this paper, we propose a Global Intent Prediction and Decomposition (GIPD) framework that enables robots to perform cooperative behavior without relying on communication. Each robot independently infers a globally consistent intent based solely on its local observations, ensuring implicit alignment across the system. Given the inferred global intent, robots autonomously determine their responsibilities and select the most appropriate tasks. They then base their local decision-making on the global intent, selected tasks, and individual observations, thereby facilitating effective execution and cooperation. We validate our approach using the MPE and SMAC benchmarks. Additionally, real-world experiments involving multiple ships demonstrate the effectiveness and practical applicability of the proposed GIPD method.

IROS Conference 2025 Conference Paper

Hierarchical Collision-Free Configuration Planning for a Soft Manipulator

  • Yi Shen
  • Ruochen Tai
  • Feiyu Hu
  • Zhe Liu

Soft manipulators (SMs) have shown great potential for interactive tasks in confined environments. However, avoiding obstacles of SMs may conflict with the manipulator’s configuration, the planned trajectory for tracking control, and the target position for grasping. To coordinate configuration planning, tracking control, and target grasping in obstacle avoidance, this study proposes a hierarchical configuration planning framework with three levels: behavior planning, configuration planning, and shape/position control. At the behavior planning level, a Discrete Event System (DES)-based planner is designed to orchestrate mode transitions among obstacle avoidance, tracking control, and target grasping. The configuration planning level adopts the Bézier curve to model the SM backbone curve and constructs a repulsive potential field to quantify obstacle effects on the entire manipulator configuration. Under the constraints of grasping distance and material physical limit, the control points of the Bézier curve corresponding to the optimal configuration that minimizes the repulsive potential energy are computed. Experiments demonstrate the effectiveness of the proposed framework in achieving collision-free configuration planning for object grasping and placement in confined operational scenarios.

IROS Conference 2025 Conference Paper

Hybrid Transformer-Mamba Model for 3D Semantic Segmentation

  • Xinyu Wang
  • Jinghua Hou
  • Zhe Liu
  • Yingying Zhu

Transformer-based methods have demonstrated remarkable capabilities in 3D semantic segmentation through their powerful attention mechanisms, but the quadratic complexity limits their modeling of long-range dependencies in large-scale point clouds. While recent Mamba-based approaches offer efficient processing with linear complexity, they struggle with feature representation when extracting 3D features. However, effectively combining these complementary strengths remains an open challenge in this field. In this paper, we propose HybridTM, the first hybrid architecture that integrates Transformer and Mamba for 3D semantic segmentation. In addition, we propose the Inner Layer Hybrid Strategy, which combines attention and Mamba at a finer granularity, enabling simultaneous capture of long-range dependencies and fine-grained local features. Extensive experiments demonstrate the effectiveness and generalization of our HybridTM on diverse indoor and outdoor datasets. Furthermore, our HybridTM achieves state-of-the-art performance on ScanNet, ScanNet200, and nuScenes benchmarks. The code will be made available at https://github.com/deepinact/HybridTM.

AAAI Conference 2025 Conference Paper

Improving Integrated Gradient-based Transferable Adversarial Examples by Refining the Integration Path

  • Yuchen Ren
  • Zhengyu Zhao
  • Chenhao Lin
  • Bo Yang
  • Lu Zhou
  • Zhe Liu
  • Chao Shen

Transferable adversarial examples are known to cause threats in practical, black-box attack scenarios. A notable approach to improving transferability is using integrated gradients (IG), originally developed for model interpretability. In this paper, we find that existing IG-based attacks have limited transferability due to their naive adoption of IG in model interpretability. To address this limitation, we focus on the IG integration path and refine it in three aspects: multiplicity, monotonicity, and diversity, supported by theoretical analyses. We propose the Multiple Monotonic Diversified Integrated Gradients (MuMoDIG) attack, which can generate highly transferable adversarial examples on different CNN and ViT models and defenses. Experiments validate that MuMoDIG outperforms the latest IG-based attack by up to 37.3% and other state-of-the-art attacks by 8.4%. In general, our study reveals that migrating established techniques to improve transferability may require non-trivial efforts.

IROS Conference 2025 Conference Paper

IoU-Aware Clustering for Anchor Configuration Determination in Efficient Defect Detection

  • Yuhao Zhao
  • Hongxuan Ma
  • Wei Zou
  • Zhe Liu
  • Hu Su
  • Song Liu 0003

Deep-learning-based object detection has gained widespread application in surface defect inspection, with anchor-based detectors achieving remarkable success by utilizing dense anchors to align with defects. Determining the optimal anchor configuration, i. e. , sizes and aspect ratios of anchor boxes, remains a critical challenge, particularly when addressing defects with significant shape variations. While previous studies have predominantly focused on developing more efficient network architectures and learning strategies, the problem of anchor configuration determination has not been thoroughly explored. To address this gap, this paper proposes the IoU-Aware Clustering (IAC) algorithm, which autonomously learns suitable anchor configurations by extracting shape priors from diverse defects. IAC takes the training bounding boxes as potential clustering centers and selects a subset that aligns with the shape distribution of the training samples. The algorithm involves only a single hyper-parameter, the anchor number k, making it highly adaptable to various scenarios. Experimental results demonstrate that IAC can effectively generate anchor configurations tailored to defect shapes, significantly improving the mean Average Precision (mAP) by 6. 9% and 14. 4% on two industrial defect datasets with substantial shape variations.

JBHI Journal 2025 Journal Article

LiMT: A Multi-Task Liver Image Benchmark Dataset

  • Zhe Liu
  • Kai Han
  • Siqi Ma
  • Yan Zhu
  • Jun Chen
  • Chongwen Lyu
  • Xinyi Qiu
  • Chengxuan Qian

Computer-aided diagnosis (CAD) technology can assist clinicians in evaluating liver lesions and intervening with treatment in time. Although CAD technology has advanced in recent years, the application scope of existing datasets remains relatively limited, typically supporting only single tasks, which has somewhat constrained the development of CAD technology. To address the above limitation, in this paper, we construct a multi-task liver dataset (LiMT) used for liver and tumor segmentation, multi-label lesion classification, and lesion detection based on arterial phase-enhanced computed tomography (CT), potentially providing an exploratory solution that is able to explore the correlation between tasks and does not need to worry about the heterogeneity between task-specific datasets during training. The dataset includes CT volumes from 150 different cases, comprising four types of liver diseases as well as normal cases. Each volume has been carefully annotated and calibrated by experienced clinicians. This public multi-task dataset may become a valuable resource for the medical imaging research community in the future. In addition, this paper not only provides relevant baseline experimental results but also reviews existing datasets and methods related to liver-related tasks. Our dataset is available at https://drive.google.com/drive/folders/1l9HRK13uaOQTNShf5pwgSz3OTanWjkag? usp=sharing.

AAAI Conference 2025 Conference Paper

Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language Navigation

  • Yiyuan Pan
  • Yunzhe Xu
  • Zhe Liu
  • Hesheng Wang

Humans navigate unfamiliar environments using episodic simulation and episodic memory, which facilitate a deeper understanding of the complex relationships between environments and objects. Developing an imaginative memory system inspired by human mechanisms can enhance the navigation performance of embodied agents in unseen environments. However, existing Vision-and-Language Navigation (VLN) agents lack a memory mechanism of this kind. To address this, we propose a novel architecture that equips agents with a reality-imagination hybrid memory system. This system enables agents to maintain and expand their memory through both imaginative mechanisms and navigation actions. Additionally, we design tailored pre-training tasks to develop the agent's imaginative capabilities. Our agent can imagine high-fidelity RGB images for future scenes, achieving state-of-the-art results in a Success rate weighted by Path Length (SPL).

IROS Conference 2025 Conference Paper

PneuChip: A Compact Pneumatic Controller for Large-scale Soft Artificial Muscles

  • Zheng Wang
  • Zhe Liu
  • Yimo Wang
  • Hongying Zhang

Pneumatic soft actuators are known for their versatility and reliability; however, their control presents a major challenge as systems scale beyond tens of actuators. Traditional rigid pneumatic valves add bulk, weight, and complexity, while most soft valves fail to generate programmable independent output states. We propose PneuChip, a compact pneumatic controller designed for large-scale soft actuators. The PneuChip functions as a two-dimensional array with rows and columns, controlled by m + n input signals to generate ${2^{m + n}} - {2^m} - {2^n} + 2$ distinct output states, enabling programmable control over m × n soft actuators. To validate its effectiveness, we implemented PneuChip in a muscular-skeletal robotic arm comprising 24 Miura-Ori inspired, negative pressure-actuated artificial muscles and a rigid two-link skeleton connected by a ball joint. A 4×6 PneuChip was fabricated and integrated to control the arm’s 3 degree of freedoms (DOFs) motion. Within a 120° rotation range, the robot arm achieved 946 distinct positions with smooth state transitions, paving the way for future applications in trajectory tracking and dexterous manipulation. The compact design and high controllability of PneuChip promise to notably simplify complex pneumatic systems, significantly enhancing the practicality of large-scale soft robots for various applications.

NeurIPS Conference 2025 Conference Paper

Seeing through Uncertainty: Robust Task-Oriented Optimization in Visual Navigation

  • Yiyuan Pan
  • Yunzhe Xu
  • Zhe Liu
  • Hesheng Wang

Visual navigation is a fundamental problem in embodied AI, yet practical deployments demand long-horizon planning capabilities to address multi-objective tasks. A major bottleneck is data scarcity: policies learned from limited data often overfit and fail to generalize OOD. Existing neural network-based agents typically increase architectural complexity that paradoxically become counterproductive in the small-sample regime. This paper introduce NeuRO, a integrated learning-to-optimize framework that tightly couples perception networks with downstream task-level robust optimization. Specifically, NeuRO addresses core difficulties in this integration: (i) it transforms noisy visual predictions under data scarcity into convex uncertainty sets using Partially Input Convex Neural Networks (PICNNs) with conformal calibration, which directly parameterize the optimization constraints; and (ii) it reformulates planning under partial observability as a robust optimization problem, enabling uncertainty-aware policies that transfer across environments. Extensive experiments on both unordered and sequential multi-object navigation tasks demonstrate that NeuRO establishes SoTA performance, particularly in generalization to unseen environments. Our work thus presents a significant advancement for developing robust, generalizable autonomous agents.

AAAI Conference 2025 Conference Paper

Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification

  • Haojian Huang
  • Chuanyu Qin
  • Zhe Liu
  • Kaijing Ma
  • Jin Chen
  • Han Fang
  • Chao Ban
  • Hao Sun

Multi-view classification (MVC) faces inherent challenges due to domain gaps and inconsistencies across different views, often resulting in uncertainties during the fusion process. While Evidential Deep Learning (EDL) has been effective in addressing view uncertainty, existing methods predominantly rely on the Dempster-Shafer combination rule, which is sensitive to conflicting evidence and often neglects the critical role of neighborhood structures within multi-view data. To address these limitations, we propose a Trusted Unified Feature-NEighborhood Dynamics (TUNED) model for robust MVC. This method effectively integrates local and global feature-neighborhood (F-N) structures for robust decision-making. Specifically, we begin by extracting local F-N structures within each view. To further mitigate potential uncertainties and conflicts in multi-view fusion, we employ a selective Markov random field that adaptively manages cross-view neighborhood dependencies. Additionally, we employ a shared parameterized evidence extractor that learns global consensus conditioned on local F-N structures, thereby enhancing the global integration of multi-view features. Experiments on benchmark datasets show that our method improves accuracy and robustness over existing approaches, particularly in scenarios with high uncertainty and conflicting views.

IJCAI Conference 2025 Conference Paper

UniCT Depth: Event-Image Fusion Based Monocular Depth Estimation with Convolution-Compensated ViT Dual SA Block

  • Luoxi Jing
  • Dianxi Shi
  • Zhe Liu
  • Songchang Jin
  • Chunping Qiu
  • Ziteng Qiao
  • Yuxian Li
  • Jianqiang Xia

Depth estimation plays a crucial role in 3D scene understanding and is extensively used in a wide range of vision tasks. Image-based methods struggle in challenging scenarios, while event cameras offer high dynamic range and temporal resolution but face difficulties with sparse data. Combining event and image data provides significant advantages, yet effective integration remains challenging. Existing CNN-based fusion methods struggle with occlusions and depth disparities due to limited receptive fields, while Transformer-based fusion methods often lack deep modality interaction. To address these issues, we propose UniCT Depth, an event-image fusion method that unifies CNNs and Transformers to model local and global features. We propose the Convolution-compensated ViT Dual SA (CcViT-DA) Block, designed for the encoder, which integrates Context Modeling Self-Attention (CMSA) to capture spatial dependencies and Modal Fusion Self-Attention (MFSA) for effective cross-modal fusion. Furthermore, we design the tailored Detail Compensation Convolution (DCC) Block to improve texture details and enhances edge representations. Extensive experiments show that UniCT Depth outperforms existing image, event, and fusion-based monocular depth estimation methods across key metrics.

NeurIPS Conference 2025 Conference Paper

Wonder Wins Ways: Curiosity-Driven Exploration through Multi-Agent Contextual Calibration

  • Yiyuan Pan
  • Zhe Liu
  • Hesheng Wang

Autonomous exploration in complex multi-agent reinforcement learning (MARL) with sparse rewards critically depends on providing agents with effective intrinsic motivation. While artificial curiosity offers a powerful self-supervised signal, it often confuses environmental stochasticity with meaningful novelty. Moreover, existing curiosity mechanisms exhibit a uniform novelty bias, treating all unexpected observations equally. However, peer behavior novelty, which encode latent task dynamics, are often overlooked, resulting in suboptimal exploration in decentralized, communication-free MARL settings. To this end, inspired by how human children adaptively calibrate their own exploratory behaviors via observing peers, we propose a novel approach to enhance multi-agent exploration. We introduce CERMIC, a principled framework that empowers agents to robustly filter noisy surprise signals and guide exploration by dynamically calibrating their intrinsic curiosity with inferred multi-agent context. Additionally, CERMIC generates theoretically-grounded intrinsic rewards, encouraging agents to explore state transitions with high information gain. We evaluate CERMIC on benchmark suites including VMAS, Meltingpot, and SMACv2. Empirical results demonstrate that exploration with CERMIC significantly outperforms SoTA algorithms in sparse-reward environments.

NeurIPS Conference 2024 Conference Paper

A Huber Loss Minimization Approach to Mean Estimation under User-level Differential Privacy

  • Puning Zhao
  • Lifeng Lai
  • Li Shen
  • Qingming Li
  • Jiafei Wu
  • Zhe Liu

Privacy protection of users' entire contribution of samples is important in distributed systems. The most effective approach is the two-stage scheme, which finds a small interval first and then gets a refined estimate by clipping samples into the interval. However, the clipping operation induces bias, which is serious if the sample distribution is heavy-tailed. Besides, users with large local sample sizes can make the sensitivity much larger, thus the method is not suitable for imbalanced users. Motivated by these challenges, we propose a Huber loss minimization approach to mean estimation under user-level differential privacy. The connecting points of Huber loss can be adaptively adjusted to deal with imbalanced users. Moreover, it avoids the clipping operation, thus significantly reducing the bias compared with the two-stage approach. We provide a theoretical analysis of our approach, which gives the noise strength needed for privacy protection, as well as the bound of mean squared error. The result shows that the new method is much less sensitive to the imbalance of user-wise sample sizes and the tail of sample distributions. Finally, we perform numerical experiments to validate our theoretical analysis.

AAAI Conference 2024 Conference Paper

Attribute-Missing Graph Clustering Network

  • Wenxuan Tu
  • Renxiang Guan
  • Sihang Zhou
  • Chuan Ma
  • Xin Peng
  • Zhiping Cai
  • Zhe Liu
  • Jieren Cheng

Deep clustering with attribute-missing graphs, where only a subset of nodes possesses complete attributes while those of others are missing, is an important yet challenging topic in various practical applications. It has become a prevalent learning paradigm in existing studies to perform data imputation first and subsequently conduct clustering using the imputed information. However, these ``two-stage" methods disconnect the clustering and imputation processes, preventing the model from effectively learning clustering-friendly graph embedding. Furthermore, they are not tailored for clustering tasks, leading to inferior clustering results. To solve these issues, we propose a novel Attribute-Missing Graph Clustering (AMGC) method to alternately promote clustering and imputation in a unified framework, where we iteratively produce the clustering-enhanced nearest neighbor information to conduct the data imputation process and utilize the imputed information to implicitly refine the clustering distribution through model optimization. Specifically, in the imputation step, we take the learned clustering information as imputation prompts to help each attribute-missing sample gather highly correlated features within its clusters for data completion, such that the intra-class compactness can be improved. Moreover, to support reliable clustering, we maximize inter-class separability by conducting cost-efficient dual non-contrastive learning over the imputed latent features, which in turn promotes greater graph encoding capability for clustering sub-network. Extensive experiments on five datasets have verified the superiority of AMGC against competitors.

AAAI Conference 2024 Conference Paper

DVSAI: Diverse View-Shared Anchors Based Incomplete Multi-View Clustering

  • Shengju Yu
  • Siwei Wang
  • Pei Zhang
  • Miao Wang
  • Ziming Wang
  • Zhe Liu
  • Liming Fang
  • En Zhu

In numerous real-world applications, it is quite common that sample information is partially available for some views due to machine breakdown or sensor failure, causing the problem of incomplete multi-view clustering (IMVC). While several IMVC approaches using view-shared anchors have successfully achieved pleasing performance improvement, (1) they generally construct anchors with only one dimension, which could deteriorate the multi-view diversity, bringing about serious information loss; (2) the constructed anchors are typically with a single size, which could not sufficiently characterize the distribution of the whole samples, leading to limited clustering performance. For generating view-shared anchors with multi-dimension and multi-size for IMVC, we design a novel framework called Diverse View-Shared Anchors based Incomplete multi-view clustering (DVSAI). Concretely, we associate each partial view with several potential spaces. In each space, we enable anchors to communicate among views and generate the view-shared anchors with space-specific dimension and size. Consequently, spaces with various scales make the generated view-shared anchors enjoy diverse dimensions and sizes. Subsequently, we devise an integration scheme with linear computational and memory expenditures to integrate the outputted multi-scale unified anchor graphs such that running spectral algorithm generates the spectral embedding. Afterwards, we theoretically demonstrate that DVSAI owns linear time and space costs, thus well-suited for tackling large-size datasets. Finally, comprehensive experiments confirm the effectiveness and advantages of DVSAI.

EAAI Journal 2024 Journal Article

Fermatean fuzzy similarity measures based on Tanimoto and Sørensen coefficients with applications to pattern classification, medical diagnosis and clustering analysis

  • Zhe Liu

Fermatean fuzzy sets (FFSs) have emerged as a powerful tool for handling uncertain information and have been successfully applied in various domains. However, the existing similarity measures for FFSs often face challenges in accurately capturing the similarity or difference between FFSs, leading to counterintuitive results. To address this issue, we propose sixteen new similarity measures for FFSs inspired by the Tanimoto and Sørensen coefficients. We analyze the properties of these measures and demonstrate their effectiveness through comparative examples with existing measures for FFSs, showing that our proposed measures outperform them in processing fuzzy information derived from FFSs. Additionally, we apply the proposed similarity measures to pattern recognition, medical diagnosis and clustering analysis to demonstrate their practical applicability. The results indicate that the proposed similarity measures can yield more meaningful and effective outcomes.

AAAI Conference 2024 Conference Paper

Hawkes-Enhanced Spatial-Temporal Hypergraph Contrastive Learning Based on Criminal Correlations

  • Ke Liang
  • Sihang Zhou
  • Meng Liu
  • Yue Liu
  • Wenxuan Tu
  • Yi Zhang
  • Liming Fang
  • Zhe Liu

Crime prediction is a crucial yet challenging task within urban computing, which benefits public safety and resource optimization. Over the years, various models have been proposed, and spatial-temporal hypergraph learning models have recently shown outstanding performances. However, three correlations underlying crime are ignored, thus hindering the performance of previous models. Specifically, there are two spatial correlations and one temporal correlation, i.e., (1) co-occurrence of different types of crimes (type spatial correlation), (2) the closer to the crime center, the more dangerous it is around the neighborhood area (neighbor spatial correlation), and (3) the closer between two timestamps, the more relevant events are (hawkes temporal correlation). To this end, we propose Hawkes-enhanced Spatial-Temporal Hypergraph Contrastive Learning framework (HCL), which mines the aforementioned correlations via two specific strategies. Concretely, contrastive learning strategies are designed for two spatial correlations, and hawkes process modeling is adopted for temporal correlations. Extensive experiments demonstrate the promising capacities of HCL from four aspects, i.e., superiority, transferability, effectiveness, and sensitivity.

NeurIPS Conference 2024 Conference Paper

LION: Linear Group RNN for 3D Object Detection in Point Clouds

  • Zhe Liu
  • Jinghua Hou
  • Xinyu Wang
  • Xiaoqing Ye
  • Jingdong Wang
  • Hengshuang Zhao
  • Xiang Bai

The benefit of transformers in large-scale 3D point cloud perception tasks, such as 3D object detection, is limited by their quadratic computation cost when modeling long-range relationships. In contrast, linear RNNs have low computational complexity and are suitable for long-range modeling. Toward this goal, we propose a simple and effective window-based framework built on Linear group RNN (i. e. , perform linear RNN for grouped features) for accurate 3D object detection, called LION. The key property is to allow sufficient feature interaction in a much larger group than transformer-based methods. However, effectively applying linear group RNN to 3D object detection in highly sparse point clouds is not trivial due to its limitation in handling spatial modeling. To tackle this problem, we simply introduce a 3D spatial feature descriptor and integrate it into the linear group RNN operators to enhance their spatial features rather than blindly increasing the number of scanning orders for voxel features. To further address the challenge in highly sparse point clouds, we propose a 3D voxel generation strategy to densify foreground features thanks to linear group RNN as a natural property of auto-regressive models. Extensive experiments verify the effectiveness of the proposed components and the generalization of our LION on different linear group RNN operators including Mamba, RWKV, and RetNet. Furthermore, it is worth mentioning that our LION-Mamba achieves state-of-the-art on Waymo, nuScenes, Argoverse V2, and ONCE datasets. Last but not least, our method supports kinds of advanced linear RNN operators (e. g. , RetNet, RWKV, Mamba, xLSTM and TTT) on small but popular KITTI dataset for a quick experience with our linear RNN-based framework.

EAAI Journal 2024 Journal Article

Novel α -divergence measures on picture fuzzy sets and interval-valued picture fuzzy sets with diverse applications

  • Sijia Zhu
  • Zhe Liu
  • Gözde Ulutagay
  • Muhammet Deveci
  • Dragan Pamučar

Currently, many studies have developed distance or divergence measures between intuitionistic fuzzy sets (IFSs) and interval-valued fuzzy sets (IvFSs). As a generalization of IFSs, picture fuzzy sets (PFSs) provide a more nuanced representation of uncertain and ambiguous information. Interval-valued picture fuzzy sets (IvPFSs) combine the concepts of IvIFSs and PFSs, providing a highly effective means of representing and processing uncertain, ambiguous and incomplete information. How to better measure the differences between PFSs and IvPFSs is still an open issue. This paper proposes some novel α -divergence measures for PFSs and IvPFSs, respectively. We demonstrate the basic properties of the proposed divergence measures, including non-negativity, non-degeneracy and symmetry. Besides, we analyze some special cases of the proposed divergence measures that degenerate into or are related to several well-known divergences. Then, we construct some numerical examples to demonstrate the effectiveness of the proposed measures concerning existing measures. Finally, the proposed α -divergence measures are applied to pattern recognition, multi-attribute decision-making (MADM) and clustering, demonstrating that these measures possess a high confidence level and can produce trustworthy results, especially in comparable situations.

AAAI Conference 2023 Conference Paper

Auto-Weighted Multi-View Clustering for Large-Scale Data

  • Xinhang Wan
  • Xinwang Liu
  • Jiyuan Liu
  • Siwei Wang
  • Yi Wen
  • Weixuan Liang
  • En Zhu
  • Zhe Liu

Multi-view clustering has gained broad attention owing to its capacity to exploit complementary information across multiple data views. Although existing methods demonstrate delightful clustering performance, most of them are of high time complexity and cannot handle large-scale data. Matrix factorization-based models are a representative of solving this problem. However, they assume that the views share a dimension-fixed consensus coefficient matrix and view-specific base matrices, limiting their representability. Moreover, a series of large-scale algorithms that bear one or more hyperparameters are impractical in real-world applications. To address the two issues, we propose an auto-weighted multi-view clustering (AWMVC) algorithm. Specifically, AWMVC first learns coefficient matrices from corresponding base matrices of different dimensions, then fuses them to obtain an optimal consensus matrix. By mapping original features into distinctive low-dimensional spaces, we can attain more comprehensive knowledge, thus obtaining better clustering results. Moreover, we design a six-step alternative optimization algorithm proven to be convergent theoretically. Also, AWMVC shows excellent performance on various benchmark datasets compared with existing ones. The code of AWMVC is publicly available at https://github.com/wanxinhang/AAAI-2023-AWMVC.

JBHI Journal 2023 Journal Article

Disentangled and Side-Aware Unsupervised Domain Adaptation for Cross-Dataset Subjective Tinnitus Diagnosis

  • Yun Li
  • Zhe Liu
  • Lina Yao
  • Jessica J. M. Monaghan
  • David McAlpine

EEG-based tinnitus classification is a valuable tool for tinnitus diagnosis, research, and treatments. Most current works are limited to a single dataset where data patterns are similar. But EEG signals are highly non-stationary, resulting in model's poor generalization to new users, sessions or datasets. Thus, designing a model that can generalize to new datasets is beneficial and indispensable. To mitigate distribution discrepancy across datasets, we propose to achieve Disentangled and Side-aware Unsupervised Domain Adaptation (DSUDA) for cross-dataset tinnitus diagnosis. A disentangled auto-encoder is developed to decouple class-irrelevant information from the EEG signals to improve the classifying ability. The side-aware unsupervised domain adaptation module adapts the class-irrelevant information as domain variance to a new dataset and excludes the variance to obtain the class-distill features for the new dataset classification. It also aligns signals of left and right ears to overcome inherent EEG pattern difference. We compare DSUDA with state-of-the-art methods, and our model achieves significant improvements over competitors regarding comprehensive evaluation criteria. The results demonstrate our model can successfully generalize to a new dataset and effectively diagnose tinnitus.

AAAI Conference 2023 Conference Paper

Let the Data Choose: Flexible and Diverse Anchor Graph Fusion for Scalable Multi-View Clustering

  • Pei Zhang
  • Siwei Wang
  • Liang Li
  • Changwang Zhang
  • Xinwang Liu
  • En Zhu
  • Zhe Liu
  • Lu Zhou

In the past few years, numerous multi-view graph clustering algorithms have been proposed to enhance the clustering performance by exploring information from multiple views. Despite the superior performance, the high time and space expenditures limit their scalability. Accordingly, anchor graph learning has been introduced to alleviate the computational complexity. However, existing approaches can be further improved by the following considerations: (i) Existing anchor-based methods share the same number of anchors across views. This strategy violates the diversity and flexibility of multi-view data distribution. (ii) Searching for the optimal anchor number within hyper-parameters takes much extra tuning time, which makes existing methods impractical. (iii) How to flexibly fuse multi-view anchor graphs of diverse sizes has not been well explored in existing literature. To address the above issues, we propose a novel anchor-based method termed Flexible and Diverse Anchor Graph Fusion for Scalable Multi-view Clustering (FDAGF) in this paper. Instead of manually tuning optimal anchor with massive hyper-parameters, we propose to optimize the contribution weights of a group of pre-defined anchor numbers to avoid extra time expenditure among views. Most importantly, we propose a novel hybrid fusion strategy for multi-size anchor graphs with theoretical proof, which allows flexible and diverse anchor graph fusion. Then, an efficient linear optimization algorithm is proposed to solve the resultant problem. Comprehensive experimental results demonstrate the effectiveness and efficiency of our proposed framework. The source code is available at https://github.com/Jeaninezpp/FDAGF.

NeurIPS Conference 2023 Conference Paper

Query-based Temporal Fusion with Explicit Motion for 3D Object Detection

  • Jinghua Hou
  • Zhe Liu
  • Dingkang Liang
  • Zhikang Zou
  • Xiaoqing Ye
  • Xiang Bai

Effectively utilizing temporal information to improve 3D detection performance is vital for autonomous driving vehicles. Existing methods either conduct temporal fusion based on the dense BEV features or sparse 3D proposal features. However, the former does not pay more attention to foreground objects, leading to more computation costs and sub-optimal performance. The latter implements time-consuming operations to generate sparse 3D proposal features, and the performance is limited by the quality of 3D proposals. In this paper, we propose a simple and effective Query-based Temporal Fusion Network (QTNet). The main idea is to exploit the object queries in previous frames to enhance the representation of current object queries by the proposed Motion-guided Temporal Modeling (MTM) module, which utilizes the spatial position information of object queries along the temporal dimension to construct their relevance between adjacent frames reliably. Experimental results show our proposed QTNet outperforms BEV-based or proposal-based manners on the nuScenes dataset. Besides, the MTM is a plug-and-play module, which can be integrated into some advanced LiDAR-only or multi-modality 3D detectors and even brings new SOTA performance with negligible computation cost and latency on the nuScenes dataset. These experiments powerfully illustrate the superiority and generalization of our method. The code is available at https: //github. com/AlmoonYsl/QTNet.

AAAI Conference 2023 Conference Paper

StereoDistill: Pick the Cream from LiDAR for Distilling Stereo-Based 3D Object Detection

  • Zhe Liu
  • Xiaoqing Ye
  • Xiao Tan
  • Errui Ding
  • Xiang Bai

In this paper, we propose a cross-modal distillation method named StereoDistill to narrow the gap between the stereo and LiDAR-based approaches via distilling the stereo detectors from the superior LiDAR model at the response level, which is usually overlooked in 3D object detection distillation. The key designs of StereoDistill are: the X-component Guided Distillation~(XGD) for regression and the Cross-anchor Logit Distillation~(CLD) for classification. In XGD, instead of empirically adopting a threshold to select the high-quality teacher predictions as soft targets, we decompose the predicted 3D box into sub-components and retain the corresponding part for distillation if the teacher component pilot is consistent with ground truth to largely boost the number of positive predictions and alleviate the mimicking difficulty of the student model. For CLD, we aggregate the probability distribution of all anchors at the same position to encourage the highest probability anchor rather than individually distill the distribution at the anchor level. Finally, our StereoDistill achieves state-of-the-art results for stereo-based 3D detection on the KITTI test benchmark and extensive experiments on KITTI and Argoverse Dataset validate the effectiveness.

AAAI Conference 2023 Conference Paper

TransLO: A Window-Based Masked Point Transformer Framework for Large-Scale LiDAR Odometry

  • Jiuming Liu
  • Guangming Wang
  • Chaokang Jiang
  • Zhe Liu
  • Hesheng Wang

Recently, transformer architecture has gained great success in the computer vision community, such as image classification, object detection, etc. Nonetheless, its application for 3D vision remains to be explored, given that point cloud is inherently sparse, irregular, and unordered. Furthermore, existing point transformer frameworks usually feed raw point cloud of N×3 dimension into transformers, which limits the point processing scale because of their quadratic computational costs to the input size N. In this paper, we rethink the structure of point transformer. Instead of directly applying transformer to points, our network (TransLO) can process tens of thousands of points simultaneously by projecting points onto a 2D surface and then feeding them into a local transformer with linear complexity. Specifically, it is mainly composed of two components: Window-based Masked transformer with Self Attention (WMSA) to capture long-range dependencies; Masked Cross-Frame Attention (MCFA) to associate two frames and predict pose estimation. To deal with the sparsity issue of point cloud, we propose a binary mask to remove invalid and dynamic points. To our knowledge, this is the first transformer-based LiDAR odometry network. The experiment results on the KITTI odometry dataset show that our average rotation and translation RMSE achieves 0.500°/100m and 0.993% respectively. The performance of our network surpasses all recent learning-based methods and even outperforms LOAM on most evaluation sequences.Codes will be released on https://github.com/IRMVLab/TransLO.

JBHI Journal 2022 Journal Article

An Effective Semi-Supervised Approach for Liver CT Image Segmentation

  • Kai Han
  • Lu Liu
  • Yuqing Song
  • Yi Liu
  • Chengjian Qiu
  • Yangyang Tang
  • Qiaoying Teng
  • Zhe Liu

Despite the substantial progress made by deep networks in the field of medical image segmentation, they generally require sufficient pixel-level annotated data for training. The scale of training data remains to be the main bottleneck to obtain a better deep segmentation model. Semi-supervised learning is an effective approach that alleviates the dependence on labeled data. However, most existing semi-supervised image segmentation methods usually do not generate high-quality pseudo labels to expand training dataset. In this paper, we propose a deep semi-supervised approach for liver CT image segmentation by expanding pseudo-labeling algorithm under the very low annotated-data paradigm. Specifically, the output features of labeled images from the pretrained network combine with corresponding pixel-level annotations to produce class representations according to the mean operation. Then pseudo labels of unlabeled images are generated by calculating the distances between unlabeled feature vectors and each class representation. To further improve the quality of pseudo labels, we adopt a series of operations to optimize pseudo labels. A more accurate segmentation network is obtained by expanding the training dataset and adjusting the contributions between supervised and unsupervised loss. Besides, the novel random patch based on prior locations is introduced for unlabeled images in the training procedure. Extensive experiments show our method has achieved more competitive results compared with other semi-supervised methods when fewer labeled slices of LiTS dataset are available.

NeurIPS Conference 2022 Conference Paper

Hilbert Distillation for Cross-Dimensionality Networks

  • Dian Qin
  • Haishuai Wang
  • Zhe Liu
  • HONGJIA XU
  • Sheng Zhou
  • Jiajun Bu

3D convolutional neural networks have revealed superior performance in processing volumetric data such as video and medical imaging. However, the competitive performance by leveraging 3D networks results in huge computational costs, which are far beyond that of 2D networks. In this paper, we propose a novel Hilbert curve-based cross-dimensionality distillation approach that facilitates the knowledge of 3D networks to improve the performance of 2D networks. The proposed Hilbert Distillation (HD) method preserves the structural information via the Hilbert curve, which maps high-dimensional (>=2) representations to one-dimensional continuous space-filling curves. Since the distilled 2D networks are supervised by the curves converted from dimensionally heterogeneous 3D features, the 2D networks are given an informative view in terms of learning structural information embedded in well-trained high-dimensional representations. We further propose a Variable-length Hilbert Distillation (VHD) method to dynamically shorten the walking stride of the Hilbert curve in activation feature areas and lengthen the stride in context feature areas, forcing the 2D networks to pay more attention to learning from activation features. The proposed algorithm outperforms the current state-of-the-art distillation techniques adapted to cross-dimensionality distillation on two classification tasks. Moreover, the distilled 2D networks by the proposed method achieve competitive performance with the original 3D networks, indicating the lightweight distilled 2D networks could potentially be the substitution of cumbersome 3D networks in the real-world scenario.

AAMAS Conference 2022 Conference Paper

The Holy Grail of Multi-Robot Planning: Learning to Generate Online-Scalable Solutions from Offline-Optimal Experts

  • Amanda Prorok
  • Jan Blumenkamp
  • Qingbiao Li
  • Ryan Kortvelesy
  • Zhe Liu
  • Ethan Stump

Many multi-robot planning problems are burdened by the curse of dimensionality, which compounds the difficulty of applying solutions to large-scale problem instances. The use of learning-based methods in multi-robot planning holds great promise as it enables us to offload the online computational burden of expensive centralized, yet optimal solvers, to an offline learning procedure. The hope is that by training a policy to copy an optimal pattern generated by a small-scale (centralized) system, we can transfer that policy to much larger, decentralized systems while maintaining near-optimal performance. Yet, a number of issues impede us from leveraging this idea to its full potential. This blue-sky paper elaborates some of the key challenges that remain.

EAAI Journal 2021 Journal Article

MLANet: Multi-Layer Anchor-free Network for generic lesion detection

  • Zhe Liu
  • Xi Xie
  • Yuqing Song
  • Yang Zhang
  • Xuesheng Liu
  • Jiawen Zhang
  • Victor S. Sheng

In medical image processing, detecting lesions from computed tomography (CT) scans becomes an important research problem with increasing attention. However, this problem is nontrivial because lesions from different organs and parts reflect different characteristics as well as different sizes. Most conventional methods only use a single-scale architecture to detect lesion areas. To get rid of the drawbacks above in medical imaging, a multi-scale framework called MLANet is proposed. To deal with the scale imbalance problem, we design a new backbone—a mixed hourglass network, in which each hourglass module share different input sizes and orders to extract features from different scales. And then the information is sent to the proposed Strengthen Weighted Feature Pyramid Network (SWFPN), a multi-layer weighted feature fusion module, to combine more semantic and spatial information, especially for the case where the number of layers is small. Finally, a Center-to-Corner (C2C) transformation is proposed to deal with the inaccurate size prediction of lesions. It is a non-linear transformation function, aiming to make the predictions more stable and accurate. MLANet is an end-to-end network and is easy to train. In our experiment, it achieves 65. 2% AP50, as well as 88. 3% in the sensitivity of FPs@4. 0 on the DeepLesion dataset, which exceeds many state-of-the-art detectors.

AAAI Conference 2021 Conference Paper

Task Aligned Generative Meta-learning for Zero-shot Learning

  • Zhe Liu
  • Yun Li
  • Lina Yao
  • Xianzhi Wang
  • Guodong Long

Zero-shot learning (ZSL) refers to the problem of learning to classify instances from novel classes (unseen) that are absent in the training set (seen). Most ZSL methods infer the correlation between visual features and attributes to train the classifier for unseen classes. They may have a strong bias towards seen classes during training. Meta-learning has been introduced to mitigate the basis, but meta-ZSL methods are inapplicable when tasks used for training are sampled from diverse distributions. In this regard, we propose a novel Task-aligned Generative Meta-learning model for Zeroshot learning (TGMZ), aiming to mitigate the potentially biased training and to enable meta-ZSL to accommodate realworld datasets that contain diverse distributions. Specifically, TGMZ incorporates an attribute-conditioned task-wise distribution alignment network that projects tasks into a unified distribution to deliver an unbiased model. Our experiments show TGMZ achieves a relative improvement of 2. 1%, 3. 0%, 2. 5%, and 7. 6% over state-of-the-art algorithms on AWA1, AWA2, CUB, and aPY datasets, respectively. Overall, TGMZ outperforms competitors by 3. 6% in the generalized zero-shot learning (GZSL) setting and 7. 9% in our proposed fusion- ZSL setting.

JBHI Journal 2020 Journal Article

Adversarial Representation Learning for Robust Patient-Independent Epileptic Seizure Detection

  • Xiang Zhang
  • Lina Yao
  • Manqing Dong
  • Zhe Liu
  • Yu Zhang
  • Yong Li

Epilepsy is a chronic neurological disorder characterized by the occurrence of spontaneous seizures, which affects about one percent of the worlds population. Most of the current seizure detection approaches strongly rely on patient history records and thus fail in the patient-independent situation of detecting the new patients. To overcome such limitation, we propose a robust and explainable epileptic seizure detection model that effectively learns from seizure states while eliminates the inter-patient noises. A complex deep neural network model is proposed to learn the pure seizure-specific representation from the raw non-invasive electroencephalography (EEG) signals through adversarial training. Furthermore, to enhance the explainability, we develop an attention mechanism to automatically learn the importance of each EEG channels in the seizure diagnosis procedure. The proposed approach is evaluated over the Temple University Hospital EEG (TUH EEG) database. The experimental results illustrate that our model outperforms the competitive state-of-the-art baselines with low latency. Moreover, the designed attention mechanism is demonstrated ables to provide fine-grained information for pathological analysis. We propose an effective and efficient patient-independent diagnosis approach of epileptic seizure based on raw EEG signals without manually feature engineering, which is a step toward the development of large-scale deployment for real-life use.

YNIMG Journal 2020 Journal Article

Fidelity imposed network edit (FINE) for solving ill-posed image reconstruction

  • Jinwei Zhang
  • Zhe Liu
  • Shun Zhang
  • Hang Zhang
  • Pascal Spincemaille
  • Thanh D. Nguyen
  • Mert R. Sabuncu
  • Yi Wang

Deep learning (DL) is increasingly used to solve ill-posed inverse problems in medical imaging, such as reconstruction from noisy and/or incomplete data, as DL offers advantages over conventional methods that rely on explicit image features and hand engineered priors. However, supervised DL-based methods may achieve poor performance when the test data deviates from the training data, for example, when it has pathologies not encountered in the training data. Furthermore, DL-based image reconstructions do not always incorporate the underlying forward physical model, which may improve performance. Therefore, in this work we introduce a novel approach, called fidelity imposed network edit (FINE), which modifies the weights of a pre-trained reconstruction network for each case in the testing dataset. This is achieved by minimizing an unsupervised fidelity loss function that is based on the forward physical model. FINE is applied to two important inverse problems in neuroimaging: quantitative susceptibility mapping (QSM) and under-sampled image reconstruction in MRI. Our experiments demonstrate that FINE can improve reconstruction accuracy.

AAAI Conference 2020 Conference Paper

TANet: Robust 3D Object Detection from Point Clouds with Triple Attention

  • Zhe Liu
  • Xin Zhao
  • Tengteng Huang
  • Ruolan Hu
  • Yu Zhou
  • Xiang Bai

In this paper, we focus on exploring the robustness of the 3D object detection in point clouds, which has been rarely discussed in existing approaches. We observe two crucial phenomena: 1) the detection accuracy of the hard objects, e. g. , Pedestrians, is unsatisfactory, 2) when adding additional noise points, the performance of existing approaches decreases rapidly. To alleviate these problems, a novel TANet is introduced in this paper, which mainly contains a Triple Attention (TA) module, and a Coarse-to-Fine Regression (CFR) module. By considering the channel-wise, point-wise and voxel-wise attention jointly, the TA module enhances the crucial information of the target while suppresses the unstable cloud points. Besides, the novel stacked TA further exploits the multi-level feature attention. In addition, the CFR module boosts the accuracy of localization without excessive computation cost. Experimental results on the validation set of KITTI dataset demonstrate that, in the challenging noisy cases, i. e. , adding additional random noisy points around each object, the presented approach goes far beyond state-of-theart approaches. Furthermore, for the 3D object detection task of the KITTI benchmark, our approach ranks the first place on Pedestrian class, by using the point clouds as the only input. The running speed is around 29 frames per second.

AAAI Conference 2019 Conference Paper

3D Object Detection Using Scale Invariant and Feature Reweighting Networks

  • Xin Zhao
  • Zhe Liu
  • Ruolan Hu
  • Kaiqi Huang

3D object detection plays an important role in a large number of real-world applications. It requires us to estimate the localizations and the orientations of 3D objects in real scenes. In this paper, we present a new network architecture which focuses on utilizing the front view images and frustum point clouds to generate 3D detection results. On the one hand, a PointSIFT module is utilized to improve the performance of 3D segmentation. It can capture the information from different orientations in space and the robustness to different scale shapes. On the other hand, our network obtains the useful features and suppresses the features with less information by a SENet module. This module reweights channel features and estimates the 3D bounding boxes more effectively. Our method is evaluated on both KITTI dataset for outdoor scenes and SUN-RGBD dataset for indoor scenes. The experimental results illustrate that our method achieves better performance than the state-of-the-art methods especially when point clouds are highly sparse.

NeurIPS Conference 2014 Conference Paper

Blossom Tree Graphical Models

  • Zhe Liu
  • John Lafferty

We combine the ideas behind trees and Gaussian graphical models to form a new nonparametric family of graphical models. Our approach is to attach nonparanormal blossoms", with arbitrary graphs, to a collection of nonparametric trees. The tree edges are chosen to connect variables that most violate joint Gaussianity. The non-tree edges are partitioned into disjoint groups, and assigned to tree nodes using a nonparametric partial correlation statistic. A nonparanormal blossom is then "grown" for each group using established methods based on the graphical lasso. The result is a factorization with respect to the union of the tree branches and blossoms, defining a high-dimensional joint density that can be efficiently estimated and evaluated on test points. Theoretical properties and experiments with simulated and real data demonstrate the effectiveness of blossom trees. "

v2026.09.13