Arrow Research search

Author name cluster

Lei Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

176 papers
2 author rows

Possible papers

176

EAAI Journal 2026 Journal Article

A Temporal-spatial Causal Variational Network for accurate sintering temperature forecasting in rotary kilns

  • Kai Wang
  • Hua Chen
  • Xiaogang Zhang
  • Qianyu Chen
  • Yuqi Cai
  • Lei Zhang

Accurate forecasting of sintering temperatures (ST) is pivotal to the high-efficiency, low-energy operation of rotary kilns. The complexity of coupled multivariable process data in industrial environments makes it difficult to uncover patterns and structures in the data, leading to unsatisfactory predictive performance. To accurately analyze temporal–spatial relationships among thermal process variables in rotary kilns, we analyze the causal association among variables and construct a causal graph of the sintering process according to the physicochemical mechanism of sintering. An autoregressive Causal Hidden Markov Model is introduced to model the causal relationships of variables and propagates to generate ST forecasting. In implementation, a generative recurrent neural network, Temporal–spatial Causal Variational Network (TCVN) is designed to generate the representation of hidden variables and extract ST-related features robustly. Each time step in TCVN is composed of a Causal Variational Module (CVM) that integrates a Graph Convolutional Network (GCN) with a Variational Autoencoder (VAE) based on the constructed causal graph. The experiments on real-world data demonstrate that the proposed approach effectively improves the forecasting accuracy of ST with horizons of 1, 3, 6, and 12 steps, confirming the superiority of the proposed model. • A TCVN is proposed for accurate sintering temperature forecasting in rotary kilns. • A causal graph according to the mechanism of sintering is designed. • A CVM is designed to learn the hidden variables in the causal graph. • Detailed experiments are conducted to validate the performance of the TCVN.

AAAI Conference 2026 Conference Paper

AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D Generation

  • Xinyue Liang
  • Zhiyuan Ma
  • Lingchen Sun
  • Yanjun Guo
  • Lei Zhang

Single-image-to-3D models typically follow a sequential generation and reconstruction workflow. However, intermediate multi-view images synthesized by pre-trained generation models often lack cross-view consistency (CVC), significantly degrading 3D reconstruction performance. While recent methods attempt to refine CVC by feeding reconstruction results back into the multi-view generator, these approaches struggle with noisy and unstable reconstruction outputs that limit effective CVC improvement. We introduce AlignCVC, a novel framework that fundamentally re-frames single-image-to-3D generation through distribution alignment rather than relying on strict regression losses. Our key insight is to align both generated and reconstructed multi-view distributions toward the ground-truth multi-view distribution, establishing a principled foundation for improved CVC. Observing that generated images exhibit weak CVC while reconstructed images display strong CVC due to explicit rendering, we propose a soft-hard alignment strategy with distinct objectives for generation and reconstruction models. This approach not only enhances generation quality but also dramatically accelerates inference to as few as 4 steps. As a plug-and-play paradigm, our method, namely AlignCVC, seamlessly integrates various combinations of multiview generation models with 3D reconstruction models. Extensive experiments demonstrate the effectiveness and efficiency of AlignCVC for single-image-to-3D generation.

EAAI Journal 2026 Journal Article

An improved domain adaption method for roughness prediction of milling surfaces under variable processes

  • Lei Zhang
  • Rushan Zhang
  • Chao Liu
  • Zhixun Cui
  • Jianhua Liu

The cutting processes of complex products are complicated and various, and the processing data is a small sample. These characteristics lead to the poor generalization and overfitting problems of the roughness prediction model, further resulting in a reduction in prediction accuracy. To address this issue, a domain adaption method combining Multi-Representation Adaptation Network with Deep Residual Shrinkage Network (DRSN-MRAN) is proposed for roughness prediction of milling surfaces under variable processes. Firstly, regarding to the issues of tedious noise reduction and limited feature extraction, the DRSN is proposed to extract underlying feature information. Subsequently, the MRAN is proposed to address the distribution differences of multi-scale features in the source and target domains and obtain an integrated loss function for the prediction model. In the MRAN, the Conditional Maximum Mean Discrepancy (CMMD) based domain adaptor is introduced to construct the domain adaptive loss function and align the feature distributions between the source and target domains in different spaces and the same category. Finally, a multi-process milling experiment is designed and conducted to obtain a small sample of milling roughness dataset, and the proposed method is verified. It is demonstrated that the DRSN-MRAN can effectively extract domain-invariant features between the source and target domains with limited samples and accurately predict the roughness of milling surfaces under variable processes.

AAAI Conference 2026 Conference Paper

BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object Detection

  • Guowen Zhang
  • Chenhang He
  • Liyi Chen
  • Lei Zhang

Integrating LiDAR and camera information in the bird's eye view (BEV) representation has demonstrated its effectiveness in 3D object detection. However, because of the fundamental disparity in geometric accuracy between these sensors, indiscriminate fusion in previous methods often leads to degraded performance. In this paper, we propose BEVDilation, a novel LiDAR-centric framework that prioritizes LiDAR information in the fusion. By formulating image BEV features as implicit guidance rather than naive concatenation, our strategy effectively alleviates the spatial misalignment caused by image depth estimation errors. Furthermore, the image guidance can effectively help the LiDAR-centric paradigm to address the sparsity and semantic limitations of point clouds. Specifically, we propose a Sparse Voxel Dilation Block that mitigates the inherent point sparsity by densifying foreground voxels through image priors. Moreover, we introduce a Semantic-Guided BEV Dilation Block to enhance the LiDAR feature diffusion processing with image semantic guidance and long-range context capture. On the challenging nuScenes benchmark, BEVDilation achieves better performance than state-of-the-art methods while maintaining competitive computational efficiency. Importantly, our LiDAR-centric strategy demonstrates greater robustness to depth noise compared to naive fusion.

CLeaR Conference 2026 Conference Paper

Causal and Active Learning-Based Counterfactual Chest X-ray Generation for Supporting Clinical Decision-Making in Lung Disease

  • Yifei Zhu
  • Greta Mohr
  • Lei Zhang
  • Christopher Sainsbury
  • Feng Dong
  • John D Maclay
  • David J Lowe
  • David Lagnado

Lung diseases such as lung cancer are major contributors to global morbidity, requiring accurate diagnostic decisions for optimal patient outcomes. While deep learning has advanced medical imaging, the lack of causal inference limits its clinical utility. This study proposes a causal generative framework for counterfactual analysis of Chest X-rays, guided by expert model supervision to ensure clinical plausibility. To solve data imbalance and enhance robustness, we introduce a recurrent active learning strategy that utilises "forgetting rates" to select informative samples. Experimental results demonstrate effectiveness improvements of 9. 25% on the MIMIC dataset and 13. 40% on ChestXray8. Furthermore, two-stage human expert evaluations confirm that the model generates highly realistic synthetic data that maintains a clinical heavy-tailed distribution. These high-quality counterfactuals not only improve diagnostic accuracy but also facilitate confidence calibration for clinicians through interpretable evidence. Our findings demonstrate that integrating causal modeling with expert supervision and active learning provides a robust, clinically meaningful tool for pulmonary diagnostic decision-making.

EAAI Journal 2026 Journal Article

Deep reinforcement learning-based dynamic integrated scheduling of automated guided vehicles and yard cranes for container terminal loading operations

  • Yuxuan Zhang
  • Liang Chen
  • Moshi Zhou
  • Xiangyu Bao
  • Funing Jia
  • Changhui Liu
  • Lei Zhang
  • Yu Zheng

Integrated scheduling of automated guided vehicles (AGVs) and yard cranes (YCs) is crucial for enhancing loading efficiency in container terminals. However, most existing integrated scheduling models are deterministic and static, which limits their effectiveness in uncertain environments. This paper attempts to find reactive scheduling policies that respond to actual observed information rather than relying on determined handling and transport times. We model the loading operation with uncertain handling and transport times as a semi-open queuing network (SOQN) known as LO-SOQN. LO-SOQN responds instantly to AGV and YC scheduling decisions, which provides quantitative metrics for optimizing dynamic scheduling policies. We formulate the problem as a Markov decision process (MDP) and propose a deep reinforcement learning (DRL) approach to find near-optimal policies. The proposed DRL approach ensures policy generalization by learning uniform state representation, which allows it to be applied to flexible equipment configurations. A series of simulation experiments evaluates the performance of our approach under both fixed and flexible equipment configurations. Compared to the inventory-based and robust fluid policies, the proposed approach reduces the long-term average turnaround time by 10. 28 %–18. 94 %. Furthermore, the generalization curve indicates the feasibility and effectiveness of training a policy network that generalizes to various equipment configurations.

AAAI Conference 2026 Conference Paper

DeepSenseMoE: Harnessing Power of Time Series Foundation Models for Few-Shot Human Activity Recognition

  • Zenan Fu
  • Dongzhou Cheng
  • Lei Zhang
  • Wenbo Huang
  • Zhenghao Chen
  • Hao Wu

Recent advances in Time Series Foundation Models (TSFMs) have fundamentally revolutionized general time series analysis across domains like finance, retail, weather, and power. However, how to unlock the hidden capacity of general-purpose TSFMs for wearable activity recognition still remains largely unexplored, given severe sensor annotation scarcity and highly heterogeneous sensor data. To address these challenges, we propose DeepSenseMoE—a novel multi-scale convolution-based Mixture of Experts (MoE) module for parameter-efficient fine-tuning of general-purpose TSFMs to sensor-based activity recognition. DeepSenseMoE integrates three key innovations: (1) Multi-scale convolutional experts with different filter sizes responsible for capturing varying sensor contexts; (2) Shared-expert isolation mechanism compressing common activity knowledge into a single shared expert while reducing redundancy among routed experts; and (3) Hierarchical supervised contrastive alignment guiding experts to further learn discriminative activity features. Extensive experiments on three challenging HAR benchmarks demonstrate DeepSenseMoE's superiority, achieving up to 9.5% accuracy gains over state-of-the-art under few-shot and full-supervised settings, with only <1% additional trainable parameters. We hope that this work may establish a solid foundation to accelerate development and deployment of powerful TSFMs in data-scarce wearable activity recognition tasks while reducing the reliance on labeled sensor data.

AAAI Conference 2026 Conference Paper

DualFete: Revisiting Teacher-Student Interactions from a Feedback Perspective for Semi-supervised Medical Image Segmentation

  • Le Yi
  • Wei Huang
  • Lei Zhang
  • Kefu Zhao
  • Yan Wang
  • Zizhou Wang

The teacher-student paradigm has emerged as a canonical framework in semi-supervised learning. When applied to medical image segmentation, the paradigm faces challenges due to inherent image ambiguities, making it particularly vulnerable to erroneous supervision. Crucially, the student's iterative reconfirmation of these errors leads to self-reinforcing bias. While some studies attempt to mitigate this bias, they often rely on external modifications to the conventional teacher-student framework, overlooking its intrinsic potential for error correction. In response, this work introduces a feedback mechanism into the teacher-student framework to counteract error reconfirmations. Here, the student provides feedback on the changes induced by the teacher's pseudo-labels, enabling the teacher to refine these labels accordingly. We specify that this interaction hinges on two key components: the feedback attributor, which designates pseudo-labels triggering the student's update, and the feedback receiver, which determines where to apply this feedback. Building on this, a dual-teacher feedback model is further proposed, which allows more dynamics in the feedback loop and fosters more gains by resolving disagreements through cross-teacher supervision while avoiding consistent errors. Comprehensive evaluations on three medical image benchmarks demonstrate the method's effectiveness in addressing error propagation in semi-supervised medical image segmentation.

AAAI Conference 2026 Conference Paper

Fast Multi-view Consistent 3D Editing with Video Priors

  • Liyi Chen
  • Ruihuang Li
  • Guowen Zhang
  • Pengfei Wang
  • Lei Zhang

Text-driven 3D editing enables user-friendly 3D object or scene editing with text instructions. Due to the lack of multi-view consistency priors, existing methods typically resort to employ 2D generation or editing models to process per-view individually, followed by iterative 2D-3D-2D updating. However, these methods are not only time-consuming but also prone to yielding over-smoothed results, since iterative process averages the different editing signals gathered from different views. In this paper, we propose, an early and pioneering work of generative Video Prior based 3D Editing, ViP3DE in short, to repurpose the temporal consistency priors from pre-trained video generation models to achieve consistent 3D editing within a single forward pass. Our key insight is to condition the video generation model on a single edited view to generate other consistent edited views for 3D updating directly, thereby bypassing iterative editing paradigm. First, 3D updating requires edited views to be paired with specific camera poses. To this end, we propose \textit{motion-preserved noise blending} for the video model to generate edited views at predefined camera poses. In addition, we introduce \textit{geometrically aware denoising} to further enhance multi-view consistency by integrating 3D geometric priors into video models. Extensive experiments demonstrate that our proposed ViP3DE can achieve high-quality 3D editing results even within a single forward pass, significantly outperforming existing methods in both editing quality and editing time cost.

AAAI Conference 2026 Conference Paper

Geometric Correspondence Constrained Pseudo-Label Alignment for Source-Free Domain Adaptive Fundus Image Segmentation

  • Zhouhongyuan Hu
  • Lei Zhang
  • Lituan Wang
  • Zhenwei Zhang
  • Minjuan Zhu
  • Zhenbin Wang

Source-free unsupervised domain adaptation (SF-UDA), which relies only on a pre-trained source model and unlabeled target data, has gained significant attention. Pseudo-labeling, valued for its simplicity and effectiveness, is a key approach in SF-UDA. However, existing methods neglect the consistency priors of anatomical features across samples, leading them fail to revise of high-confidence noise in structurally inconsistent regions, ultimately manifesting as significant discrepancies in pseudo-labeled samples especially in limited source data scenarios. Motivated by this insight, we propose a novel Geometric Correspondence Constrained (GCC) pseudo-labeling framework. GCC first stratifies pseudo-labeled samples into high/low-quality subsets. It then refines low-quality samples by leveraging the anatomical features inherent in high-quality samples while injecting Gaussian perturbation to perturb high-confidence noise towards the decision boundaries. This process effectively mitigates high-confidence noise disruptive effect and preserves critical prior anatomical knowledge, making it particularly powerful for scenarios with limited source data. Experiments on cross-domain fundus image datasets demonstrate that our method achieves state-of-the-art performance.

AAAI Conference 2026 Conference Paper

JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation Promotion

  • Haoyu Wang
  • Lei Zhang
  • Wenrui Liu
  • Dengyang Jiang
  • Wei Wei
  • Chen Ding

Given the inherently costly and time-intensive nature of pixel-level annotation, the generation of synthetic datasets comprising sufficiently diverse synthetic images paired with ground-truth pixel-level annotations has garnered increasing attention recently for training high-performance semantic segmentation models. However, existing methods necessitate to either predict pseudo annotations after image generation or generate images conditioned on manual annotation masks, which incurs image-annotation semantic inconsistency or scalability problem. To migrate both problems with one stone, we present a novel dataset generative diffusion framework for semantic segmentation, termed JoDiffusion. Firstly, given a standard latent diffusion model, JoDiffusion incorporates an independent annotation variational auto-encoder (VAE) network to map annotation masks into the latent space shared by images. Then, the diffusion model is tailored to capture the joint distribution of each image and its annotation mask conditioned on a text prompt. By doing these, JoDiffusion enables simultaneously generating paired images and semantically consistent annotation masks solely conditioned on text prompts, thereby demonstrating superior scalability. Additionally, a mask optimization strategy is developed to mitigate the annotation noise produced during generation. Experiments on Pascal VOC, COCO, and ADE20K datasets show that the annotated dataset generated by JoDiffusion yields substantial performance improvements in semantic segmentation compared to existing methods.

AAAI Conference 2026 Conference Paper

LayerEdit: Disentangled Multi-Object Editing via Conflict-Aware Multi-Layer Learning

  • Fengyi Fu
  • Mengqi Huang
  • Lei Zhang
  • Zhendong Mao

Text-driven multi-object image editing which aims to precisely modify multiple objects within an image based on text descriptions, has recently attracted considerable interest. Existing works primarily follow the localize-editing paradigm, focusing on independent object localization and editing while neglecting critical inter-object interactions. However, this work points out that the neglected attention entanglements in inter-object conflict regions, inherently hinder disentangled multi-object editing, leading to either inter-object editing leakage or intra-object editing constraints. We thereby propose a novel multi-layer disentangled editing framework LayerEdit, a training-free method which, for the first time, through precise object-layered decomposition and coherent fusion, enables conflict-free object-layered editing. Specifically, LayerEdit introduces a novel “decompose-editing-fusion” framework, consisting of: (1) Conflict-aware Layer Decomposition module, which utilizes an attention-aware IoU scheme and time-dependent region removing, to enhance conflict awareness and suppression for layer decomposition. (2) Object-layered Editing module, to establish coordinated intra-layer text guidance and cross-layer geometric mapping, achieving disentangled semantic and structural modifications. (3) Transparency-guided Layer Fusion module, to facilitate structure-coherent inter-object layer fusion through precise transparency guidance learning. Extensive experiments verify the superiority of LayerEdit over existing methods, showing unprecedented intra-object controllability and inter-object coherence in complex multi-object scenarios.

EAAI Journal 2026 Journal Article

Multitasking optimization for personalized exercise group recommendation in E-learning environments

  • Haipeng Yang
  • Sibo Liu
  • Zihao Chen
  • Yuanyuan Ge
  • Lei Zhang

Personalized exercise group recommendation (PEGR) is to select a set of exercises from a large exercise bank for students, which plays an important role in E-learning. Due to the complexity of real application scenarios, PEGR is usually modeled as a large-scale constrained multi-objective optimization problem and solved by multi-objective evolutionary algorithms (MOEAs). However, the “curse of dimensionality” and the complex constraints handling are the two challenges encountered when designing MOEAs to solve the PEGR problem. To this end, we propose a novel evolutionary tri-tasking algorithm named ETT-PEGR to tackle the challenges of solving the PEGR, in which two auxiliary tasks are constructed to help solve the original task through knowledge transfer. Specifically, the first concept-recommended auxiliary task is designed to recommend knowledge concepts instead of exercises to students, which can help accelerate the convergence speed of the original task since the number of concepts is much smaller than that of exercises. The second constraint-ignored auxiliary task is designed to help the solutions of the original task to cross the infeasible region. In addition, a novel knowledge transfer mechanism based on different encoding strategies is proposed for the original task and the two auxiliary tasks, which can effectively realize the knowledge transfer between them. Experimental results on four popular datasets show that ETT-PEGR outperforms the state-of-the-art algorithms for PEGR.

JBHI Journal 2026 Journal Article

Optimizing Accuracy-Efficiency Trade-Offs of On-Device Activity Inference With Star Operation

  • Guangjie Chen
  • Zenan Fu
  • Yetong Sha
  • Di Xiong
  • Lei Zhang
  • Hao Wu
  • Aiguo Song

Lightweight convolution-based neural networks (CNNs) are well suited for sensor-based human activity recognition (HAR) applications on resource-constrained edge devices with faster inference speed. However, the convolutional kernels are often limited to a small window range, which can only capture local details in time series sensor data, thus preventing further performance boost. Though Introducing self-attention into convolution can help to handle long-range dependence well, it might significantly slow down actual activity inference speed, due to high computational cost. In this paper, we introduce a new learning paradigm (star operation) and then present a lightweight Dual-Branch High-Order Interactions (DbHoi) block, which is computationally friendly for mobile HAR deployment. The proposed DbHoi block may implicitly transform raw sensor inputs into high-dimensional non-linear features, but actually operate in a low-dimensional feature space (analogs to the design principle of polynomial kernel tricks), without incurring extra computational overhead. Extensive experiments are conducted on three public HAR benchmarks including UCI-HAR, UniMiB-SHAR, and OPPORTUNITY, which demonstrate that our suggested DbHoi can consistently surpass various meticulously designed lightweight networks such as MobileNet, ShuffleNet, and GhostNet. Detailed ablation studies, visualizing representations, and on-device latency analyses further validate our insights with regards to the star operation, while underscoring its practical merit in real-world HAR deployment.

AAAI Conference 2026 Conference Paper

Otter: Mitigating Background Distractions of Wide-Angle Few-Shot Action Recognition with Enhanced RWKV

  • Wenbo Huang
  • Jinghui Zhang
  • Zhenghao Chen
  • Guang Li
  • Lei Zhang
  • Yang Cao
  • Fang Dong
  • Takahiro Ogawa

Wide-angle videos in few-shot action recognition (FSAR) effectively express actions within specific scenarios. However, without a global understanding of both subjects and background, recognizing actions in such samples remains challenging because of the background distractions. Receptance Weighted Key Value (RWKV), which learns interaction between various dimensions, shows promise for global modeling. While directly applying RWKV to wide-angle FSAR may fail to highlight subjects due to excessive background information. Additionally, temporal relation degraded by frames with similar backgrounds is difficult to reconstruct, further impacting performance. Therefore, we design the CompOund SegmenTation and Temporal REconstructing RWKV (Otter). Specifically, the Compound Segmentation Module (CSM) is devised to segment and emphasize key patches in each frame, effectively highlighting subjects against background information. The Temporal Reconstruction Module (TRM) is incorporated into the temporal-enhanced prototype construction to enable bidirectional scanning, allowing better reconstruct temporal relation. Furthermore, a regular prototype is combined with the temporal-enhanced prototype to simultaneously enhance subject emphasis and temporal modeling, improving wide-angle FSAR performance. Extensive experiments on benchmarks such as SSv2, Kinetics, UCF101, and HMDB51 demonstrate that Otter achieves state-of-the-art performance. Extra evaluation on the VideoBadminton dataset further validates the superiority of Otter in wide-angle FSAR.

AAAI Conference 2026 Conference Paper

Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding

  • Jiaqi Tang
  • Jianmin Chen
  • Wei Wei
  • Xiaogang Xu
  • Runtao Liu
  • Xiangyu Wu
  • Qipeng Xie
  • Jiafei Wu

Multimodal Large Language Models struggle to maintain reliable performance under extreme real-world visual degradations, which impede their practical robustness. Existing robust MLLMs predominantly rely on implicit training/adaptation that focuses solely on visual encoder generalization, suffering from limited interpretability and isolated optimization. To overcome these limitations, we propose Robust-R1, a novel framework that explicitly models visual degradations through structured reasoning chains. Our approach integrates: (i) supervised fine-tuning for degradation-aware reasoning foundations, (ii) reward-driven alignment for accurately perceiving degradation parameters, and (iii) dynamic reasoning depth scaling adapted to degradation intensity. To facilitate this approach, we introduce a specialized 11K dataset featuring realistic degradations synthesized across four critical real-world visual processing stages, each annotated with structured chains connecting degradation parameters, perceptual influence, pristine semantic reasoning chain, and conclusion. Comprehensive evaluations demonstrate state-of-theart robustness: Robust-R1 outperforms all general and robust baselines on the real-world degradation benchmark R-Bench, while maintaining superior anti-degradation performance under multi-intensity adversarial degradations on MMMB, MMStar, and RealWorldQA.

AAAI Conference 2026 Conference Paper

SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features

  • Jinyuan Qu
  • Hongyang Li
  • Xingyu Chen
  • Shilong Liu
  • Yukai Shi
  • Tianhe Ren
  • Ruitao Jing
  • Lei Zhang

In this paper, we present SegDINO3D, a novel Transformer encoder-decoder framework for 3D instance segmentation. As 3D training data is generally not as sufficient as 2D training images, SegDINO3D is designed to fully leverage 2D representation from a pre-trained 2D detection model, including both image-level and object-level features, for improving 3D representation. SegDINO3D takes both a point cloud and its associated 2D images as input. In the encoder stage, it first enriches each 3D point by retrieving 2D image features from its corresponding image views and then leverages a 3D encoder for 3D context fusion. In the decoder stage, it formulates 3D object queries as 3D anchor boxes and performs cross-attention from 3D queries to 2D object queries obtained from 2D images using the 2D detection model. These 2D object queries serve as a compact object-level representation of 2D images, effectively avoiding the challenge of keeping thousands of image feature maps in the memory while faithfully preserving the knowledge of the pre-trained 2D model. The introducing of 3D box queries also enables the model to modulate cross-attention using the predicted boxes for more precise querying. SegDINO3D achieves the state-of-the-art performance on the ScanNetV2 and ScanNet200 3D instance segmentation benchmarks. Notably, on the challenging ScanNet200 dataset, SegDINO3D significantly outperforms prior methods by +8.7 and +6.8 mAP on the validation and hidden test sets, respectively, demonstrating its superiority.

AAAI Conference 2026 Conference Paper

T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object Detection

  • Jiazhou Zhou
  • Qing Jiang
  • Kanghao Chen
  • Lutao Jiang
  • Yuanhuiyi Lyu
  • Ying-Cong Chen
  • Lei Zhang

Object detection methods have evolved from closed-set to open-set paradigms over the years. Current open-set object detectors, however, remain constrained by their exclusive reliance on positive indicators based on given prompts like text descriptions or visual exemplars. This positive-only paradigm experiences consistent vulnerability to visually similar but semantically different distractors. We propose T-Rex-Omni, a novel framework that addresses this limitation by incorporating negative visual prompts to negate hard negative distractors. Specifically, we first introduce a unified visual prompt encoder that jointly processes positive and negative visual prompts. Next, a training-free Negating Negative Computing (NNC) module is proposed to dynamically suppress negative responses during the probability computing stage. To further boost performance through fine-tuning, our Negating Negative Hinge (NNH) loss enforces discriminative margins between positive and negative embeddings. T-Rex-Omni supports flexible deployment in both positive-only and joint positive-negative inference modes, accommodating either user-specified or automatically generated negative examples. Extensive experiments demonstrate remarkable zero-shot detection performance, significantly narrowing the performance gap between visual-prompted and text-prompted methods while showing particular strength in long-tailed scenarios (51.2 AP_r on LVIS-minival). This work establishes negative prompts as a crucial new dimension for advancing open-set visual recognition systems.

EAAI Journal 2026 Journal Article

The accidental explosion tracing model of architecture glass damage based on shuffle attention

  • Hao Liu
  • Zhen Qing Wang
  • Shuai Qin
  • Qiang Zhao
  • Lei Zhang

Explosion tracing is the basis of hazard analysis and risk analysis of accidental explosion. The accidental explosion damage effects data has typical multi-source heterogeneous characteristics. The method of data fusion can be used to aggregate the redundant or complementary information on multiple sensors to obtain the more completed information. The architecture glass damage accidental explosion tracing machine learning model with Shuffle attention based on the experiment and simulate data of tempered glass plate under blast loading has been exhibited. Two branch models, blast wave propagation tracing model (Process-model) and glass plate dynamic response tracing model (Response-model), were established based on multi-source data fusion. Then the final decision tracing model (Decision-model) has been constructed based on the decision level. The Mean Absolute Percentage Error (MAPE) of the Decision-model was reduced to 0. 1605 compared with two branch models. The data of three different positions of the glass plate were used in the tracing process respectively. The results indicated that the data of the peripheral area showed the largest error. Considering the incompleteness of the actual explosion accident investigation data, in order to verify the real-word practical application of the model, an accidental explosion verification test of emulsion explosive was carried out. The MAPE of the actual measured imperfect dataset is 0. 2423. The results show that the tracing model can still ensure its prediction accuracy even when part input data is missing in practical applications. It provides a reliable analysis tool for accident explosion risk assessment.

AAAI Conference 2026 Conference Paper

Towards Better Code Understanding in Decoder-Only Models with Contrastive Learning

  • Jiayi Lin
  • Yanlin Wang
  • Yibiao Yang
  • Lei Zhang
  • Yutao Xie

Recent advances in large-scale code generation models have led to remarkable progress in producing high-quality code. These models are trained in a self-supervised manner on extensive unlabeled code corpora using a decoder-only architecture. However, despite their generative strength, decoder-only models often exhibit limited performance on code understanding tasks such as code search and clone detection, primarily due to their generation-oriented training objectives. While training large encoder-only models from scratch on massive code datasets can improve understanding ability but remains computationally expensive and time-consuming. In this paper, we explore a more efficient alternative by transferring knowledge from pre-trained decoder-only code generation models to code understanding tasks. We investigate how decoder-only architectures can be effectively adapted to learn discriminative and semantically meaningful code representations. To this end, we propose CL4D, a contrastive learning framework tailored to strengthen the representation capabilities of decoder-only models. Extensive experiments on multiple benchmark datasets demonstrate that CL4D achieves competitive or superior performance compared to existing methods on representative code understanding tasks, including code search and clone detection. Further analysis reveals that CL4D substantially improves the semantic alignment of code representations by reducing the distance between semantically similar code snippets. These findings highlight the feasibility of leveraging decoder-only models as a unified backbone for both code generation and understanding.

YNIMG Journal 2026 Journal Article

Transcutaneous auricular vagus nerve stimulation facilitates visuomotor association learning: Behavioral and electrophysiological evidence

  • Long Chen
  • Chenghu Tang
  • Huixin Gao
  • Lei Zhang
  • Shengcui Cheng
  • Zhongpeng Wang
  • Shuang Liu
  • Dong Ming

Associating visual cues with appropriate motor responses is a fundamental adaptive skill. Transcutaneous auricular vagus nerve stimulation (taVNS) may enhance visuomotor association (VMA) learning, though its neural mechanisms remain unclear. Electroencephalogram (EEG), with its millisecond temporal resolution, offers unique advantages for elucidating the neurodynamic of VMA plasticity. This single-blind, sham-controlled, between-subjects study investigated whether taVNS facilitates VMA learning through behavioral and EEG analysis. Participants (each group N = 19) performed a VMA task (associating five oracle pictures with five keyboard keys) before and after 20-min active/sham taVNS. Behavioral results revealed that compared to the sham group, the active group exhibited shorter reaction time, higher response accuracy and larger learning curve integration, confirming the positive effect of taVNS on VMA learning. Neurophysiologically, taVNS reduced the P200 and P300 amplitudes, enhanced N170 negativity and attenuated error-related negativity. Cross-regional-frequency phase-amplitude coupling results demonstrated enhanced synchronization of frontal-parietal-occipital neural cross-frequency activity. Additionally, parietal-occipital θ, α, β band inter-trial phase coherence was enhanced in the active group. These findings demonstrate that taVNS enhances VMA acquisition through optimizing visual and error processing efficiency. This study establishes a neurophysiological basis for taVNS's cognitive enhancement potential, suggesting its utility in rehabilitative paradigms targeting associative learning deficits.

EAAI Journal 2025 Journal Article

A dual-population based two-archive coevolution algorithm for constrained multi-objective optimization problems

  • Miao Chen
  • Shijie Zhao
  • Tianran Zhang
  • Lei Zhang

How to balance the objectives and constraints better is the key to solving constrained multi-objective optimization problems (CMOPs). Many evolutionary algorithms struggle to fully converge to the entire Pareto front, especially in CMOPs with narrow and complex feasible regions, which posing significant challenges in solving CMOPs. To handle this problem, the paper proposes a dual-population based two-archive coevolution algorithm (DPTAC). The main population evolves towards the true Pareto front while accounting for considering the original problem. The auxiliary population ignores the constraints and approximates unconstrained Pareto front. To assist main population in crossing larger infeasible regions, enhancing its diversity, and discovering more feasible regions, a two-archive strategy is proposed, which stores the potentially valuable non-dominated infeasible solutions and non-dominated solutions generated by the evolution of the main population and the auxiliary population respectively. In addition, a removal mechanism is introduced and integrated into the auxiliary population to reduce computational resource waste. This can help the main population have more computational resources in the late stage of evolution to find narrow feasible regions and improve the convergence of the population. Experimental results demonstrate that DPTAC outperforms 9 state-of-the-art constrained multi-objective evolutionary algorithms (CMOEAs) across 5 test suites comprising 62 benchmark functions and 6 real-world problems, confirming its superior competitiveness.

EAAI Journal 2025 Journal Article

A novel multimodal deep learning framework for predicting residual strength of corroded rectangular hollow-section columns

  • Yu-Jia Zhang
  • Lei Zhang
  • Yu Zhou
  • Tian-Xiang Li
  • Reece Lincoln
  • Jing-Zhong Tong
  • Jia-Jia Shen

Corrosion, recognized as a thermodynamically spontaneous process, is one of the key issues affecting the health of rectangular hollow steel section columns under working conditions, and has attracted much attention in recent years. Traditional approaches, such as multi-layer perceptron, often rely solely on the degree of volume loss to predict residual strength, overlooking the spatial complexity of actual corrosion patterns. To address these limitations, this study presents a novel multimodal deep learning network for accurately predicting the residual strength of corroded hollow steel section columns with random, nonuniform corrosion distributions. Our approach integrates (i) image‐based corrosion distributions on four steel walls, and (ii) tabular geometric parameters, through five distinct data-fusion methods proposed in this work, three employing Late Fusion (via a novel multi‐head attention module) and two using Early Fusion (via pixel−level merging). The image information extraction core is built upon a lightweight convolutional neural network and a channel−spatial attention block, while the tabular extraction module leverages a revised multi-layer perceptron architecture. After Bayesian hyperparameter optimization, the best‐performing model achieves a coefficient of determination of 0. 971 on the test set, surpassing conventional machine learning and other multimodal fusion techniques by 0. 01–0. 161. Further analysis shows that the reverse visualization technique highlights corrosion−critical regions that closely coincide with the experimentally validated failure zones. Consequently, the proposed framework not only predicts residual strength with high accuracy but also localizes vulnerable areas for targeted reinforcement. This methodology holds promise for large‐scale corrosion monitoring and structural health assessment of steel infrastructure.

EAAI Journal 2025 Journal Article

A three-stage adaptive memetic algorithm for multi-objective optimization of flexible assembly job-shop scheduling problem

  • Chenlu Zhang
  • Jiamei Feng
  • Mingchuan Zhang
  • Lei Yang
  • Lei Zhang
  • Lin Wang
  • Junlong Zhu
  • Qingtao Wu

The flexible assembly job-shop scheduling problem (FAJSP) widely arises in the manufacturing industry. Various approaches have been designed in recent years to address this problem. However, existing methods have rarely considered assembly process constraints and task assembly wait time. For this reason, this paper proposes a three-stage adaptive memetic algorithm (TA-MA) to solve the FAJSP with process route constraints. Specifically, the proposed algorithm combines memetic algorithms and reinforcement learning. The optimization objectives are completion time, equipment load, and assembly operation waiting time. Moreover, a two-layer integer coding method is proposed to encode the problem, and a reinforcement learning method is introduced to assist the solution search of the memetic algorithm. Further, a three-stage search framework is designed to reasonably equilibrium TA-MA’s exploration and mining capabilities as iterations advance. Finally, the effectiveness of the proposed algorithm is assessed through a series of experiments. The outcomes demonstrate that the proposed algorithm is effective and outperforms existing algorithms.

AAAI Conference 2025 Conference Paper

Adversarial Contrastive Graph Augmentation with Counterfactual Regularization

  • Tao Long
  • Lei Zhang
  • Liang Zhang
  • Laizhong Cui

With the advancement of graph representation learning, self-supervised graph contrastive learning (GCL) has emerged as a key technique in the field. In GCL, positive and negative samples are generated through data augmentation. While recent works have introduced model-based methods to enhance positive graph augmentations, they often overlook the importance of negative samples, relying instead on rule-based methods that can fail to capture meaningful graph patterns. To address this issue, we propose a novel model-based adversarial contrastive graph augmentation (ACGA) method that automatically generates both positive graph samples with minimal sufficient information and hard negative graph samples. Additionally, we provide a theoretical framework to analyze the process of positive and negative graph augmentation in self-supervised GCL. We evaluate our ACGA method through extensive experiments on representative benchmark datasets, and the results demonstrate that ACGA outperforms state-of-the-art baselines.

ICLR Conference 2025 Conference Paper

Autoregressive Pretraining with Mamba in Vision

  • Sucheng Ren
  • Xianhang Li
  • Haoqin Tu
  • Feng Wang
  • Fangxun Shu
  • Lei Zhang
  • Jieru Mei
  • Linjie Yang

The vision community has started to build with the recently developed state space model, Mamba, as the new backbone for a range of tasks. This paper shows that Mamba's visual capability can be significantly enhanced through autoregressive pretraining, a direction not previously explored. Efficiency-wise, the autoregressive nature can well capitalize on the Mamba's unidirectional recurrent structure, enabling faster overall training speed compared to other training strategies like mask modeling. Performance-wise, autoregressive pretraining equips the Mamba architecture with markedly higher accuracy over its supervised-trained counterparts and, more importantly, successfully unlocks its scaling potential to large and even huge model sizes. For example, with autoregressive pretraining, a base-size Mamba attains 83.2\% ImageNet accuracy, outperforming its supervised counterpart by 2.0\%; our huge-size Mamba, the largest Vision Mamba to date, attains 85.0\% ImageNet accuracy (85.5\% when finetuned with $384\times384$ inputs), notably surpassing all other Mamba variants in vision. The code is available at \url{https://github.com/OliverRensu/ARM}.

IJCAI Conference 2025 Conference Paper

Beyond Statistical Analysis: Multimodal Framework for Time Series Forecasting with LLM-Driven Temporal Pattern

  • Jiahong Xiong
  • Chengsen Wang
  • Haifeng Sun
  • Yuhan Jing
  • Qi Qi
  • Zirui Zhuang
  • Lei Zhang
  • Jianxin Liao

Accurate forecasting of time series is crucial for many applications in the real world. Conventional methods primarily rely on statistical analysis of historical data, often leading to overfitting and failing to account for background information and constraints imposed by external events. Therefore, introducing large language models (LLMs) with robust textual capabilities holds significant potential. However, due to the inherent limitations of LLMs in handling numerical data, they do not exhibit advantages in precise numerical prediction tasks. Therefore, we propose a framework to integrate LLMs with conventional methods synergistically. Rather than directly outputting numerical predictions, we leverage the capabilities of the LLMs to generate textual temporal patterns, thereby fully utilizing their inherent knowledge and reasoning abilities. Additionally, we introduce a memory network designed to decode these textual representations into a format that numerical models can effectively interpret. This approach not only capitalizes on the strengths of the LLM in text processing but also bridges the gap between textual and numerical data, enhancing the overall predictive performance of the model. Our experimental results demonstrate the framework's effectiveness, achieving state-of-the-art performance on various benchmark datasets.

NeurIPS Conference 2025 Conference Paper

BurstDeflicker: A Benchmark Dataset for Flicker Removal in Dynamic Scenes

  • Lishen Qu
  • Zhihao Liu
  • Shihao Zhou
  • LUO YAQI
  • Jie Liang
  • Hui Zeng
  • Lei Zhang
  • Jufeng Yang

Flicker artifacts in short-exposure images are caused by the interplay between the row-wise exposure mechanism of rolling shutter cameras and the temporal intensity variations of alternating current (AC)-powered lighting. These artifacts typically appear as uneven brightness distribution across the image, forming noticeable dark bands. Beyond compromising image quality, this structured noise also affects high-level tasks, such as object detection and tracking, where reliable lighting is crucial. Despite the prevalence of flicker, the lack of a large-scale, realistic dataset has been a significant barrier to advancing research in flicker removal. To address this issue, we present BurstDeflicker, a scalable benchmark constructed using three complementary data acquisition strategies. First, we develop a Retinex-based synthesis pipeline that redefines the goal of flicker removal and enables controllable manipulation of key flicker-related attributes (e. g. , intensity, area, and frequency), thereby facilitating the generation of diverse flicker patterns. Second, we capture 4, 000 real-world flicker images from different scenes, which help the model better understand the spatial and temporal characteristics of real flicker artifacts and generalize more effectively to wild scenarios. Finally, due to the non-repeatable nature of dynamic scenes, we propose a green-screen method to incorporate motion into image pairs while preserving real flicker degradation. Comprehensive experiments demonstrate the effectiveness of our dataset and its potential to advance research in flicker removal.

AAAI Conference 2025 Conference Paper

ChatTime: A Unified Multimodal Time Series Foundation Model Bridging Numerical and Textual Data

  • Chengsen Wang
  • Qi Qi
  • Jingyu Wang
  • Haifeng Sun
  • Zirui Zhuang
  • Jinming Wu
  • Lei Zhang
  • Jianxin Liao

Human experts typically integrate numerical and textual multimodal information to analyze time series. However, most traditional deep learning predictors rely solely on unimodal numerical data, using a fixed-length window for training and prediction on a single dataset, and cannot adapt to different scenarios. The powered pre-trained large language model has introduced new opportunities for time series analysis. Yet, existing methods are either inefficient in training, incapable of handling textual information, or lack zero-shot forecasting capability. In this paper, we innovatively model time series as a foreign language and construct ChatTime, a unified framework for time series and text processing. As an out-of-the-box multimodal time series foundation model, ChatTime provides zero-shot forecasting capability and supports bimodal input/output for both time series and text. We design a series of experiments to verify the superior performance of ChatTime across multiple tasks and scenarios, and create four multimodal datasets to address data gaps. The experimental results demonstrate the potential and utility of ChatTime.

AAAI Conference 2025 Conference Paper

CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility

  • Bojia Zi
  • Shihao Zhao
  • Xianbiao Qi
  • Jianan Wang
  • Yukai Shi
  • Qianyu Chen
  • Bin Liang
  • Rong Xiao

Video inpainting is a crucial task with diverse applications, including fine-grained video editing, video recovery, and video dewatermarking. However, most existing video inpainting methods primarily focus on visual content completion while neglecting text information. There are only a limited number of text-guided video inpainting techniques, and these techniques struggle with maintaining visual quality and exhibit poor semantic representation capabilities. In this paper, we introduce CoCoCo, a text-guided video inpainting diffusion framework. To address the aforementioned challenges, we enhance both the training data and model structure. Specifically, we devise an instance-aware region selection strategy for masked area sampling and develop a novel motion block that incorporates efficient 3D full attention and textual cross attention. Additionally, our CoCoCo framework can be seamlessly integrated with various personalized text-to-image diffusion models through a delicate training-free transfer mechanism. Comprehensive experiments demonstrate that CoCoCo can create high-quality visual content with enhanced temporal consistency, improved text controllability, and better compatibility with personalized image models.

AAAI Conference 2025 Conference Paper

CustomContrast: A Multilevel Contrastive Perspective for Subject-Driven Text-to-Image Customization

  • Nan Chen
  • Mengqi Huang
  • Zhuowei Chen
  • Yang Zheng
  • Lei Zhang
  • Zhendong Mao

Subject-driven text-to-image (T2I) customization has drawn significant interest in academia and industry. This task enables pre-trained models to generate novel images based on unique subjects. Existing studies adopt a self-reconstructive perspective, focusing on capturing all details of a single image, which will misconstrue the specific image's irrelevant attributes (e.g., view, pose, and background) as the subject intrinsic attributes. This misconstruction leads to both overfitting or underfitting of irrelevant and intrinsic attributes of the subject, i.e., these attributes are over-represented or under-represented simultaneously, causing a trade-off between similarity and controllability. In this study, we argue an ideal subject representation can be achieved by a cross-differential perspective, i.e., decoupling subject intrinsic attributes from irrelevant attributes via contrastive learning, which allows the model to focus more on intrinsic attributes through intra-consistency (features of the same subject are spatially closer) and inter-distinctiveness (features of different subjects have distinguished differences). Specifically, we propose CustomContrast, a novel framework, which includes a Multilevel Contrastive Learning (MCL) paradigm and a Multimodal Feature Injection (MFI) Encoder. The MCL paradigm is used to extract intrinsic features of subjects from high-level semantics to low-level appearance through crossmodal semantic contrastive learning and multiscale appearance contrastive learning. To facilitate contrastive learning, we introduce the MFI encoder to capture cross-modal representations. Extensive experiments show the effectiveness of CustomContrast in subject similarity and text controllability.

NeurIPS Conference 2025 Conference Paper

DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing

  • Chenxi Xie
  • Minghan Li
  • Shuai Li
  • Yuhui Wu
  • Qiaosi Yi
  • Lei Zhang

Leveraging the powerful generation capability of large-scale pretrained text-to-image models, training-free methods have demonstrated impressive image editing results. Conventional diffusion-based methods, as well as recent rectified flow (RF)-based methods, typically reverse synthesis trajectories by gradually adding noise to clean images, during which the noisy latent at the current timestep is used to approximate that at the next timesteps, introducing accumulated drift and degrading reconstruction accuracy. Considering the fact that in RF the noisy latent is estimated through direct interpolation between Gaussian noises and clean images at each timestep, we propose Direct Noise Alignment (DNA), which directly refines the desired Gaussian noise in the noise domain, significantly reducing the error accumulation in previous methods. Specifically, DNA estimates the velocity field of the interpolated noised latent at each timestep and adjusts the Gaussian noise by computing the difference between the predicted and expected velocity field. We validate the effectiveness of DNA and reveal its relationship with existing RF-based inversion methods. Additionally, we introduce a Mobile Velocity Guidance (MVG) to control the target prompt-guided generation process, balancing image background preservation and target object editability. DNA and MVG collectively constitute our proposed method, namely DNAEdit. Finally, we introduce DNA-Bench, a long-prompt benchmark, to evaluate the performance of advanced image editing models. Experimental results demonstrate that our DNAEdit achieves superior performance to state-of-the-art text-guided editing methods. Our code, model, and benchmark will be made publicly available.

NeurIPS Conference 2025 Conference Paper

Do LVLMs Truly Understand Video Anomalies? Revealing Hallucination via Co-Occurrence Patterns

  • Menghao Zhang
  • Huazheng Wang
  • Pengfei Ren
  • Kangheng Lin
  • Qi Qi
  • Haifeng Sun
  • Zirui Zhuang
  • Lei Zhang

Large Vision-Language Models (LVLMs) pretrained on large-scale multimodal data have shown promising capabilities in Video Anomaly Detection (VAD). However, their ability to reason about abnormal events based on scene semantics remains underexplored. In this paper, we investigate LVLMs’ behavior in VAD from a visual-textual co-occurrence perspective, focusing on whether their decisions are driven by statistical shortcuts between visual instances and textual phrases. By analyzing visual-textual co-occurrence in pretraining data and conducting experiments under different data settings, we reveal a hallucination phenomenon: LVLMs tend to rely on co-occurrence patterns between visual instances and textual phrases associated with either normality or abnormality, leading to incorrect predictions when these high-frequency objects appear in semantically mismatched contexts. To address this issue, we propose VAD-DPO, a direct preference optimization method supervised with counter-example pairs. By constructing visually similar but semantically contrasting video clips, VAD-DPO encourages the model to align its predictions with the semantics of scene rather than relying on co-occurrence patterns. Extensive experiments on six benchmark datasets demonstrate the effectiveness of VAD-DPO in enhancing both anomaly detection and reasoning performance, particularly in scene-dependent scenarios.

NeurIPS Conference 2025 Conference Paper

DP²O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution

  • Rongyuan Wu
  • Lingchen Sun
  • Zhengqiang Zhang
  • Shihao Wang
  • Tianhe Wu
  • Qiaosi Yi
  • Shuai Li
  • Lei Zhang

Benefiting from pre-trained text-to-image (T2I) diffusion models, real-world image super-resolution (Real-ISR) methods can synthesize rich and realistic details. However, due to the inherent stochasticity of T2I models, different noise inputs often lead to outputs with varying perceptual quality. Although this randomness is sometimes seen as a limitation, it also introduces a wider perceptual quality range, which can be exploited to improve Real-ISR performance. To this end, we introduce Direct Perceptual Preference Optimization for Real-ISR (DP²O-SR), a framework that aligns generative models with perceptual preferences without requiring costly human annotations. We construct a hybrid reward signal by combining full-reference and no-reference image quality assessment (IQA) models trained on large-scale human preference datasets. This reward encourages both structural fidelity and natural appearance. To better utilize perceptual diversity, we move beyond the standard best-vs-worst selection and construct multiple preference pairs from outputs of the same model. Our analysis reveals that the optimal selection ratio depends on model capacity: smaller models benefit from broader coverage, while larger models respond better to stronger contrast in supervision. Furthermore, we propose hierarchical preference optimization, which adaptively weights training pairs based on intra-group reward gaps and inter-group diversity, enabling more efficient and stable learning. Extensive experiments across both diffusion- and flow-based T2I backbones demonstrate that DP²O-SR significantly improves perceptual quality and generalizes well to real-world benchmarks.

EAAI Journal 2025 Journal Article

Enhancing the absolute positioning accuracy of welding robots based on joint error compensation

  • Bingqi Jia
  • Haihong Pan
  • Yukang Cai
  • Lei Zhang
  • Xuhong Chen
  • Lin Chen

This paper presented a novel method for improving the absolute positioning accuracy of robot end-effectors through joint space error prediction, addressing the issue of low accuracy in welding robots. A sampling method that considered both Cartesian and joint spaces was proposed, and a joint error compensation model based on Gaussian process regression (GPR) was established, utilizing the expected joint positions and joint errors as input features. A compensation strategy that integrated the error model into the robot controller was introduced. Experiments involving a laser tracker for positioning error compensation, including single-point multi-pose, spatial multi-point, and continuous welding trajectory tests, were conducted. The experimental results demonstrated that the proposed compensation method effectively reduced the mean absolute error (MAE) of welding robot positioning from approximately 0. 6 mm to within 0. 25 mm. This method significantly enhanced the absolute positioning accuracy of welding robot end-effectors, meeting the requirements for offline programming in welding applications.

AIIM Journal 2025 Journal Article

EvidenceMap: Learning evidence analysis to unleash the power of small language models for biomedical question answering

  • Chang Zong
  • Jian Wan
  • Siliang Tang
  • Lei Zhang

When addressing professional questions in the biomedical domain, humans typically acquire multiple pieces of information as evidence and engage in multifaceted analysis to provide high-quality answers. Current LLM-based question answering methods lack a detailed definition and learning process for evidence analysis, leading to the risk of error propagation and hallucinations while using evidence. Although increasing the parameter size of LLMs can alleviate these issues, it also presents challenges in training and deployment with limited resources. In this study, we propose EvidenceMap, which aims to enable a lightweight pre-trained language model to explicitly learn multiple aspects of biomedical evidence, including supportive evaluation, logical correlation and content summarization, thereby latently guiding a generative model (around 3B parameters) to provide textual responses. Experimental results demonstrate that our method, learning evidence analysis by fine-tuning a model with only 66M parameters, exceeds the RAG method with an 8B LLM by 19. 9% and 5. 7% in reference-based quality and accuracy, respectively. The code and dataset for reproducing our framework and experiments are available at https: //github. com/ZUST-BIT/EvidenceMap.

AAAI Conference 2025 Conference Paper

Fine-Tuning Language Models with Collaborative and Semantic Experts

  • Jiaxi Yang
  • Binyuan Hui
  • Min Yang
  • Jian Yang
  • Lei Zhang
  • Qiang Qu
  • Junyang Lin

Recent advancements in large language models (LLMs) have broadened their application scope but revealed challenges in balancing capabilities across general knowledge, coding, and mathematics. To address this, we introduce a Collaborative and Semantic Experts (CoE) approach for supervised fine-tuning (SFT), which employs a two-phase training strategy. Initially, expert training fine-tunes the feed-forward network on specialized datasets, developing distinct experts in targeted domains. Subsequently, expert leveraging synthesizes these trained experts into a structured model with semantic guidance to activate specific experts, enhancing performance and interpretability. Evaluations on comprehensive benchmarks across MMLU, HumanEval, GSM8K, MT-Bench, and AlpacaEval confirm CoE's efficacy, demonstrating improved performance and expert collaboration in diverse tasks, significantly outperforming traditional SFT methods.

AAAI Conference 2025 Conference Paper

GapMatch: Bridging Instance and Model Perturbations for Enhanced Semi-Supervised Medical Image Segmentation

  • Wei Huang
  • Lei Zhang
  • Zizhou Wang
  • Yan Wang

Medical image segmentation provides detailed understanding and aids in diagnosis, treatment planning, and monitoring of diseases. Due to the high cost of acquiring labeled data in the field of medical image analysis, semi-supervised segmentation methods have garnered increasing attention. Benefiting from their simplicity and effectiveness, consistency regularization-based methods have emerged as a significant research focus by utilizing perturbations. However, existing methods typically consider perturbation strategies from only a single perspective: either instance perturbation or model perturbation, thus ignoring the potential benefit of effectively combining both. In response, we propose a unified perturbation framework named GapMatch, which bridges instance and model perturbations to broaden the perturbation space and employs dual perturbation to impose consistency regularization on the model. Specifically, GapMatch involves using instance perturbation to update the decision boundary and model perturbation to further optimize it. These two steps mutually reinforce each other in an iterative manner, effectively pushing the decision boundary towards low-density regions while maximizing the class margin. Extensive experimental results on two popular medical image benchmarks demonstrate the effectiveness and generality of the proposed method.

AAAI Conference 2025 Conference Paper

GaussianSR: High Fidelity 2D Gaussian Splatting for Arbitrary-Scale Image Super-Resolution

  • Jintong Hu
  • Bin Xia
  • Bin Chen
  • Wenming Yang
  • Lei Zhang

Implicit neural representations (INRs) have revolutionized arbitrary-scale super-resolution (ASSR) by modeling images as continuous functions. Most existing INR-based ASSR networks first extract features from the given low-resolution image using an encoder, and then render the super-resolved result via a multi-layer perceptron decoder. Although these approaches have shown promising results, their performance is constrained by the limited representation ability of discrete latent codes in the encoded features. In this paper, we propose a novel ASSR method named GaussianSR that overcomes this limitation through 2D Gaussian Splatting (2DGS). Unlike traditional methods that treat pixels as discrete points, GaussianSR represents each pixel as a continuous Gaussian field. The encoded features are simultaneously refined and upsampled by rendering the mutually stacked Gaussian fields. As a result, long-range dependencies are established to enhance representation ability. In addition, a classifier is developed to dynamically assign Gaussian kernels to all pixels to further improve flexibility. All components of GaussianSR (i.e. encoder, classifier, Gaussian kernels, and decoder) are jointly learned end-to-end. Experiments demonstrate that GaussianSR achieves superior ASSR performance with fewer parameters than existing methods while enjoying interpretable and content-aware feature aggregations.

AAAI Conference 2025 Conference Paper

Generalizable Sensor-Based Activity Recognition via Categorical Concept Invariant Learning

  • Di Xiong
  • Shuoyuan Wang
  • Lei Zhang
  • Wenbo Huang
  • Chaolei Han

Human Activity Recognition (HAR) aims to recognize activities by training models on massive sensor data. In real-world deployment, a crucial aspect of HAR that has been largely overlooked is that the test sets may have different distributions from training sets due to inter-subject variability including age, gender, behavioral habits, etc., which leads to poor generalization performance. One promising solution is to learn domain-invariant representations to enable a model to generalize on an unseen distribution. However, most existing methods only consider the feature-invariance of the penultimate layer for domain-invariant learning, which leads to suboptimal results. In this paper, we propose a Categorical Concept Invariant Learning (CCIL) framework for generalizable activity recognition, which introduces a concept matrix to regularize the model in the training stage by simultaneously concertrating on feature-invariance and logit-invariance. Our key idea is that the concept matrix for samples belonging to the same activity category should be similar. Extensive experiments on four public HAR benchmarks demonstrate that our CCIL substantially outperforms the state-of-the-art approaches under cross-person, cross-dataset, cross-position, and one-person-to-another settings.

NeurIPS Conference 2025 Conference Paper

GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and Generation

  • Zhengqiang Zhang
  • Rongyuan Wu
  • Lingchen Sun
  • Lei Zhang

Effective and efficient tokenization plays an important role in image representation and generation. Conventional methods, constrained by uniform 2D/1D grid tokenization, are inflexible to represent regions with varying shapes and textures and at different locations, limiting their efficacy of feature representation. In this work, we propose GPSToken, a novel G aussian P arameterized S patially-adaptive Token ization framework, to achieve non-uniform image tokenization by leveraging parametric 2D Gaussians to dynamically model the shape, position, and textures of different image regions. We first employ an entropy-driven algorithm to partition the image into texture-homogeneous regions of variable sizes. Then, we parameterize each region as a 2D Gaussian (mean for position, covariance for shape) coupled with texture features. A specialized transformer is trained to optimize the Gaussian parameters, enabling continuous adaptation of position/shape and content-aware feature extraction. During decoding, Gaussian parameterized tokens are reconstructed into 2D feature maps through a differentiable splatting-based renderer, bridging our adaptive tokenization with standard decoders for end-to-end training. GPSToken disentangles spatial layout (Gaussian parameters) from texture features to enable efficient two-stage generation: structural layout synthesis using lightweight networks, followed by structure-conditioned texture generation. Experiments demonstrate the state-of-the-art performance of GPSToken, which achieves rFID and FID scores of 0. 65 and 1. 50 on image reconstruction and generation tasks using 128 tokens, respectively. Codes and models of GPSToken can be found at https: //github. com/xtudbxk/GPSToken.

AAAI Conference 2025 Conference Paper

Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs

  • Lei Zhang
  • Yunshui Li
  • Jiaming Li
  • Xiaobo Xia
  • Jiaxi Yang
  • Run Luo
  • Minzheng Wang
  • Longze Chen

Some of the latest released Code Large Language Models (Code LLMs) have been trained on repository-level code data, enabling them to perceive repository structures and utilize cross-file code information. This capability allows us to directly concatenate the content of repository code files in prompts to achieve repository-level code completion. However, in real development scenarios, directly concatenating all code repository files in a prompt can easily exceed the context window of Code LLMs, leading to a significant decline in completion performance. Additionally, overly long prompts can increase completion latency, negatively impacting the user experience. In this study, we conducted extensive experiments, including completion error analysis, topology dependency analysis, and cross-file content analysis, to investigate the factors affecting repository-level code completion. Based on the conclusions drawn from these preliminary experiments, we proposed a strategy called **Hierarchical Context Pruning (HCP)** to construct high-quality completion prompts. We applied the **HCP** to six Code LLMs and evaluated them on the CrossCodeEval dataset. The experimental results showed that, compared to previous methods, the prompts constructed using our **HCP** strategy achieved higher completion accuracy on five out of six Code LLMs. Additionally, the **HCP** managed to keep the prompt length around 8k tokens (whereas the full repository code is approximately 50k tokens), significantly improving completion throughput. Our code and data will be publicly available.

AAAI Conference 2025 Conference Paper

Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text Detection

  • Jiaqi Chen
  • Xiaoye Zhu
  • Tianyang Liu
  • Ying Chen
  • Chen Xinhui
  • Yiwen Yuan
  • Chak Tou Leong
  • Zuchao Li

Large Language Models (LLMs) have revolutionized text generation, making detecting machine-generated text increasingly challenging. Although past methods have achieved good performance on detecting pure machine-generated text, those detectors have poor performance on distinguishing machine-revised text (rewriting, expansion, and polishing), which can have only minor changes from its original human prompt. As the content of text may originate from human prompts, detecting machine-revised text often involves identifying distinctive machine styles, e.g., worded favored by LLMs. However, existing methods struggle to detect machine-style phrasing hidden within the content contributed by humans. We propose the “Imitate Before Detect” (ImBD) approach, which first imitates the machine-style token distribution, and then compares the distribution of the text to be tested with the machine-style distribution to determine whether the text has been machine-revised. To this end, we introduce Style Preference Optimization (SPO), which aligns a scoring LLM model to the preference of text styles generated by machines. The aligned scoring model is then used to calculate the style-conditional probability curvature (Style-CPC), quantifying the log probability difference between the original and conditionally sampled texts for effective detection. We conduct extensive comparisons across various scenarios, encompassing text revisions by six LLMs, four distinct text domains, and three machine revision types. Compared to existing state-of-the-art methods, our method yields a 13% increase in AUC for detecting text revised by open-source LLMs, and improves performance by 5% and 19% for detecting GPT-3.5 and GPT-4o revised text, respectively. Notably, our method surpasses the commercially trained GPT-Zero with just 1,000 samples and five minutes of SPO, demonstrating its efficiency and effectiveness.

NeurIPS Conference 2025 Conference Paper

InstructRestore: Region-Customized Image Restoration with Human Instructions

  • Shuaizheng Liu
  • Jianqi Ma
  • Lingchen Sun
  • Xiangtao Kong
  • Lei Zhang

Despite the significant progress in diffusion prior-based image restoration for real-world scenarios, most existing methods apply uniform processing to the entire image, lacking the capability to perform region-customized image restoration according to user preferences. In this work, we propose a new framework, namely InstructRestore, to perform region-adjustable image restoration following human instructions. To achieve this, we first develop a data generation engine to produce training triplets, each consisting of a high-quality image, the target region description, and the corresponding region mask. With this engine and careful data screening, we construct a comprehensive dataset comprising 536, 945 triplets to support the training and evaluation of this task. We then examine how to integrate the low-quality image features under the ControlNet architecture to adjust the degree of image details enhancement. Consequently, we develop a ControlNet-like model to identify the target region and allocate different integration scales to the target and surrounding regions, enabling region-customized image restoration that aligns with user instructions. Experimental results demonstrate that our proposed InstructRestore approach enables effective human-instructed image restoration, including restoration with controllable bokeh blur effects and region-specific restoration with continuous intensity control. Our work advances the investigation of interactive image restoration and enhancement techniques. Data, code, and models are publicly available at https: //github. com/shuaizhengliu/InstructRestore. git.

EAAI Journal 2025 Journal Article

Learning multi-color curve for image harmonization

  • Jingrong Yuan
  • Hao Wu
  • Lidong Xie
  • Lei Zhang
  • Jichen Xing

Due to the varying shooting conditions, composite images often lack realism between the foreground and the back ground. As an important and challenging visual task, image harmonization can effectively improve visual effect of composite images. Currently, image harmonization methods have achieved satisfied performance on public dataset. However, in some challenging examples with substantial color disparities between the foreground and the background, existing methods get poor results. To solve this problem, we propose a Multi-color Curve Net that processes images through multiple color spaces to capture richer color information. Our Multi-color Curve Net performs multi-stage curve learning in different color spaces with the encoder composed of modified Transformer blocks. Simultaneously, we introduce a Multi-color Integration Module to effectively fuse the information extracted from different color spaces and further improve the results by a lightweight Fine-grained Optimization Module. The Multi-color Curve Net gains high performance while maintaining a small parameter scale. Experiments on benchmark demonstrate that the Multi-color Curve Net outperforms state-of-the-art methods in terms of peak signal to-noise ratio (PSNR), structural similarity (SSIM) and foreground mean squared error (fMSE) with fewer parameters. The code for our method is available at https: //github. com/gmrj2024/MC2Net.

JBHI Journal 2025 Journal Article

Learning Sensor Sample-Reweighting for Dynamic Early-Exit Activity Recognition Via Meta Learning

  • Zenan Fu
  • Lei Zhang
  • Wenbo Huang
  • Dongzhou Cheng
  • Hao Wu
  • Aiguo Song

During recent years, dynamic early-exit has provided a promising paradigm to improve the computational efficiency of deep neural networks by constructing multiple classifiers to let easy samples exit at shallow layers while avoiding redundant computations at deep exits, which has been seldom explored in the context of latency-aware human activity recognition (HAR) deployed on wearable devices. Particularly, most existing early-exit strategies have always treated all activity samples equally at each exit during training, which ignore such dynamic early-exit behavior at test-time, causing a potential mismatch between training and test. Intuitively, easy activity samples that often exit earlier at test-time should place more emphasis on the training loss of shallow classifiers, while hard activity samples should contribute more to the training loss of deep classifiers. To bridge this gap, this paper introduces a sample-reweighting approach for efficient activity inference, which employs a weight-predicting network to reweight the training loss of different activity samples at every exit. From a perspective of meta learning, a new optimization objective function is designed to jointly optimize both weight-predicting network and backbone network. We perform extensive experiments on three popular HAR benchmarks including UCI-HAR, WISDM, and UniMiB-SHAR, which demonstrate that while incorporating such test-time early-exit behavior into conventional training pipeline, it can consistently improve the accuracy-efficiency trade-offs under budgeted batch classification and anytime prediction patterns. Moreover, our approach has a natural advantage in handing class-imbalance HAR problem. Detailed ablation studies, visualized illustrations, and real hardware deployment are provided to support our statement.

AAAI Conference 2025 Conference Paper

Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-Sequence

  • Wenbo Huang
  • Jinghui Zhang
  • Guang Li
  • Lei Zhang
  • Shuoyuan Wang
  • Fang Dong
  • Jiahui Jin
  • Takahiro Ogawa

In few-shot action recognition (FSAR), long sub-sequences of video naturally express entire actions more effectively. However, the high computational complexity of mainstream Transformer-based methods limits their application. Recent Mamba demonstrates efficiency in modeling long sequences, but directly applying Mamba to FSAR overlooks the importance of local feature modeling and alignment. Moreover, long sub-sequences within the same class accumulate intra-class variance, which adversely impacts FSAR performance. To solve these challenges, we propose a Matryoshka MAmba and CoNtrasTive LeArning framework (Manta). Firstly, the Matryoshka Mamba introduces multiple Inner Modules to enhance local feature representation, rather than directly modeling global features. An Outer Module captures dependencies of timeline between these local features for implicit temporal alignment. Secondly, a hybrid contrastive learning paradigm, combining both supervised and unsupervised methods, is designed to mitigate the negative effects of intra-class variance accumulation. The Matryoshka Mamba and the hybrid contrastive learning paradigm operate in two parallel branches within Manta, enhancing Mamba for FSAR of long sub-sequence. Manta achieves new state-of-the-art performance on prominent benchmarks, including SSv2, Kinetics, UCF101, and HMDB51. Extensive empirical studies prove that Manta significantly improves FSAR of long sub-sequence from multiple perspectives.

AAAI Conference 2025 Conference Paper

MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis

  • Wanggui He
  • Siming Fu
  • Mushui Liu
  • Xierui Wang
  • Wenyi Xiao
  • Fangxun Shu
  • Yi Wang
  • Lei Zhang

Auto-regressive models have made significant progress in the realm of text-to-image synthesis, yet devising an appropriate model architecture and training strategy to achieve a satisfactory level remains an important avenue of exploration. In this work, we introduce MARS, a novel framework for T2I generation that incorporates a specially designed Semantic Vision-Language Integration Expert (SemVIE). This innovative component integrates pre-trained LLMs by independently processing linguistic and visual information—freezing the textual component while fine-tuning the visual component. This methodology preserves the NLP capabilities of LLMs while imbuing them with exceptional visual understanding. Building upon the powerful base of the pre-trained Qwen-7B, MARS stands out with its bilingual generative capabilities corresponding to both English and Chinese language prompts and the capacity for joint image and text generation. The flexibility of this framework lends itself to migration towards any-to-any task adaptability. Furthermore, MARS employs a multi-stage training strategy that first establishes robust image-text alignment through complementary bidirectional tasks and subsequently concentrates on refining the T2I generation process, significantly augmenting text-image synchrony and the granularity of image details. Notably, MARS requires only 9% of the GPU days needed by SD1.5, yet it achieves remarkable results across a variety of benchmarks, illustrating the training efficiency and the potential for swift deployment in various applications.

NeurIPS Conference 2025 Conference Paper

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM

  • Bowen Dong
  • Minheng Ni
  • Zitong Huang
  • Guanglei Yang
  • Wangmeng Zuo
  • Lei Zhang

Multimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse causes. Existing benchmarks fail to adequately distinguish between perception-induced hallucinations and reasoning-induced hallucinations. This failure constitutes a significant issue and hinders the diagnosis of multimodal reasoning failures within MLLMs. To address this, we propose the MIRAGE benchmark, which isolates reasoning hallucinations by constructing questions where input images are correctly perceived by MLLMs yet reasoning errors persist. MIRAGE introduces multi-granular evaluation metrics: accuracy, factuality, and LLMs hallucination score for hallucination quantification. Our analysis reveals strong correlations between question types and specific hallucination patterns, particularly systematic failures of MLLMs in spatial reasoning involving complex relationships (\emph{e. g. }, complex geometric patterns across images). This highlights a critical limitation in the reasoning capabilities of current MLLMs and provides targeted insights for hallucination mitigation on specific types. To address these challenges, we propose Logos, a method that combines curriculum reinforcement fine-tuning to encourage models to generate logic-consistent reasoning chains by stepwise reducing learning difficulty, and collaborative hint inference to reduce reasoning complexity. Logos establishes a baseline on MIRAGE, and reduces the logical hallucinations in original base models. Link: \url{https: //bit. ly/25mirage}.

NeurIPS Conference 2025 Conference Paper

MobileODE: An Extra Lightweight Network

  • Le Yu
  • Jun Wu
  • Bo Gou
  • Xiangde Min
  • Lei Zhang
  • Zhang Yi
  • Tao He

Depthwise-separable convolution has emerged as a significant milestone in the lightweight development of Convolutional Neural Networks (CNNs) over the past decade. This technique consists of two key components: depthwise convolution, which captures spatial information, and pointwise convolution, which enhances channel interactions. In this paper, we propose a novel method to lightweight CNNs through the discretization of Ordinary Differential Equations (ODEs). Specifically, we optimize depthwise-separable convolution by replacing the pointwise convolution with a discrete ODE module, termed the \emph{\textbf{C}hannelwise \textbf{O}DE \textbf{S}olver (COS)}. The COS module is constructed by a simple yet efficient direct differentiation Euler algorithm, using learnable increment parameters. This replacement reduces parameters by over $98. 36$\% compared to conventional pointwise convolution. By integrating COS into MobileNet, we develop a new extra lightweight network called MobileODE. With carefully designed basic and inverse residual blocks, the resulting MobileODEV1 and MobileODEV2 reduce channel interaction parameters by $71. 0$\% and $69. 2$\%, respectively, compared to MobileNetV1, while achieving higher accuracy across various tasks, including image classification, object detection, and semantic segmentation. The code is available at {\url{https: //github. com/cashily/MobileODE}}.

AAAI Conference 2025 Conference Paper

Multi-Edge Reinforced Collaborative Data Acquisition for Continuous Video Analytics by Prioritizing Quality over Quantity

  • Lei Zhang
  • Guanyu Gao
  • Haiyan Yin
  • Huaizheng Zhang

Edge computing-based video analytics faces data drift issues due to the occurrence of unseen objects or scenes in ever-changing environments. To maintain accuracy, continuous learning (CL) retrains stale models periodically with newly obtained data. However, it leads to unaffordable costs, as we must keep labeling drift data and retraining models. Regarding this concern, we first investigate video patterns across multiple cameras within an area and reveal significant data redundancies. We find that many of the same objects can be captured by multiple edge cameras or appear many times on the same edges. Our quantitative findings suggest that selecting a subset of high-quality data for CL is preferable over using a larger quantity. Yet, existing efforts for data acquisition have only focused on a single static dataset. These methods are not suitable for multi-edge video analytics scenarios, where videos are captured from multiple sources with non-iid data distribution. Hence, we propose a multi-edge collaborative active video acquisition (AVA) framework to collaboratively learn a reinforced video acquisition strategy to identify informative video frames from multiple edge nodes that best enhance model accuracy, avoiding redundancy across edges. Extensive experiments on three video datasets demonstrate that, our method achieves comparable performance to full-set video training while utilizing only 20% of the data in classification tasks. In object detection tasks, our methods can maintain productive accuracy with a reduction of nearly 70% in training video frames.

NeurIPS Conference 2025 Conference Paper

One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution

  • Yujing Sun
  • Lingchen Sun
  • Shuaizheng Liu
  • Rongyuan Wu
  • Zhengqiang Zhang
  • Lei Zhang

It is a challenging problem to reproduce rich spatial details while maintaining temporal consistency in real-world video super-resolution (Real-VSR), especially when we leverage pre-trained generative models such as stable diffusion (SD) for realistic details synthesis. Existing SD-based Real-VSR methods often compromise spatial details for temporal coherence, resulting in suboptimal visual quality. We argue that the key lies in how to effectively extract the degradation-robust temporal consistency priors from the low-quality (LQ) input video and enhance the video details while maintaining the extracted consistency priors. To achieve this, we propose a Dual LoRA Learning (DLoRAL) paradigm to train an effective SD-based one-step diffusion model, achieving realistic frame details and temporal consistency simultaneously. Specifically, we introduce a Cross-Frame Retrieval (CFR) module to aggregate complementary information across frames, and train a Consistency-LoRA (C-LoRA) to learn robust temporal representations from degraded inputs. After consistency learning, we fix the CFR and C-LoRA modules and train a Detail-LoRA (D-LoRA) to enhance spatial details while aligning with the temporal space defined by C-LoRA to keep temporal coherence. The two phases alternate iteratively for optimization, collaboratively delivering consistent and detail-rich outputs. During inference, the two LoRA branches are merged into the SD model, allowing efficient and high-quality video restoration in a single diffusion step. Experiments show that DLoRAL achieves strong performance in both accuracy and speed. Code and models will be released.

NeurIPS Conference 2025 Conference Paper

OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis

  • Run Luo
  • Ting-En Lin
  • Haonan Zhang
  • Yuchuan Wu
  • Xiong Liu
  • Yongbin Li
  • Longze Chen
  • Jiaming Li

Recent advancements in omnimodal learning have significantly improved understanding and generation across images, text, and speech, yet these developments remain predominantly confined to proprietary models. The lack of high-quality omnimodal datasets and the challenges of real-time emotional speech synthesis have notably hindered progress in open-source research. To address these limitations, we introduce OpenOmni, a two-stage training framework that integrates omnimodal alignment and speech generation to develop a state-of-the-art omnimodal large language model. In the alignment phase, a pretrained speech model undergoes further training on image-text tasks, enabling (near) zero-shot generalization from vision to speech, outperforming models trained on tri-modal datasets. In the speech generation phase, a lightweight decoder is trained on speech tasks with direct preference optimization, which enables real-time emotional speech synthesis with high fidelity. Extensive experiments demonstrate that OpenOmni surpasses state-of-the-art models across omnimodal, vision-language, and speech-language benchmarks. It achieves a 4-point absolute improvement on OmniBench over the leading open-source model VITA, despite using 5$\times$ fewer training examples and a smaller model size (7B vs. 7$\times$8B). Besides, OpenOmni achieves real-time speech generation with less than 1 second latency at non-autoregressive mode, reducing inference time by 5$\times$ compared to autoregressive methods, and improves emotion classification accuracy by 7. 7\%. The codebase is available at https: //github. com/RainBowLuoCS/OpenOmni.

NeurIPS Conference 2025 Conference Paper

PASS: Path-selective State Space Model for Event-based Recognition

  • Jiazhou Zhou
  • Kanghao Chen
  • Lei Zhang
  • Lin Wang

Event cameras are bio-inspired sensors that capture intensity changes asynchronously with distinct advantages, such as high temporal resolution. Existing methods for event-based object/action recognition predominantly sample and convert event representation at every fixed temporal interval (or frequency). However, they are constrained to processing a limited number of event lengths and show poor frequency generalization, thus not fully leveraging the event's high temporal resolution. In this paper, we present our PASS framework, exhibiting superior capacity for spatiotemporal event modeling towards a larger number of event lengths and generalization across varying inference temporal frequencies. Our key insight is to learn adaptively encoded event features via the state space models (SSMs), whose linear complexity and generalization on input frequency make them ideal for processing high temporal resolution events. Specifically, we propose a Path-selective Event Aggregation and Scan (PEAS) module to encode events into features with fixed dimensions by adaptively scanning and selecting aggregated event presentation. On top of it, we introduce a novel Multi-faceted Selection Guiding (MSG) loss to minimize the randomness and redundancy of the encoded features during the PEAS selection process. Our method outperforms prior methods on five public datasets and shows strong generalization across varying inference frequencies with less accuracy drop (ours -8. 62% v. s. -20. 69% for the baseline). Moreover, our model exhibits strong long spatiotemporal modeling for a broader distribution of event length (1-10^9), precise temporal perception, and effective generalization for real-world scenarios. Code and checkpoints will be released upon acceptance.

NeurIPS Conference 2025 Conference Paper

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

  • Weifeng Lin
  • Xinyu Wei
  • Ruichuan An
  • Tianhe Ren
  • Tingwei Chen
  • Renrui Zhang
  • Ziyu Guo
  • Wentao Zhang

We present Perceive Anything Model (PAM), a conceptually straightforward and efficient framework for comprehensive region-level visual understanding in images and videos. Our approach extends the powerful segmentation model SAM 2 by integrating Large Language Models (LLMs), enabling simultaneous object segmentation with the generation of diverse, region-specific semantic outputs, including categories, label definition, functional explanations, and detailed captions. A key component, Semantic Perceiver, is introduced to efficiently transform SAM 2's rich visual features, which inherently carry general vision, localization, and semantic priors into multi-modal tokens for LLM comprehension. To support robust multi-granularity understanding, we also develop a dedicated data refinement and augmentation pipeline, yielding a high-quality dataset of 1. 5M image and 0. 6M video region-semantic annotations, including novel region-level streaming video caption data. PAM is designed for lightweightness and efficiency, while also demonstrates strong performance across a diverse range of region understanding tasks. It runs 1. 2$-$2. 4$\times$ faster and consumes less GPU memory than prior approaches, offering a practical solution for real-world applications. We believe that our effective approach will serve as a strong baseline for future research in region-level visual understanding.

ICRA Conference 2025 Conference Paper

Physics-Informed Split Koopman Operators for Data-Efficient Soft Robotic Simulation

  • Eron Ristich
  • Lei Zhang
  • Yi Ren
  • Jiefeng Sun

Koopman operator theory provides a powerful data-driven technique for modeling nonlinear dynamical systems in a linear framework, in comparison to computationally expensive and highly nonlinear physics-based simulations. However, Koopman operator-based models for soft robots are very high dimensional and require considerable amounts of data to properly resolve. Inspired by physics-informed techniques from machine learning, we present a novel physics-informed Koopman operator identification method that improves simulation accuracy for small dataset sizes. Through Strang splitting, the method takes advantage of both continuous and discrete Koopman operator approximation to obtain information both from trajectory and phase space data. The method is validated on a tendon-driven soft robotic arm, showing orders of magnitude improvement over standard methods in terms of the shape error. We envision this method can significantly reduce the data requirement of Koopman operators for systems with partially known physical models, and thus reduce the cost of obtaining data. More info: https://sunrobotics.lab.asu.edu/blog/2024/ristich-icra-2025/

EAAI Journal 2025 Journal Article

Point cloud semantic segmentation network based on graph convolution and attention mechanism

  • Nan Yang
  • Yong Wang
  • Lei Zhang
  • Bin Jiang

Point cloud data provides rich three-dimensional spatial information. Accurate three-dimensional point cloud semantic segmentation algorithms enhance environmental understanding and perception, with wide-ranging applications in autonomous driving and scene analysis. However, Graph Neural Networks often struggle to retain semantic relationships among neighboring points during feature extraction, potentially leading to the omission of critical features during aggregation. To address these challenges, we propose a novel network, the Feature-Enhanced Residual Attention Network. This network includes an innovative graph convolution module, the Neighborhood-Enhanced Convolutional Aggregation Module, which utilizes K-Nearest Neighbor and Dilated K-Nearest Neighbor techniques to construct diverse dynamic graphs and aggregate features, thereby prioritizing essential information. This approach significantly enhances the expressiveness and generalization capabilities of the network. Additionally, we introduce a new spatial attention module designed to capture semantic relationships among points. Experimental results demonstrate that the Feature-Enhanced Residual Attention Network outperforms benchmark models, achieving an average intersection ratio of 61. 3% and an overall accuracy of 86. 7% on the Stanford Large-Scale Three-dimensional Indoor Spaces dataset, thereby significantly improving semantic segmentation performance.

NeurIPS Conference 2025 Conference Paper

Polyline Path Masked Attention for Vision Transformer

  • Zhongchen Zhao
  • Chaodong Xiao
  • Hui Lin
  • Qi Xie
  • Lei Zhang
  • Deyu Meng

Global dependency modeling and spatial position modeling are two core issues of the foundational architecture design in current deep learning frameworks. Recently, Vision Transformers (ViTs) have achieved remarkable success in computer vision, leveraging the powerful global dependency modeling capability of the self-attention mechanism. Furthermore, Mamba2 has demonstrated its significant potential in natural language processing tasks by explicitly modeling the spatial adjacency prior through the structured mask. In this paper, we propose Polyline Path Masked Attention (PPMA) that integrates the self-attention mechanism of ViTs with an enhanced structured mask of Mamba2, harnessing the complementary strengths of both architectures. Specifically, we first ameliorate the traditional structured mask of Mamba2 by introducing a 2D polyline path scanning strategy and derive its corresponding structured mask, polyline path mask, which better preserves the adjacency relationships among image tokens. Notably, we conduct a thorough theoretical analysis on the structural characteristics of the proposed polyline path mask and design an efficient algorithm for the computation of the polyline path mask. Next, we embed the polyline path mask into the self-attention mechanism of ViTs, enabling explicit modeling of spatial adjacency prior. Extensive experiments on standard benchmarks, including image classification, object detection, and segmentation, demonstrate that our model outperforms previous state-of-the-art approaches based on both state-space models and Transformers. For example, our proposed PPMA-T/S/B models achieve 48. 7%/51. 1%/52. 3% mIoU on the ADE20K semantic segmentation task, surpassing RMT-T/S/B by 0. 7%/1. 3%/0. 3%, respectively. Code is available at https: //github. com/zhongchenzhao/PPMA.

IJCAI Conference 2025 Conference Paper

Prompt-Free Conditional Diffusion for Multi-object Image Augmentation

  • Haoyu Wang
  • Lei Zhang
  • Wei Wei
  • Chen Ding
  • Yanning Zhang

Diffusion model has underpinned much recent advances of dataset augmentation in various computer vision tasks. However, when involving generating multi-object images as real scenarios, most existing methods either rely entirely on text condition, resulting in a deviation between the generated objects and the original data, or rely too much on the original images, resulting in a lack of diversity in the generated images, which is of limited help to downstream tasks. To mitigate both problems with one stone, we propose a prompt-free conditional diffusion framework for multi-object image augmentation. Specifically, we introduce a local-global semantic fusion strategy to extract semantics from images to replace text, and inject knowledge into the diffusion model through LoRA to alleviate the category deviation between the original model and the target dataset. In addition, we design a reward model based counting loss to assist the traditional reconstruction loss for model training. By constraining the object counts of each category instead of pixel-by-pixel constraints, bridging the quantity deviation between the generated data and the original data while improving the diversity of the generated data. Experimental results demonstrate the superiority of the proposed method over several representative state-of-the-art baselines and showcase strong downstream task gain and out-of-domain generalization capabilities. Code is available at \href{https: //github. com/00why00/PFCD}{here}.

NeurIPS Conference 2025 Conference Paper

Registration is a Powerful Rotation-Invariance Learner for 3D Anomaly Detection

  • Yuyang Yu
  • Zhengwei Chen
  • Xuemiao Xu
  • Lei Zhang
  • Haoxin Yang
  • Yongwei Nie
  • Shengfeng He

3D anomaly detection in point-cloud data is critical for industrial quality control, aiming to identify structural defects with high reliability. However, current memory bank-based methods often suffer from inconsistent feature transformations and limited discriminative capacity, particularly in capturing local geometric details and achieving rotation invariance. These limitations become more pronounced when registration fails, leading to unreliable detection results. We argue that point-cloud registration plays an essential role not only in aligning geometric structures but also in guiding feature extraction toward rotation-invariant and locally discriminative representations. To this end, we propose a registration-induced, rotation-invariant feature extraction framework that integrates the objectives of point-cloud registration and memory-based anomaly detection. Our key insight is that both tasks rely on modeling local geometric structures and leveraging feature similarity across samples. By embedding feature extraction into the registration learning process, our framework jointly optimizes alignment and representation learning. This integration enables the network to acquire features that are both robust to rotations and highly effective for anomaly detection. Extensive experiments on the Anomaly-ShapeNet and Real3D-AD datasets demonstrate that our method consistently outperforms existing approaches in effectiveness and generalizability.

NeurIPS Conference 2025 Conference Paper

Rethinking Out-of-Distribution Detection and Generalization with Collective Behavior Dynamics

  • Zhenbin Wang
  • Lei Zhang
  • Wei Huang
  • Zhao Zhang
  • Zizhou Wang

Out-of-distribution (OOD) problems commonly occur when models process data with a distribution significantly deviates from the in-distribution (InD) training data. In this paper, we hypothesize that a $\textit{field}$ or $\textit{potential}$ more essential than features exists, and features are not the ultimate essence of the data but rather manifestations of them during training. we investigate OOD problems from the perspective of collective behavior dynamics. With this in mind, we first treat the output of the feature extractor as charged particles and investigate their collective behavior dynamics within a self-consistent electric field. Then, to characterize the relationship between OOD problems and dynamical equations, we introduce the $\textit{basin of attraction}$ and prove that its boundary can be represented as the zero level set of a differentiable function of the potential, $\textit{i. e. }$, the spatial integral of field. We further demonstrate that: $\textit{i)}$ InD and OOD inputs can be effectively separated based on whether they are steady state solutions for specific field conditions, enabling robust OOD detection and outperforming prior methods over three benchmarks. $\textit{ii)}$ the generalization capability correlates positively with the basin of attraction. By analyzing the dynamics of perturbations, we propose that the potential is well-characterized by a Fourier-domain form of the Poisson equation. Evaluated on six benchmark datasets, our method rivals the SoTA approaches for OOD generalization and can be seamlessly integrated with them to deliver additional gains.

AAAI Conference 2025 Conference Paper

SLRL: Semi-Supervised Local Community Detection Based on Reinforcement Learning

  • Li Ni
  • Rui Ye
  • Wenjian Luo
  • Yiwen Zhang
  • Lei Zhang
  • Victor S. Sheng

Most existing semi-supervised community detection algorithms leverage known communities to learn community structures, subsequently identifying communities that align with these learned community structures. However, differences in community structures may render the community structures learned by these methods inappropriate for the community containing the given node of interest. As a result, the identified community may exclude the given node or be of poor quality. Inspired by the success of reinforcement learning, we propose a Semi-supervised Local community detection method based on Reinforcement Learning, named SLRL, which only explores parts of the network surrounding the given node. It first extracts the local structure around a given node with an extractor, followed by selecting communities that are similar to this local structure to distill useful communities. These selected communities are employed to train the expander, which expands the community containing a given node. Experimental results demonstrate that SLRL outperforms state-of-the-art algorithms on five real-world datasets.

AAAI Conference 2025 Conference Paper

SyncNoise: Geometrically Consistent Noise Prediction for Instruction-based 3D Editing

  • Ruihuang Li
  • Liyi Chen
  • Zhengqiang Zhang
  • Varun Jampani
  • Vishal M. Patel
  • Lei Zhang

Text-based 2D diffusion models have demonstrated impressive capabilities in image generation and editing. Meanwhile, the 2D diffusion models also exhibit substantial potentials for 3D editing tasks. However, how to achieve consistent edits across multiple viewpoints remains a challenge. While the iterative dataset update method is capable of achieving global consistency, it suffers from slow convergence and over-smoothed textures. We propose SyncNoise, a novel geometry-guided multi-view consistent noise editing approach for high-fidelity 3D scene editing. SyncNoise synchronously edits multiple views with 2D diffusion models while enforcing multi-view noise predictions to be geometrically consistent, which ensures global consistency in both semantic structure and low-frequency appearance. To further enhance local consistency in high-frequency details, we set a group of anchor views and propagate them to their neighboring frames through cross-view reprojection. To improve the reliability of multi-view correspondences, we introduce depth supervision during training to enhance the reconstruction of precise geometries. Our method achieves high-quality 3D editing results respecting the textual instructions, especially in scenes with complex textures, by enhancing geometric consistency at the noise and pixel levels.

TAAS Journal 2025 Journal Article

The Comp-TSSs Scheme for Anomaly Detection in AI-Powered Autonomous Driving

  • Jiuzhen Zeng
  • Laurence T. Yang
  • Chao Wang
  • Lei Zhang

Given the vulnerability of vehicular networks to security attacks and the criticality of secure AI-powered autonomous driving, this paper emphasizes the security issue concerning vehicular networks in AI-powered autonomous vehicles. The novel complementary tensor summary statistics named as Comp-TSSs, is proposed for the statistical depiction of discrepancy between normal and abnormal volume instances in vehicular networks. This suggested Comp-TSSs enhances vehicular network security by incorporating reconstruction and regularization statistic terms derived from TPCA, which is extended from PCA through a fresh perspective of fully diagonalizing the covariance tensor. Comp-TSSs effectively captures multi-dimensional correlations in vehicular network volume data, providing complementary measures for representation residuals and weighted distances of instances projected in the principal tensor subspace. Building upon Comp-TSSs, a non-parametric statistic framework is developed for real-time detection of diverse volume anomalies, ensuring the security of AI-powered autonomous driving. The theoretical analyses concerning its detection performance and parameter selection are provided as well. Extensive experiments on synthetic and real-world datasets validate our superior vehicular network security monitoring system for AI-powered autonomous vehicles. It demonstrates higher true positive rates, lower false alarm rates, and minimal detection delays, even when both of the energy and variance anomalies are present.

NeurIPS Conference 2025 Conference Paper

The Underappreciated Power of Vision Models for Graph Structural Understanding

  • Xinjian Zhao
  • Wei Pang
  • Zhongkai Xue
  • Xiangru Jian
  • Lei Zhang
  • Yaoyao Xu
  • Xiaozhuang Song
  • Shu Wu

Graph Neural Networks operate through bottom-up message-passing, fundamentally differing from human visual perception, which intuitively captures global structures first. We investigate the underappreciated potential of vision models for graph understanding, finding they achieve performance comparable to GNNs on established benchmarks while exhibiting distinctly different learning patterns. These divergent behaviors, combined with limitations of existing benchmarks that conflate domain features with topological understanding, motivate our introduction of GraphAbstract. This benchmark evaluates models' ability to perceive global graph properties as humans do: recognizing organizational archetypes, detecting symmetry, sensing connectivity strength, and identifying critical elements. Our results reveal that vision models significantly outperform GNNs on tasks requiring holistic structural understanding and maintain generalizability across varying graph scales, while GNNs struggle with global pattern abstraction and degrade with increasing graph size. This work demonstrates that vision models possess remarkable yet underutilized capabilities for graph structural understanding, particularly for problems requiring global topological awareness and scale-invariant reasoning. These findings open new avenues to leverage this underappreciated potential for developing more effective graph foundation models for tasks dominated by holistic pattern recognition.

EAAI Journal 2025 Journal Article

User feedback information analysis based on collaborative filtering and revised rough numbers: A study on product green design elements extraction

  • Yan Xuan
  • Lei Zhang

In the era of the knowledge economy, user requirements (URs) have become one of the core driving forces that promote the development of enterprise product collaborative design. However, since user interests often drive URs and they are not professional designers, expecting them to translate requirements into design information is unrealistic. This paper presents an Artificial Intelligence (AI)-driven method to extract the product green design elements (PGDEs) from the user feedback preference information (UFPI). Firstly, sentiment analysis and Latent Dirichlet Allocation (LDA) topic modeling are implemented to determine the URs in online comments. Secondly, based on Large Language models (LLMs), specifically SparkDesk, the product green design requirements (PGDRs) that designers must consider during the product life cycle are generated. Third, the product green design elements (PGDEs) set composed of URs and PGDRs is constructed. Then, based on the collaborative filtering (CF) algorithm, the UFPI is obtained, and the essential elements of product green design are determined. Finally, based on the revised rough number theory, the decision and optimization of green design elements are completed. This study demonstrates the practical application of AI technologies in green design decision-making and provides theoretical and methodological support for the decision to PGDEs.

NeurIPS Conference 2025 Conference Paper

VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to Rank

  • Tianhe Wu
  • Jian Zou
  • Jie Liang
  • Lei Zhang
  • Kede Ma

DeepSeek-R1 has demonstrated remarkable effectiveness in incentivizing reasoning and generalization capabilities of large language models (LLMs) through reinforcement learning. Nevertheless, the potential of reasoning-induced computation has not been thoroughly explored in the context of image quality assessment (IQA), a task depending critically on visual reasoning. In this paper, we introduce VisualQuality-R1, a reasoning-induced no-reference IQA (NR-IQA) model, and we train it with reinforcement learning to rank, a learning algorithm tailored to the intrinsically relative nature of visual quality. Specifically, for a pair of images, we employ group relative policy optimization to generate multiple quality scores for each image. These estimates are used to compute comparative probabilities of one image having higher quality than the other under the Thurstone model. Rewards for each quality estimate are defined using continuous fidelity measures rather than discretized binary labels. Extensive experiments show that the proposed VisualQuality-R1 consistently outperforms discriminative deep learning-based NR-IQA models as well as a recent reasoning-induced quality regression method. Moreover, VisualQuality-R1 is capable of generating contextually rich, human-aligned quality descriptions, and supports multi-dataset training without requiring perceptual scale realignment. These features make VisualQuality-R1 especially well-suited for reliably measuring progress in a wide range of image processing tasks like super-resolution and image generation.

JBHI Journal 2025 Journal Article

VLD-Net: Localization and Detection of the Vertebrae From X-Ray Images by Reinforcement Learning With Adaptive Exploration Mechanism and Spine Anatomy Information

  • Shun Xiang
  • Lei Zhang
  • Yuanquan Wang
  • Shoujun Zhou
  • Xing Zhao
  • Tao Zhang
  • Shuo Li

Accurate and efficient vertebrae localization and detection in X-ray images are essential for diagnosing and treating spinal diseases. However, most existing methods struggle with the complexity of spine X-ray images, yielding inaccurate results due to insufficient utilization of spinal anatomy information and neglect of individual vertebra characteristics. In this paper, we propose an innovative Vertebrae Localization and Detection Network (VLD-Net) to accurately assist physicians in diagnosing spine-related diseases from X-ray images. Our VLD-Net, for the first time, defines vertebrae localization as a top-bottom sequential decision-making process, employing deep reinforcement learning (DRL) to fully leverage the anatomical information of the spine. Simultaneously, it also prioritizes the distinct characteristics of each vertebra for accurate detection. Specifically, VLD-Net combines three key components: 1) An advanced vertebrae localization module based on DRL is proposed, effectively leveraging anatomical information of the spine. 2) A novel adaptive exploration mechanism is coined to understand the behavior of the DRL agent during training, pinpointing how to effectively achieve the trade-off between exploration and exploitation. 3) An innovative vertebra-focused module is proposed to accurately detect vertebral landmarks, using the attention region of each vertebra as input to enhance focus on the target and reduce interference from surrounding tissue. Extensive experiments on two public spine datasets demonstrate that the VLD-Net outperforms the state-of-the-art methods in accuracy and robustness.

NeurIPS Conference 2024 Conference Paper

AdaNeg: Adaptive Negative Proxy Guided OOD Detection with Vision-Language Models

  • Yabin ZHANG
  • Lei Zhang

Recent research has shown that pre-trained vision-language models are effective at identifying out-of-distribution (OOD) samples by using negative labels as guidance. However, employing consistent negative labels across different OOD datasets often results in semantic misalignments, as these text labels may not accurately reflect the actual space of OOD images. To overcome this issue, we introduce \textit{adaptive negative proxies}, which are dynamically generated during testing by exploring actual OOD images, to align more closely with the underlying OOD label space and enhance the efficacy of negative proxy guidance. Specifically, our approach utilizes a feature memory bank to selectively cache discriminative features from test images, representing the targeted OOD distribution. This facilitates the creation of proxies that can better align with specific OOD datasets. While task-adaptive proxies average features to reflect the unique characteristics of each dataset, the sample-adaptive proxies weight features based on their similarity to individual test samples, exploring detailed sample-level nuances. The final score for identifying OOD samples integrates static negative labels with our proposed adaptive proxies, effectively combining textual and visual knowledge for enhanced performance. Our method is training-free and annotation-free, and it maintains fast testing speed. Extensive experiments across various benchmarks demonstrate the effectiveness of our approach, abbreviated as AdaNeg. Notably, on the large-scale ImageNet benchmark, our AdaNeg significantly outperforms existing methods, with a 2. 45\% increase in AUROC and a 6. 48\% reduction in FPR95. Codes are available at \url{https: //github. com/YBZh/OpenOOD-VLM}.

JBHI Journal 2024 Journal Article

Adaptive Annotation Correlation Based Multi-Annotation Learning for Calibrated Medical Image Segmentation

  • Wei Huang
  • Lei Zhang
  • Xin Shu
  • Zizhou Wang
  • Zhang Yi

Medical image segmentation is a fundamental task in many clinical applications, yet current automated segmentation methods rely heavily on manual annotations, which are inherently subjective and prone to annotation bias. Recently, modeling annotator preference has garnered great interest, and several methods have been proposed in the past two years. However, the existing methods completely ignore the potential correlation between annotations, such as complementary and discriminative information. In this work, the A daptive annotation C orrela T ion based mult I -ann O tation Lear N ing ( ACTION ) method is proposed for calibrated medical image segmentation. ACTION employs consensus feature learning and dynamic adaptive weighting to leverage complementary information across annotations and emphasize discriminative information within each annotation based on their correlations, respectively. Meanwhile, memory accumulation-replay is proposed to accumulate the prior knowledge and integrate it into the model to enable the model to accommodate the multi-annotation setting. Two medical image benchmarks with different modalities are utilized to evaluate the performance of ACTION, and extensive experimental results demonstrate that it achieves superior performance compared to several state-of-the-art methods.

ICML Conference 2024 Conference Paper

DNA-SE: Towards Deep Neural-Nets Assisted Semiparametric Estimation

  • Qinshuo Liu
  • Zixin Wang
  • Xi-An Li 0004
  • Xinyao Ji
  • Lei Zhang
  • Liu Lin
  • Zhonghua Liu

Semiparametric statistics play a pivotal role in a wide range of domains, including but not limited to missing data, causal inference, and transfer learning, to name a few. In many settings, semiparametric theory leads to (nearly) statistically optimal procedures that yet involve numerically solving Fredholm integral equations of the second kind. Traditional numerical methods, such as polynomial or spline approximations, are difficult to scale to multi-dimensional problems. Alternatively, statisticians may choose to approximate the original integral equations by ones with closed-form solutions, resulting in computationally more efficient, but statistically suboptimal or even incorrect procedures. To bridge this gap, we propose a novel framework by formulating the semiparametric estimation problem as a bi-level optimization problem; and then we propose a scalable algorithm called D eep N eural-Nets A ssisted S emiparametric E stimation ($\mathsf{DNA\mbox{-}SE}$) by leveraging the universal approximation property of Deep Neural-Nets (DNN) to streamline semiparametric procedures. Through extensive numerical experiments and a real data analysis, we demonstrate the numerical and statistical advantages of $\mathsf{DNA\mbox{-}SE}$ over traditional methods. To the best of our knowledge, we are the first to bring DNN into semiparametric statistics as a numerical solver of integral equations in our proposed general framework.

EAAI Journal 2024 Journal Article

Dynamic instance-aware layer-bit-select network on human activity recognition using wearable sensors

  • Nanfu Ye
  • Lei Zhang
  • Dongzhou Cheng
  • Can Bu
  • Songming Sun
  • Hao Wu
  • Aiguo Song

During recent years, deep convolutional neural networks have achieved remarkable success in a wide range of sensor-based human activity recognition (HAR) applications, which often require high computational cost and memory footprint, hence hindering practical HAR deployment on resource-limited mobile and wearable devices. Quantization has provided an effective solution to compress models and accelerate activity inference in real-world situations. However, most previous quantization schemes are static, which always utilize the same bit-width for all activity samples in a given layer. Intuitively, since activity samples are highly diverse according to their difficulty level, it is rather unrealistic to maintain a consistent bit-width quantization configuration for different activity samples. Based on dynamic quantization strategy, this paper introduces a novel Layer-Bit-Select Network named LBSNet to adaptively determine the optimal bit-widths of each convolutional layer according to the difficulty level of recognized activities. To achieve this goal, we design a lightweight Bit-selector, which is jointly optimized with a given main network. In such a way, easy activities such as sitting may be allocated to lower bit-widths, while high bit-widths may handle more complicated or hard activities like falls. Extensive experiments are conducted on several mainstream HAR benchmarks including WISDM, UCI-HAR, UniMiB-SHAR, and PAMAP2 to validate the effectiveness of our proposed approach. For instance, it can achieve round 5. 3 × speedup and 6. 2 × model size compression, with merely 0. 6% accuracy drop on WISDM dataset, compared to full-precision model. This approach has great potential to yield more computation-efficient and faster activity inference on mobile embedded platforms.

AAAI Conference 2024 Conference Paper

Dynamic Weighted Combiner for Mixed-Modal Image Retrieval

  • Fuxiang Huang
  • Lei Zhang
  • Xiaowei Fu
  • Suqi Song

Mixed-Modal Image Retrieval (MMIR) as a flexible search paradigm has attracted wide attention. However, previous approaches always achieve limited performance, due to two critical factors are seriously overlooked. 1) The contribution of image and text modalities is different, but incorrectly treated equally. 2) There exist inherent labeling noises in describing users' intentions with text in web datasets from diverse real-world scenarios, giving rise to overfitting. We propose a Dynamic Weighted Combiner (DWC) to tackle the above challenges, which includes three merits. First, we propose an Editable Modality De-equalizer (EMD) by taking into account the contribution disparity between modalities, containing two modality feature editors and an adaptive weighted combiner. Second, to alleviate labeling noises and data bias, we propose a dynamic soft-similarity label generator (SSG) to implicitly improve noisy supervision. Finally, to bridge modality gaps and facilitate similarity learning, we propose a CLIP-based mutual enhancement module alternately trained by a mixed-modality contrastive loss. Extensive experiments verify that our proposed model significantly outperforms state-of-the-art methods on real-world datasets. The source code is available at https://github.com/fuxianghuang1/DWC.

JBHI Journal 2024 Journal Article

Enhancing Motor Sequence Learning via Transcutaneous Auricular Vagus Nerve Stimulation (taVNS): An EEG Study

  • Long Chen
  • Chenghu Tang
  • Zhongpeng Wang
  • Lei Zhang
  • Bin Gu
  • Xiuyun Liu
  • Dong Ming

Motor learning plays a crucial role in human life, and various neuromodulation methods have been utilized to strengthen or improve it. Transcutaneous auricular vagus nerve stimulation (taVNS) has gained increasing attention due to its non-invasive nature, affordability and ease of implementation. Although the potential of taVNS on regulating motor learning has been suggested, its actual regulatory effect has yet been fully explored. Electroencephalogram (EEG) analysis provides an in-depth understanding of cognitive processes involved in motor learning so as to offer methodological support for regulation of motor learning. To investigate the effect of taVNS on motor learning, this study recruited 22 healthy subjects to participate a single-blind, sham-controlled, and within-subject serial reaction time task (SRTT) experiment. Every subject involved in two sessions at least one week apart and received a 20-minute active/sham taVNS in each session. Behavioral indicators as well as EEG characteristics during the task state, were extracted and analyzed. The results revealed that compared to the sham group, the active group showed higher learning performance. Additionally, the EEG results indicated that after taVNS, the motor-related cortical potential amplitudes and alpha-gamma modulation index decreased significantly and functional connectivity based on partial directed coherence towards frontal lobe was enhanced. These findings suggest that taVNS can improve motor learning, mainly through enhancing cognitive and memory functions rather than simple movement learning. This study confirms the positive regulatory effect of taVNS on motor learning, which is particularly promising as it offers a potential avenue for enhancing motor skills and facilitating rehabilitation.

AAAI Conference 2024 Conference Paper

Gradual Residuals Alignment: A Dual-Stream Framework for GAN Inversion and Image Attribute Editing

  • Hao Li
  • Mengqi Huang
  • Lei Zhang
  • Bo Hu
  • Yi Liu
  • Zhendong Mao

GAN-based image attribute editing firstly leverages GAN Inversion to project real images into the latent space of GAN and then manipulates corresponding latent codes. Recent inversion methods mainly utilize additional high-bit features to improve image details preservation, as low-bit codes cannot faithfully reconstruct source images, leading to the loss of details. However, during editing, existing works fail to accurately complement the lost details and suffer from poor editability. The main reason is they inject all the lost details indiscriminately at one time, which inherently induces the position and quantity of details to overfit source images, resulting in inconsistent content and artifacts in edited images. This work argues that details should be gradually injected into both the reconstruction and editing process in a multi-stage coarse-to-fine manner for better detail preservation and high editability. Therefore, a novel dual-stream framework is proposed to accurately complement details at each stage. The Reconstruction Stream is employed to embed coarse-to-fine lost details into residual features and then adaptively add them to the GAN generator. In the Editing Stream, residual features are accurately aligned by our Selective Attention mechanism and then injected into the editing process in a multi-stage manner. Extensive experiments have shown the superiority of our framework in both reconstruction accuracy and editing quality compared with existing methods.

NeurIPS Conference 2024 Conference Paper

Homology Consistency Constrained Efficient Tuning for Vision-Language Models

  • Huatian Zhang
  • Lei Zhang
  • Yongdong Zhang
  • Zhendong Mao

Efficient transfer learning has shown remarkable performance in tuning large-scale vision-language models (VLMs) toward downstream tasks with limited data resources. The key challenge of efficient transfer lies in adjusting image-text alignment to be task-specific while preserving pre-trained general knowledge. However, existing methods adjust image-text alignment merely on a set of observed samples, e. g. , data set and external knowledge base, which cannot guarantee to keep the correspondence of general concepts between image and text latent manifolds without being disrupted and thereby a weak generalization of the adjusted alignment. In this work, we propose a Homology Consistency (HC) constraint for efficient transfer on VLMs, which explicitly constrains the correspondence of image and text latent manifolds through structural equivalence based on persistent homology in downstream tuning. Specifically, we build simplicial complex on the top of data to mimic the topology of latent manifolds, then track the persistence of the homology classes of topological features across multiple scales, and guide the directions of persistence tracks in image and text manifolds to coincide each other, with a deviating perturbation additionally. For practical application, we tailor the implementation of our proposed HC constraint for two main paradigms of adapter tuning. Extensive experiments on few-shot learning over 11 datasets and domain generalization demonstrate the effectiveness and robustness of our method.

AAAI Conference 2024 Conference Paper

Identification of Necessary Semantic Undertakers in the Causal View for Image-Text Matching

  • Huatian Zhang
  • Lei Zhang
  • Kun Zhang
  • Zhendong Mao

Image-text matching bridges vision and language, which is a fundamental task in multimodal intelligence. Its key challenge lies in how to capture visual-semantic relevance. Fine-grained semantic interactions come from fragment alignments between image regions and text words. However, not all fragments contribute to image-text relevance, and many existing methods are devoted to mining the vital ones to measure the relevance accurately. How well image and text relate depends on the degree of semantic sharing between them. Treating the degree as an effect and fragments as its possible causes, we define those indispensable causes for the generation of the degree as necessary undertakers, i.e., if any of them did not occur, the relevance would be no longer valid. In this paper, we revisit image-text matching in the causal view and uncover inherent causal properties of relevance generation. Then we propose a novel theoretical prototype for estimating the probability-of-necessity of fragments, PN_f, for the degree of semantic sharing by means of causal inference, and further design a Necessary Undertaker Identification Framework (NUIF) for image-text matching, which explicitly formalizes the fragment's contribution to image-text relevance by modeling PN_f in two ways. Extensive experiments show our method achieves state-of-the-art on benchmarks Flickr30K and MSCOCO.

JBHI Journal 2024 Journal Article

Influence of Transcutaneous Vagus Nerve Stimulation on Motor Planning: A Resting-State and Task-State EEG Study

  • Long Chen
  • Jiatong He
  • Jiasheng Zhang
  • Zhongpeng Wang
  • Lei Zhang
  • Bin Gu
  • Xiuyun Liu
  • Dong Ming

Transcutaneous vagus nerve stimulation (tVNS) shows a potential regulatory role for motor planning. Still, existing research mainly focuses on behavioral studies, and the neural modulation mechanism needs to be clarified. Therefore, we designed a multi-condition (active or sham, pre or under, difficult or easy, left-hand or right-hand) motor planning experiment to explore the effect of online tVNS (i. e. , tVNS and tasks synchronized). Twenty-eight subjects were recruited and randomly assigned to active and sham groups. Both groups performed the same tasks in the experiment and separately collected task-state EEG and 5-min eye-open resting-state EEG. The results showed that the changes in event-related potential (ERP) and movement-related cortical potential (MRCP) amplitudes were more significant for the left-hand difficult task (LD) under active-tVNS. According to the power spectrum results, active-tVNS significantly modulated the activities of the contralateral motor cortex at beta and gamma bands in the resting state. The functional connectivity based on partial directed coherence (PDC) showed significant changes in the parietal lobe after active-tVNS. These findings suggest that tVNS is a promising way to improve motor planning ability.

NeurIPS Conference 2024 Conference Paper

Many-Shot In-Context Learning

  • Rishabh Agarwal
  • Avi Singh
  • Lei Zhang
  • Bernd Bohnet
  • Luis Rosias
  • Stephanie Chan
  • Biao Zhang
  • Ankesh Anand

Large language models (LLMs) excel at few-shot in-context learning (ICL) -- learning from a few examples provided in context at inference, without any weight updates. Newly expanded context windows allow us to investigate ICL with hundreds or thousands of examples – the many-shot regime. Going from few-shot to many-shot, we observe significant performance gains across a wide variety of generative and discriminative tasks. While promising, many-shot ICL can be bottlenecked by the available amount of human-generated outputs. To mitigate this limitation, we explore two new settings: (1) "Reinforced ICL" that uses model-generated chain-of-thought rationales in place of human rationales, and (2) "Unsupervised ICL" where we remove rationales from the prompt altogether, and prompts the model only with domain-specific inputs. We find that both Reinforced and Unsupervised ICL can be quite effective in the many-shot regime, particularly on complex reasoning tasks. We demonstrate that, unlike few-shot learning, many-shot learning is effective at overriding pretraining biases, can learn high-dimensional functions with numerical inputs, and performs comparably to supervised fine-tuning. Finally, we reveal the limitations of next-token prediction loss as an indicator of downstream ICL performance.

JBHI Journal 2024 Journal Article

MaskCAE: Masked Convolutional AutoEncoder via Sensor Data Reconstruction for Self-Supervised Human Activity Recognition

  • Dongzhou Cheng
  • Lei Zhang
  • Lutong Qin
  • Shuoyuan Wang
  • Hao Wu
  • Aiguo Song

Self-supervised Human Activity Recognition (HAR) has been gradually gaining a lot of attention in ubiquitous computing community. Its current focus primarily lies in how to overcome the challenge of manually labeling complicated and intricate sensor data from wearable devices, which is often hard to interpret. However, current self-supervised algorithms encounter three main challenges: performance variability caused by data augmentations in contrastive learning paradigm, limitations imposed by traditional self-supervised models, and the computational load deployed on wearable devices by current mainstream transformer encoders. To comprehensively tackle these challenges, this paper proposes a powerful self-supervised approach for HAR from a novel perspective of denoising autoencoder, the first of its kind to explore how to reconstruct masked sensor data built on a commonly employed, well-designed, and computationally efficient fully convolutional network. Extensive experiments demonstrate that our proposed Masked Convolutional AutoEncoder (MaskCAE) outperforms current state-of-the-art algorithms in self-supervised, fully supervised, and semi-supervised situations without relying on any data augmentations, which fills the gap of masked sensor data modeling in HAR area. Visualization analyses show that our MaskCAE could effectively capture temporal semantics in time series sensor data, indicating its great potential in modeling abstracted sensor data. An actual implementation is evaluated on an embedded platform.

NeurIPS Conference 2024 Conference Paper

Meta-Exploiting Frequency Prior for Cross-Domain Few-Shot Learning

  • Fei Zhou
  • Peng Wang
  • Lei Zhang
  • Zhenghua Chen
  • Wei Wei
  • Chen Ding
  • Guosheng Lin
  • Yanning Zhang

Meta-learning offers a promising avenue for few-shot learning (FSL), enabling models to glean a generalizable feature embedding through episodic training on synthetic FSL tasks in a source domain. Yet, in practical scenarios where the target task diverges from that in the source domain, meta-learning based method is susceptible to over-fitting. To overcome this, we introduce a novel framework, Meta-Exploiting Frequency Prior for Cross-Domain Few-Shot Learning, which is crafted to comprehensively exploit the cross-domain transferable image prior that each image can be decomposed into complementary low-frequency content details and high-frequency robust structural characteristics. Motivated by this insight, we propose to decompose each query image into its high-frequency and low-frequency components, and parallel incorporate them into the feature embedding network to enhance the final category prediction. More importantly, we introduce a feature reconstruction prior and a prediction consistency prior to separately encourage the consistency of the intermediate feature as well as the final category prediction between the original query image and its decomposed frequency components. This allows for collectively guiding the network's meta-learning process with the aim of learning generalizable image feature embeddings, while not introducing any extra computational cost in the inference phase. Our framework establishes new state-of-the-art results on multiple cross-domain few-shot learning benchmarks.

NeurIPS Conference 2024 Conference Paper

One-Step Effective Diffusion Network for Real-World Image Super-Resolution

  • Rongyuan Wu
  • Lingchen Sun
  • Zhiyuan Ma
  • Lei Zhang

The pre-trained text-to-image diffusion models have been increasingly employed to tackle the real-world image super-resolution (Real-ISR) problem due to their powerful generative image priors. Most of the existing methods start from random noise to reconstruct the high-quality (HQ) image under the guidance of the given low-quality (LQ) image. While promising results have been achieved, such Real-ISR methods require multiple diffusion steps to reproduce the HQ image, increasing the computational cost. Meanwhile, the random noise introduces uncertainty in the output, which is unfriendly to image restoration tasks. To address these issues, we propose a one-step effective diffusion network, namely OSEDiff, for the Real-ISR problem. We argue that the LQ image contains rich information to restore its HQ counterpart, and hence the given LQ image can be directly taken as the starting point for diffusion, eliminating the uncertainty introduced by random noise sampling. We finetune the pre-trained diffusion network with trainable layers to adapt it to complex image degradations. To ensure that the one-step diffusion model could yield HQ Real-ISR output, we apply variational score distillation in the latent space to conduct KL-divergence regularization. As a result, our OSEDiff model can efficiently and effectively generate HQ images in just one diffusion step. Our experiments demonstrate that OSEDiff achieves comparable or even better Real-ISR results, in terms of both objective metrics and subjective evaluations, than previous diffusion model-based Real-ISR methods that require dozens or hundreds of steps. The source codes are released at https: //github. com/cswry/OSEDiff.

JBHI Journal 2024 Journal Article

Predicting miRNA-Disease Associations Based on Spectral Graph Transformer With Dynamic Attention and Regularization

  • Zhengwei Li
  • Xu Bai
  • Ru Nie
  • Yanyan Liu
  • Lei Zhang
  • Zhuhong You

Extensive research indicates that microRNAs (miRNAs) play a crucial role in the analysis of complex human diseases. Recently, numerous methods utilizing graph neural networks have been developed to investigate the complex relationships between miRNAs and diseases. However, these methods often face challenges in terms of overall effectiveness and are sensitive to node positioning. To address these issues, the researchers introduce DARSFormer, an advanced deep learning model that integrates dynamic attention mechanisms with a spectral graph Transformer effectively. In the DARSFormer model, a miRNA-disease heterogeneous network is constructed initially. This network undergoes spectral decomposition into eigenvalues and eigenvectors, with the eigenvalue scalars being mapped into a vector space subsequently. An orthogonal graph neural network is employed to refine the parameter matrix. The enhanced features are then input into a graph Transformer, which utilizes a dynamic attention mechanism to amalgamate features by aggregating the enhanced neighbor features of miRNA and disease nodes. A projection layer is subsequently utilized to derive the association scores between miRNAs and diseases. The performance of DARSFormer in predicting miRNA-disease associations (MDAs) is exemplary. It achieves an AUC of 94. 18% in a five-fold cross-validation on the HMDD v2. 0 database. Similarly, on HMDD v3. 2, it records an AUC of 95. 27%. Case studies involving colorectal, esophageal, and prostate tumors confirm 27, 28, and 26 of the top 30 associated miRNAs against the dbDEMC and miR2Disease databases, respectively.

IJCAI Conference 2024 Conference Paper

Safeguarding Sustainable Cities: Unsupervised Video Anomaly Detection through Diffusion-based Latent Pattern Learning

  • Menghao Zhang
  • Jingyu Wang
  • Qi Qi
  • Pengfei Ren
  • Haifeng Sun
  • Zirui Zhuang
  • Lei Zhang
  • Jianxin Liao

Sustainable cities requires high-quality community management and surveillance analytics, which are supported by video anomaly detection techniques. However, mainstream video anomaly detection techniques still require manually labeled data and do not apply to real-world massive videos. Without labeling, unsupervised video anomaly detection (UVAD) is challenged by the problem of pseudo-labeled noise and the openness of anomaly detection. In response, a diffusion-based latent pattern learning UVAD framework is proposed, called DiffVAD. The method learns potential patterns by generating different patterns of the same event through diffusion models. The detection of anomalies is realized by evaluating the pattern distribution. The different patterns of normal events are diverse but correlated, while the different patterns of abnormal events are more diffuse. This manner of detection is equally effective for unseen normal events in the training set. In addition, we design a refinement strategy for pseudo-labels to mitigate the effects of the noise problem. Extensive experiments on six benchmark datasets demonstrate the design’s promising generalization ability and high efficiency. Specifically, DiffVAD obtains an AUC score of 81. 9% on the ShanghaiTech dataset.

ICML Conference 2024 Conference Paper

State-Constrained Zero-Sum Differential Games with One-Sided Information

  • Mukesh Ghimire
  • Lei Zhang
  • Zhe Xu 0005
  • Yi Ren

We study zero-sum differential games with state constraints and one-sided information, where the informed player (Player 1) has a categorical payoff type unknown to the uninformed player (Player 2). The goal of Player 1 is to minimize his payoff without violating the constraints, while that of Player 2 is to violate the state constraints if possible, or to maximize the payoff otherwise. One example of the game is a man-to-man matchup in football. Without state constraints, Cardaliaguet (2007) showed that the value of such a game exists and is convex to the common belief of players. Our theoretical contribution is an extension of this result to games with state constraints and the derivation of the primal and dual subdynamic principles necessary for computing behavioral strategies. Different from existing works that are concerned about the scalability of no-regret learning in games with discrete dynamics, our study reveals the underlying structure of strategies for belief manipulation resulting from information asymmetry and state constraints. This structure will be necessary for scalable learning on games with continuous actions and long time windows. We use a simplified football game to demonstrate the utility of this work, where we reveal player positions and belief states in which the attacker should (or should not) play specific random deceptive moves to take advantage of information asymmetry, and compute how the defender should respond.

ICLR Conference 2024 Conference Paper

Symbol as Points: Panoptic Symbol Spotting via Point-based Representation

  • Wenlong Liu
  • Tianyu Yang
  • Yuhan Wang
  • Qizhi Yu
  • Lei Zhang

This work studies the problem of panoptic symbol spotting, which is to spot and parse both countable object instances (windows, doors, tables, etc.) and uncountable stuff (wall, railing, etc.) from computer-aided design (CAD) drawings. Existing methods typically involve either rasterizing the vector graphics into images and using image-based methods for symbol spotting, or directly building graphs and using graph neural networks for symbol recognition. In this paper, we take a different approach, which treats graphic primitives as a set of 2D points that are locally connected and use point cloud segmentation methods to tackle it. Specifically, we utilize a point transformer to extract the primitive features and append a mask2former-like spotting head to predict the final output. To better use the local connection information of primitives and enhance their discriminability, we further propose the attention with connection module (ACM) and contrastive connection learning scheme (CCL). Finally, we propose a KNN interpolation mechanism for the mask attention module of the spotting head to better handle primitive mask downsampling, which is primitive-level in contrast to pixel-level for the image. Our approach, named SymPoint, is simple yet effective, outperforming recent state-of-the-art method GAT-CADNet by an absolute increase of 9.6% PQ and 10.4% RQ on the FloorPlanCAD dataset. The source code and models will be available at \url{https://github.com/nicehuster/SymPoint}.

NeurIPS Conference 2024 Conference Paper

TAPTRv2: Attention-based Position Update Improves Tracking Any Point

  • Hongyang Li
  • Hao Zhang
  • Shilong Liu
  • Zhaoyang Zeng
  • Feng Li
  • Tianhe Ren
  • Bohan Li
  • Lei Zhang

In this paper, we present TAPTRv2, a Transformer-based approach built upon TAPTR for solving the Tracking Any Point (TAP) task. TAPTR borrows designs from DEtection TRansformer (DETR) and formulates each tracking point as a point query, making it possible to leverage well-studied operations in DETR-like algorithms. TAPTRv2 improves TAPTR by addressing a critical issue regarding its reliance on cost-volume, which contaminates the point query’s content feature and negatively impacts both visibility prediction and cost-volume computation. In TAPTRv2, we propose a novel attention-based position update (APU) operation and use key-aware deformable attention to realize. For each query, this operation uses key-aware attention weights to combine their corresponding deformable sampling positions to predict a new query position. This design is based on the observation that local attention is essentially the same as cost-volume, both of which are computed by dot-production between a query and its surrounding features. By introducing this new operation, TAPTRv2 not only removes the extra burden of cost-volume computation, but also leads to a substantial performance improvement. TAPTRv2 surpasses TAPTR and achieves state-of-the-art performance on many challenging datasets, demonstrating the effectiveness of our approach.

NeurIPS Conference 2024 Conference Paper

Voxel Mamba: Group-Free State Space Models for Point Cloud based 3D Object Detection

  • Guowen Zhang
  • Lue Fan
  • Chenhang He
  • Zhen Lei
  • Zhaoxiang Zhang
  • Lei Zhang

Serialization-based methods, which serialize the 3D voxels and group them into multiple sequences before inputting to Transformers, have demonstrated their effectiveness in 3D object detection. However, serializing 3D voxels into 1D sequences will inevitably sacrifice the voxel spatial proximity. Such an issue is hard to be addressed by enlarging the group size with existing serialization-based methods due to the quadratic complexity of Transformers with feature sizes. Inspired by the recent advances of state space models (SSMs), we present a Voxel SSM, termed as Voxel Mamba, which employs a group-free strategy to serialize the whole space of voxels into a single sequence. The linear complexity of SSMs encourages our group-free design, alleviating the loss of spatial proximity of voxels. To further enhance the spatial proximity, we propose a Dual-scale SSM Block to establish a hierarchical structure, enabling a larger receptive field in the 1D serialization curve, as well as more complete local regions in 3D space. Moreover, we implicitly apply window partition under the group-free framework by positional encoding, which further enhances spatial proximity by encoding voxel positional information. Our experiments on Waymo Open Dataset and nuScenes dataset show that Voxel Mamba not only achieves higher accuracy than state-of-the-art methods, but also demonstrates significant advantages in computational efficiency. The source code is available at https: //github. com/gwenzhang/Voxel-Mamba.

NeurIPS Conference 2023 Conference Paper

A Comprehensive Benchmark for Neural Human Radiance Fields

  • Kenkun Liu
  • Derong Jin
  • Ailing Zeng
  • Xiaoguang Han
  • Lei Zhang

The past two years have witnessed a significant increase in interest concerning NeRF-based human body rendering. While this surge has propelled considerable advancements, it has also led to an influx of methods and datasets. This explosion complicates experimental settings and makes fair comparisons challenging. In this work, we design and execute thorough studies into unified evaluation settings and metrics to establish a fair and reasonable benchmark for human NeRF models. To reveal the effects of extant models, we benchmark them against diverse and hard scenes. Additionally, we construct a cross-subject benchmark pre-trained on large-scale datasets to assess generalizable methods. Finally, we analyze the essential components for animatability and generalizability, and make HumanNeRF from monocular videos generalizable, as the inaugural baseline. We hope these benchmarks and analyses could serve the community.

EAAI Journal 2023 Journal Article

A survey on the mechanism and countermeasures of low-frequency swaying of high-speed trains caused by aerodynamic loads

  • Chao Chang
  • Xin Ding
  • Zhuang Sun
  • Yizheng Yu
  • Lei Zhang

The carbody abnormal vibration has significant impacts on the comfort and safety of high-speed train. Field measurements were conducted to study the low-frequency swaying of the carbody on a high-speed train operating on a railway line. According to the test results, the tail vehicle of the train swayed laterally in various places along the line. The lateral stability index of vehicle clearly exceeded the limit value when the carbody swayed. The predominant frequency of the lateral acceleration of the carbody was between 1. 4 and 1. 5 Hz. Through the simulation of computational fluid dynamics, it is confirmed that the yaw moment and lift force of the tail vehicle are greater than those of the middle and the head vehicles. Furthermore, the dynamics simulation show that aerodynamics disturbances may be the intensified cause of the abnormal swaying of the tail vehicle. Therefore, the study proposes the installation of inter-vehicle dampers. and optimizes the damping values of dampers based on GA-BP optimization algorithm, so as to weaken the abnormal swaying motion. The relevant simulation results offer a solution to address the abnormal swaying phenomenon in practical situations.

ICRA Conference 2023 Conference Paper

Approximating Discontinuous Nash Equilibrial Values of Two-Player General-Sum Differential Games

  • Lei Zhang
  • Mukesh Ghimire
  • Wenlong Zhang
  • Zhe Xu 0005
  • Yi Ren

Finding Nash equilibrial policies for two-player differential games requires solving Hamilton-Jacobi-Isaacs (HJI) PDEs. Self-supervised learning has been used to approximate solutions of such PDEs while circumventing the curse of dimensionality. However, this method fails to learn discontinuous PDE solutions due to its sampling nature, leading to poor safety performance of the resulting controllers in robotics applications when player rewards are discontinuous. This paper investigates two potential solutions to this problem: a hybrid method that leverages both supervised Nash equilibria and the HJI PDE, and a value-hardening method where a sequence of HJIs are solved with a gradually hardening reward. We compare these solutions using the resulting generalization and safety performance in two vehicle interaction simulation studies with 5D and 9D state spaces, respectively. Results show that with informative supervision (e. g. , collision and near-collision demonstrations) and the low cost of self-supervised learning, the hybrid method achieves better safety performance than the supervised, self-supervised, and value hardening approaches on equal computational budget. Value hardening fails to generalize in the higher-dimensional case without informative supervision. Lastly, we show that the neural activation function needs to be continuously differentiable for learning PDEs and its choice can be case dependent.

AAAI Conference 2023 Conference Paper

Are Transformers Effective for Time Series Forecasting?

  • Ailing Zeng
  • Muxi Chen
  • Lei Zhang
  • Qiang Xu

Recently, there has been a surge of Transformer-based solutions for the long-term time series forecasting (LTSF) task. Despite the growing performance over the past few years, we question the validity of this line of research in this work. Specifically, Transformers is arguably the most successful solution to extract the semantic correlations among the elements in a long sequence. However, in time series modeling, we are to extract the temporal relations in an ordered set of continuous points. While employing positional encoding and using tokens to embed sub-series in Transformers facilitate preserving some ordering information, the nature of the permutation-invariant self-attention mechanism inevitably results in temporal information loss. To validate our claim, we introduce a set of embarrassingly simple one-layer linear models named LTSF-Linear for comparison. Experimental results on nine real-life datasets show that LTSF-Linear surprisingly outperforms existing sophisticated Transformer-based LTSF models in all cases, and often by a large margin. Moreover, we conduct comprehensive empirical studies to explore the impacts of various design elements of LTSF models on their temporal relation extraction capability. We hope this surprising finding opens up new research directions for the LTSF task. We also advocate revisiting the validity of Transformer-based solutions for other time series analysis tasks (e.g., anomaly detection) in the future.

JBHI Journal 2023 Journal Article

Convolutional Feature Descriptor Selection for Mammogram Classification

  • Dong Li
  • Lei Zhang
  • Jianwei Zhang
  • Xingyu Xie

Breast cancer was the most commonly diagnosed cancer among women worldwide in 2020. Recently, several deep learning-based classification approaches have been proposed to screen breast cancer in mammograms. However, most of these approaches require additional detection or segmentation annotations. Meanwhile, some other image-level label-based methods often pay insufficient attention to lesion areas, which are critical for diagnosis. This study designs a novel deep-learning method for automatically diagnosing breast cancer in mammography, which focuses on the local lesion areas and only utilizes image-level classification labels. In this study, we propose to select discriminative feature descriptors from feature maps instead of identifying lesion areas using precise annotations. And we design a novel adaptive convolutional feature descriptor selection (AFDS) structure based on the distribution of the deep activation map. Specifically, we adopt the triangle threshold strategy to calculate a specific threshold for guiding the activation map to determine which feature descriptors (local areas) are discriminative. Ablation experiments and visualization analysis indicate that the AFDS structure makes the model easier to learn the difference between malignant and benign/normal lesions. Furthermore, since the AFDS structure can be regarded as a highly efficient pooling structure, it can be easily plugged into most existing convolutional neural networks with negligible effort and time consumption. Experimental results on two publicly available INbreast and CBIS-DDSM datasets indicate that the proposed method performs satisfactorily compared with state-of-the-art methods.

EAAI Journal 2023 Journal Article

Design of concrete incorporating microencapsulated phase change materials for clean energy: A ternary machine learning approach based on generative adversarial networks

  • Afshin Marani
  • Lei Zhang
  • Moncef L. Nehdi

The inclusion of microencapsulated phase change materials (MPCM) in construction materials is a promising solution for increasing the energy efficiency of buildings and reducing their carbon emissions. Although MPCMs provide thermal energy storage capability in concrete, they typically decrease its compressive strength. A unified framework for the mixture design of concrete incorporating MPCM is yet to be developed to facilitate practical applications. This study proposes a mix design procedure using a novel ternary machine learning (ML) paradigm. For this purpose, the tabular generative adversarial network (TGAN) was utilized to generate large synthetic mixture design data based on the limited available experimental observations. The synthetic data is then employed to construct robust predictive ML models. The gradient boosting regressor (GBR) model trained with synthetic data outperformed the model trained with real data, achieving a testing coefficient of determination (R2) of 0. 963 and mean absolute error (MAE) of 2. 085 MPa. The TGAN-GBR model was ultimately integrated with the particle swarm optimization (PSO) algorithm to construct a powerful recommendation system for optimizing the mixture design of concrete and mortar incorporating different types of MPCMs. Extensive parametric analyses along with the employed optimization procedure accomplished the mixture design of latent heat thermal energy storage concrete with maximum MPCM inclusion and minimum cement content for various compressive strength classes. The proposed framework enables energy conservation technology in the design of eco-friendly building materials with acceptable mechanical performance.

AAAI Conference 2023 Conference Paper

DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and Grounding

  • Shilong Liu
  • Shijia Huang
  • Feng Li
  • Hao Zhang
  • Yaoyuan Liang
  • Hang Su
  • Jun Zhu
  • Lei Zhang

In this paper, we study the problem of visual grounding by considering both phrase extraction and grounding (PEG). In contrast to the previous phrase-known-at-test setting, PEG requires a model to extract phrases from text and locate objects from image simultaneously, which is a more practical setting in real applications. As phrase extraction can be regarded as a 1D text segmentation problem, we formulate PEG as a dual detection problem and propose a novel DQ-DETR model, which introduces dual queries to probe different features from image and text for object prediction and phrase mask prediction. Each pair of dual queries are designed to have shared positional parts but different content parts. Such a design effectively alleviates the difficulty of modality alignment between image and text (in contrast to a single query design) and empowers Transformer decoder to leverage phrase mask-guided attention to improve the performance. To evaluate the performance of PEG, we also propose a new metric CMAP (cross-modal average precision), analogous to the AP metric in object detection. The new metric overcomes the ambiguity of Recall@1 in many-box-to-one-phrase cases in phrase grounding. As a result, our PEG pre-trained DQ-DETR establishes new state-of-the-art results on all visual grounding benchmarks with a ResNet-101 backbone. For example, it achieves 91.04% and 83.51% in terms of recall rate on RefCOCO testA and testB with a ResNet-101 backbone.

NeurIPS Conference 2023 Conference Paper

DreamWaltz: Make a Scene with Complex 3D Animatable Avatars

  • Yukun Huang
  • Jianan Wang
  • Ailing Zeng
  • He CAO
  • Xianbiao Qi
  • Yukai Shi
  • Zheng-Jun Zha
  • Lei Zhang

We present DreamWaltz, a novel framework for generating and animating complex 3D avatars given text guidance and parametric human body prior. While recent methods have shown encouraging results for text-to-3D generation of common objects, creating high-quality and animatable 3D avatars remains challenging. To create high-quality 3D avatars, DreamWaltz proposes 3D-consistent occlusion-aware Score Distillation Sampling (SDS) to optimize implicit neural representations with canonical poses. It provides view-aligned supervision via 3D-aware skeleton conditioning which enables complex avatar generation without artifacts and multiple faces. For animation, our method learns an animatable 3D avatar representation from abundant image priors of diffusion model conditioned on various poses, which could animate complex non-rigged avatars given arbitrary poses without retraining. Extensive evaluations demonstrate that DreamWaltz is an effective and robust approach for creating 3D avatars that can take on complex shapes and appearances as well as novel poses for animation. The proposed framework further enables the creation of complex scenes with diverse compositions, including avatar-avatar, avatar-object and avatar-scene interactions. See https: //dreamwaltz3d. github. io/ for more vivid 3D avatar and animation results.

AAAI Conference 2023 Conference Paper

DRGCN: Dynamic Evolving Initial Residual for Deep Graph Convolutional Networks

  • Lei Zhang
  • Xiaodong Yan
  • Jianshan He
  • Ruopeng Li
  • Wei Chu

Graph convolutional networks (GCNs) have been proved to be very practical to handle various graph-related tasks. It has attracted considerable research interest to study deep GCNs, due to their potential superior performance compared with shallow ones. However, simply increasing network depth will, on the contrary, hurt the performance due to the over-smoothing problem. Adding residual connection is proved to be effective for learning deep convolutional neural networks (deep CNNs), it is not trivial when applied to deep GCNs. Recent works proposed an initial residual mechanism that did alleviate the over-smoothing problem in deep GCNs. However, according to our study, their algorithms are quite sensitive to different datasets. In their setting, the personalization (dynamic) and correlation (evolving) of how residual applies are ignored. To this end, we propose a novel model called Dynamic evolving initial Residual Graph Convolutional Network (DRGCN). Firstly, we use a dynamic block for each node to adaptively fetch information from the initial representation. Secondly, we use an evolving block to model the residual evolving pattern between layers. Our experimental results show that our model effectively relieves the problem of over-smoothing in deep GCNs and outperforms the state-of-the-art (SOTA) methods on various benchmark datasets. Moreover, we develop a mini-batch version of DRGCN which can be applied to large-scale data. Coupling with several fair training techniques, our model reaches new SOTA results on the large-scale ogbn-arxiv dataset of Open Graph Benchmark (OGB). Our reproducible code is available on GitHub.

JBHI Journal 2023 Journal Article

FreqSense: Adaptive Sampling Rates for Sensor-Based Human Activity Recognition Under Tunable Computational Budgets

  • Guangyu Yang
  • Lei Zhang
  • Can Bu
  • Shuaishuai Wang
  • Hao Wu
  • Aiguo Song

Recent years have witnessed great success of deep convolutional networks in sensor-based human activity recognition (HAR), yet their practical deployment remains a challenge due to the varying computational budgets required to obtain a reliable prediction. This article focuses on adaptive inference from a novel perspective of signal frequency, which is motivated by an intuition that low-frequency features are enough for recognizing “easy” activity samples, while only “hard” activity samples need temporally detailed information. We propose an adaptive resolution network by combining a simple subsampling strategy with conditional early-exit. Specifically, it is comprised of multiple subnetworks with different resolutions, where “easy” activity samples are first classified by lightweight subnetwork using the lowest sampling rate, while the subsequent subnetworks in higher resolution would be sequentially applied once the former one fails to reach a confidence threshold. Such dynamical decision process could adaptively select a proper sampling rate for each activity sample conditioned on an input if the budget varies, which will be terminated until enough confidence is obtained, hence avoiding excessive computations. Comprehensive experiments on four diverse HAR benchmark datasets demonstrate the effectiveness of our method in terms of accuracy-cost tradeoff. We benchmark the average latency on a real hardware.

NeurIPS Conference 2023 Conference Paper

Label-efficient Segmentation via Affinity Propagation

  • Wentong Li
  • Yuqian Yuan
  • Song Wang
  • Wenyu Liu
  • Dongqi Tang
  • Jian Liu
  • Jianke Zhu
  • Lei Zhang

Weakly-supervised segmentation with label-efficient sparse annotations has attracted increasing research attention to reduce the cost of laborious pixel-wise labeling process, while the pairwise affinity modeling techniques play an essential role in this task. Most of the existing approaches focus on using the local appearance kernel to model the neighboring pairwise potentials. However, such a local operation fails to capture the long-range dependencies and ignores the topology of objects. In this work, we formulate the affinity modeling as an affinity propagation process, and propose a local and a global pairwise affinity terms to generate accurate soft pseudo labels. An efficient algorithm is also developed to reduce significantly the computational cost. The proposed approach can be conveniently plugged into existing segmentation networks. Experiments on three typical label-efficient segmentation tasks, i. e. box-supervised instance segmentation, point/scribble-supervised semantic segmentation and CLIP-guided semantic segmentation, demonstrate the superior performance of the proposed approach.

AAAI Conference 2023 Conference Paper

Mind the Gap: Polishing Pseudo Labels for Accurate Semi-supervised Object Detection

  • Lei Zhang
  • Yuxuan Sun
  • Wei Wei

Exploiting pseudo labels (e.g., categories and bounding boxes) of unannotated objects produced by a teacher detector have underpinned much of recent progress in semi-supervised object detection (SSOD). However, due to the limited generalization capacity of the teacher detector caused by the scarce annotations, the produced pseudo labels often deviate from ground truth, especially those with relatively low classification confidences, thus limiting the generalization performance of SSOD. To mitigate this problem, we propose a dual pseudo-label polishing framework for SSOD. Instead of directly exploiting the pseudo labels produced by the teacher detector, we take the first attempt at reducing their deviation from ground truth using dual polishing learning, where two differently structured polishing networks are elaborately developed and trained using synthesized paired pseudo labels and the corresponding ground truth for categories and bounding boxes on the given annotated objects, respectively. By doing this, both polishing networks can infer more accurate pseudo labels for unannotated objects through sufficiently exploiting their context knowledge based on the initially produced pseudo labels, and thus improve the generalization performance of SSOD. Moreover, such a scheme can be seamlessly plugged into the existing SSOD framework for joint end-to-end learning. In addition, we propose to disentangle the polished pseudo categories and bounding boxes of unannotated objects for separate category classification and bounding box regression in SSOD, which enables introducing more unannotated objects during model training and thus further improves the performance. Experiments on both PASCAL VOC and MS-COCO benchmarks demonstrate the superiority of the proposed method over existing state-of-the-art baselines. The code can be found at https://github.com/snowdusky/DualPolishLearning.

AAAI Conference 2023 Conference Paper

MMTN: Multi-Modal Memory Transformer Network for Image-Report Consistent Medical Report Generation

  • Yiming Cao
  • Lizhen Cui
  • Lei Zhang
  • Fuqiang Yu
  • Zhen Li
  • Yonghui Xu

Automatic medical report generation is an essential task in applying artificial intelligence to the medical domain, which can lighten the workloads of doctors and promote clinical automation. The state-of-the-art approaches employ Transformer-based encoder-decoder architectures to generate reports for medical images. However, they do not fully explore the relationships between multi-modal medical data, and generate inaccurate and inconsistent reports. To address these issues, this paper proposes a Multi-modal Memory Transformer Network (MMTN) to cope with multi-modal medical data for generating image-report consistent medical reports. On the one hand, MMTN reduces the occurrence of image-report inconsistencies by designing a unique encoder to associate and memorize the relationship between medical images and medical terminologies. On the other hand, MMTN utilizes the cross-modal complementarity of the medical vision and language for the word prediction, which further enhances the accuracy of generating medical reports. Extensive experiments on three real datasets show that MMTN achieves significant effectiveness over state-of-the-art approaches on both automatic metrics and human evaluation.

NeurIPS Conference 2023 Conference Paper

MomentDiff: Generative Video Moment Retrieval from Random to Real

  • Pandeng Li
  • Chen-Wei Xie
  • Hongtao Xie
  • Liming Zhao
  • Lei Zhang
  • Yun Zheng
  • Deli Zhao
  • Yongdong Zhang

Video moment retrieval pursues an efficient and generalized solution to identify the specific temporal segments within an untrimmed video that correspond to a given language description. To achieve this goal, we provide a generative diffusion-based framework called MomentDiff, which simulates a typical human retrieval process from random browsing to gradual localization. Specifically, we first diffuse the real span to random noise, and learn to denoise the random noise to the original span with the guidance of similarity between text and video. This allows the model to learn a mapping from arbitrary random locations to real moments, enabling the ability to locate segments from random initialization. Once trained, MomentDiff could sample random temporal segments as initial guesses and iteratively refine them to generate an accurate temporal boundary. Different from discriminative works (e. g. , based on learnable proposals or queries), MomentDiff with random initialized spans could resist the temporal location biases from datasets. To evaluate the influence of the temporal location biases, we propose two ``anti-bias'' datasets with location distribution shifts, named Charades-STA-Len and Charades-STA-Mom. The experimental results demonstrate that our efficient framework consistently outperforms state-of-the-art methods on three public benchmarks, and exhibits better generalization and robustness on the proposed anti-bias datasets. The code, model, and anti-bias evaluation datasets will be released publicly.

NeurIPS Conference 2023 Conference Paper

Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset

  • Jing Lin
  • Ailing Zeng
  • Shunlin Lu
  • Yuanhao Cai
  • Ruimao Zhang
  • Haoqian Wang
  • Lei Zhang

In this paper, we present Motion-X, a large-scale 3D expressive whole-body motion dataset. Existing motion datasets predominantly contain body-only poses, lacking facial expressions, hand gestures, and fine-grained pose descriptions. Moreover, they are primarily collected from limited laboratory scenes with textual descriptions manually labeled, which greatly limits their scalability. To overcome these limitations, we develop a whole-body motion and text annotation pipeline, which can automatically annotate motion from either single- or multi-view videos and provide comprehensive semantic labels for each video and fine-grained whole-body pose descriptions for each frame. This pipeline is of high precision, cost-effective, and scalable for further research. Based on it, we construct Motion-X, which comprises 15. 6M precise 3D whole-body pose annotations (i. e. , SMPL-X) covering 81. 1K motion sequences from massive scenes. Besides, Motion-X provides 15. 6M frame-level whole-body pose descriptions and 81. 1K sequence-level semantic labels. Comprehensive experiments demonstrate the accuracy of the annotation pipeline and the significant benefit of Motion-X in enhancing expressive, diverse, and natural motion generation, as well as 3D whole-body human mesh recovery.

JBHI Journal 2023 Journal Article

ProtoHAR: Prototype Guided Personalized Federated Learning for Human Activity Recognition

  • Dongzhou Cheng
  • Lei Zhang
  • Can Bu
  • Xing Wang
  • Hao Wu
  • Aiguo Song

Federated Learning (FL) has recently attracted great interest in sensor-based human activity recognition (HAR) tasks. However, in real-world environment, sensor data on devices is non-independently and identically distributed (Non-IID), e. g. , activity data recorded by most devices is sparse, and sensor data distribution for each client may be inconsistent. As a result, the traditional FL methods in the heterogeneous environment may incur a drifted global model that causes slow convergence and a heavy communication burden. Although some FL methods are gradually being applied to HAR, they are designed for overly ideal scenarios and do not address such Non-IID problem in the real-world setting. It is still a question whether they can be applied to cross-device FL. To tackle this challenge, we propose ProtoHAR, a prototype-guided FL framework for HAR, which aims to decouple the representation and classifier in the heterogeneous FL setting efficiently. It leverages the global prototype to correct the activity feature representation to make the prototype knowledge flow among clients without leaking privacy while solving a better classifier to avoid excessive drift of the local model in personalized training. Extensive experiments are conducted on four publicly available datasets: USC-HAD, UNIMIB-SHAR, PAMAP2, and HARBOX, which are collected in both controlled environments and real-world scenarios. The results show that compared with the state-of-the-art FL algorithms, ProtoHAR achieves the best performance and faster convergence speed in HAR datasets.

AAAI Conference 2023 Conference Paper

Revisiting Unsupervised Local Descriptor Learning

  • Wufan Wang
  • Lei Zhang
  • Hua Huang

Constructing accurate training tuples is crucial for unsupervised local descriptor learning, yet challenging due to the absence of patch labels. The state-of-the-art approach constructs tuples with heuristic rules, which struggle to precisely depict real-world patch transformations, in spite of enabling fast model convergence. A possible solution to alleviate the problem is the clustering-based approach, which can capture realistic patch variations and learn more accurate class decision boundaries, but suffers from slow model convergence. This paper presents HybridDesc, an unsupervised approach that learns powerful local descriptor models with fast convergence speed by combining the rule-based and clustering-based approaches to construct training tuples. In addition, HybridDesc also contributes two concrete enhancing mechanisms: (1) a Differentiable Hyperparameter Search (DHS) strategy to find the optimal hyperparameter setting of the rule-based approach so as to provide accurate prior for the clustering-based approach, (2) an On-Demand Clustering (ODC) method to reduce the clustering overhead of the clustering-based approach without eroding its advantage. Extensive experimental results show that HybridDesc can efficiently learn local descriptors that surpass existing unsupervised local descriptors and even rival competitive supervised ones.

JBHI Journal 2023 Journal Article

SA-RPN: A Spacial Aware Region Proposal Network for Acne Detection

  • Jianwei Zhang
  • Lei Zhang
  • Junyou Wang
  • Xin Wei
  • Jiaqi Li
  • Xian Jiang
  • Dan Du

Automated detection of skin lesions offers excellent potential for interpretative diagnosis and precise treatment of acne vulgar. However, the blurry boundary and small size of lesions make it challenging to detect acne lesions with traditional object detection methods. To better understand the acne detection task, we construct a new benchmark dataset named AcneSCU, consisting of 276 facial images with 31777 instance-level annotations from clinical dermatology. To the best of our knowledge, AcneSCU is the first acne dataset with high-resolution imageries, precise annotations, and fine-grained lesion categories, which enables the comprehensive study of acne detection. More importantly, we propose a novel method called Spatial Aware Region Proposal Network (SA-RPN) to improve the proposal quality of two-stage detection methods. Specifically, the representation learning for the classification and localization task is disentangled with a double head component to promote the proposals for hard samples. Then, Normalized Wasserstein Distance of each proposal is predicted to improve the correlation between the classification scores and the proposals' intersection-over-unions (IoUs). SA-RPN can serve as a plug-and-play module to enhance standard two-stage detectors. Extensive experiments are conducted on both AcneSCU and the public dataset ACNE04, and the results show that the proposed method can consistently outperform state-of-the-art methods.

NeurIPS Conference 2023 Conference Paper

Semi-Supervised Domain Generalization with Known and Unknown Classes

  • Lei Zhang
  • Ji-Fu Li
  • Wei Wang

Semi-Supervised Domain Generalization (SSDG) aims to learn a model that is generalizable to an unseen target domain with only a few labels, and most existing SSDG methods assume that unlabeled training and testing samples are all known classes. However, a more realistic scenario is that known classes may be mixed with some unknown classes in unlabeled training and testing data. To deal with such a scenario, we propose the Class-Wise Adaptive Exploration and Exploitation (CWAEE) method. In particular, we explore unlabeled training data by using one-vs-rest classifiers and class-wise adaptive thresholds to detect known and unknown classes, and exploit them by adopting consistency regularization on augmented samples based on Fourier Transformation to improve the unseen domain generalization. The experiments conducted on real-world datasets verify the effectiveness and superiority of our method.

NeurIPS Conference 2023 Conference Paper

SMPLer-X: Scaling Up Expressive Human Pose and Shape Estimation

  • Zhongang Cai
  • Wanqi Yin
  • Ailing Zeng
  • Chen Wei
  • Qingping SUN
  • Wang Yanjun
  • Hui En Pang
  • Haiyi Mei

Expressive human pose and shape estimation (EHPS) unifies body, hands, and face motion capture with numerous applications. Despite encouraging progress, current state-of-the-art methods still depend largely on a confined set of training datasets. In this work, we investigate scaling up EHPS towards the first generalist foundation model (dubbed SMPLer-X), with up to ViT-Huge as the backbone and training with up to 4. 5M instances from diverse data sources. With big data and the large model, SMPLer-X exhibits strong performance across diverse test benchmarks and excellent transferability to even unseen environments. 1) For the data scaling, we perform a systematic investigation on 32 EHPS datasets, including a wide range of scenarios that a model trained on any single dataset cannot handle. More importantly, capitalizing on insights obtained from the extensive benchmarking process, we optimize our training scheme and select datasets that lead to a significant leap in EHPS capabilities. 2) For the model scaling, we take advantage of vision transformers to study the scaling law of model sizes in EHPS. Moreover, our finetuning strategy turn SMPLer-X into specialist models, allowing them to achieve further performance boosts. Notably, our foundation model SMPLer-X consistently delivers state-of-the-art results on seven benchmarks such as AGORA (107. 2 mm NMVE), UBody (57. 4 mm PVE), EgoBody (63. 6 mm PVE), and EHF (62. 3 mm PVE without finetuning).

EAAI Journal 2022 Journal Article

A novel fractional time-delayed grey Bernoulli forecasting model and its application for the energy production and consumption prediction

  • Yong Wang
  • Xinbo He
  • Lei Zhang
  • Xin Ma
  • Wenqing Wu
  • Rui Nie
  • Pei Chi
  • Yuyang Zhang

Energy affects the stable and sustainable development of social economy. Energy prediction plays an important role in the process of China’s energy market transformation. Scientific and reasonable energy predicting method can help government to make decisions effectively, and then adjust energy structure and industrial layout. The energy field is full of fractional order phenomenon and nonlinear disturbance. Aiming at the energy data sets with the characteristics of scarcity, complexity and nonlinear, a mathematical model including time delay term and Bernoulli equation can be used to fit this trend. A new fractional time-delayed grey Bernoulli model is proposed, and the new model has a wider application in the nonlinear field. The model is discretized by integral, and the least square estimation of the linear parameters and the approximate time response equation are obtained. The Grey Wolf Optimizer (GWO) is used to search the optimal parameters of the model. In addition, the energy prediction model is established from the perspective of renewable energy and fossil energy, and the effectiveness of the model is verified by three actual cases of renewable energy, crude oil and fossil fuel. Compared with the other seven grey models, the results show that the new model has higher prediction performance. Finally, the energy development trend in the next few years is predicted by using the proposed model, and relevant conclusions are drawn according to the prediction results.

EAAI Journal 2022 Journal Article

A novel self-adaptive fractional multivariable grey model and its application in forecasting energy production and conversion of China

  • Yong Wang
  • Li Wang
  • Lingling Ye
  • Xin Ma
  • Wenqing Wu
  • Zhongsen Yang
  • Xinbo He
  • Lei Zhang

Energy production and conversion have a significant impact on the economic development of all countries in the world. China’s energy production and conversion are large. Therefore, accurate mid-to-long term China’s energy production and conversion forecasting is becoming more and more important for integrating energy systems and energy strategic planning. For this purpose, a novel fractional grey sequence is proposed based on Grunwald–Letnikov fractional calculus. Furthermore, a novel self-adaptive fractional multivariable grey model is proposed based on the novel sequence. In this article, we compare several classical optimization algorithms and finally choose Particle Swarm Optimization (PSO) to compute the parameters. In addition, Monte-Carlo simulation and probability density analysis (PDA) are presented in this article to verify the model’s performance. Monte-Carlo simulation reduces the randomness of the results of the model runs to a certain extent. Probability density analysis visualizes this randomness through kernel density estimation (KDE). This paper compares the new model with the existing seven grey models and predicts the total energy consumption per capita, energy conversion efficiency and total renewable energy in China, respectively. The experimental results show that the new model is superior to the other seven models in terms of stability and prediction accuracy.

TCS Journal 2022 Journal Article

An adaptation-complete proof system for local reasoning about cloud storage systems

  • Zhao Jin
  • Bowen Zhang
  • Lei Zhang
  • Yongzhi Cao
  • Hanpin Wang

The rapid growth of data presents a significant challenge to the capability of traditional storage technologies to collect and manage data. Cloud storage systems (CSSs) have been proposed as a method to improve storage capacity. To safely and effectively manage cloud storage data and improve data service quality, it is necessary to verify the correctness of CSS management programs. However, the complexity of these systems renders program verification difficult. In this paper, we propose a Hoare-style proof system, in conjunction with two languages, to analyze and verify CSS management programs. The first is a modeling language that describes the program execution. The second is an assertion language based on Separation Logic (SL), used to describe the properties of the CSS file-block-location storage structure. The proof system supports modular local reasoning for CSS programs by a set of adaptation rules, which enable the condition of specifications to be applied to broader contexts. A key question that arises is whether the proof system can meet adaptation completeness. If so, arbitrary satisfiable specifications can be adjusted using the adaptation rules. To this end, we developed local predicate transformers and used their domain to interpret all types of commands. By finding the smallest local predicate transformer, we established adaptation completeness. In summary, this work provides a formalization of automatic modular reasoning patterns and lays a theoretical foundation for the compositional program verification of CSSs.

AAAI Conference 2022 Conference Paper

Co-promotion Predictions of Financing Market and Sales Market: A Cooperative-Competitive Attention Approach

  • Lei Zhang
  • Wang Xiang
  • Chuang Zhao
  • Hongke Zhao
  • Rui Li
  • Runze Wu

Market popularity prediction has always been a hot research topic, such as sales prediction and crowdfunding prediction. Most of these studies put the perspective on isolated markets, relying on the knowledge of certain market to maximize the prediction performance. However, these market-specific approaches are restricted by the knowledge limitation of isolated markets and incapable of the complicated and potential relations among different markets, especially some with strong dependence such as the financing market and sales market. Fortunately, we discover potentially symbiotic relations between the financing market and the sales market, which provides us with an opportunity to co-promote the popularity predictions of both markets. Thus, for bridgly learning the knowledge interactions between financing market and sales market, we propose a cross-market approach, namely CATN: Cooperative-competitive Attention Transfer Network, which could effectively transfer knowledge of financing capability from the crowdfunding market and sales prospect from the E-commerce market. Specifically, for capturing the complicated relations especially the cooperation or complement of items and enhancing the knowledge transfer between the two heterogeneous markets, we design a novel Cooperative Attention; meanwhile, for finely computing the relations of items especially the competition in specific same market, we further design Competitive Attentions for the two markets respectively. Besides, we also distinguish aligned features and unique features to adapt the cross-market predictions. With the real-world datasets collected from Indiegogo and Amazon, we construct extensive experiments on three types of datasets from the two markets and the results demonstrate the effectiveness and generalization of our CATN model.

JBHI Journal 2022 Journal Article

Deep Neural Network With Structural Similarity Difference and Orientation-Based Loss for Position Error Classification in the Radiotherapy of Graves’ Ophthalmopathy Patients

  • Wenjie Liu
  • Lei Zhang
  • Guyu Dai
  • Xiangbin Zhang
  • Guangjun Li
  • Zhang Yi

Identifying position errors for Graves’ ophthalmopathy (GO) patients using electronic portal imaging device (EPID) transmission fluence maps is helpful in monitoring treatment. However, most of the existing models only extract features from dose difference maps computed from EPID images, which do not fully characterize all information of the positional errors. In addition, the position error has a three-dimensional spatial nature, which has never been explored in previous work. To address the above problems, a deep neural network (DNN) model with structural similarity difference and orientation-based loss is proposed in this paper, which consists of a feature extraction network and a feature enhancement network. To capture more information, three types of Structural SIMilarity (SSIM) sub-index maps are computed to enhance the luminance, contrast, and structural features of EPID images, respectively. These maps and the dose difference maps are fed into different networks to extract radiomic features. To acquire spatial features of the position errors, an orientation-based loss function is proposed for optimal training. It makes the data distribution more consistent with the realistic 3D space by integrating the error deviations of the predicted values in the left-right, superior-inferior, anterior-posterior directions. Experimental results on a constructed dataset demonstrate the effectiveness of the proposed model, compared with other related models and existing state-of-the-art methods.

JBHI Journal 2022 Journal Article

Dual-Branch Interactive Networks on Multichannel Time Series for Human Activity Recognition

  • Yin Tang
  • Lei Zhang
  • Hao Wu
  • Jun He
  • Aiguo Song

The popularity of convolutional architecture has made sensor-based human activity recognition (HAR) become one primary beneficiary. By simply superimposing multiple convolution layers, the local features can be effectively captured from multi-channel time series sensor data, which could output high-performance activity prediction results. On the other hand, recent years have witnessed great success of Transformer model, which uses powerful self-attention mechanism to handle long-range sequence modeling tasks, hence avoiding the shortcoming of local feature representations caused by convolutional neural networks (CNNs). In this paper, we seek to combine the merits of CNN and Transformer to model multi-channel time series sensor data, which might provide compelling recognition performance with fewer parameters and FLOPs based on lightweight wearable devices. To this end, we propose a new Dual-branch Interactive Network (DIN) that inherits the advantages from both CNN and Transformer to handle multi-channel time series for HAR. Specifically, the proposed framework utilizes two-stream architecture to disentangle local and global features by performing conv-embedding and patch-embedding, where a co-attention mechanism is used to adaptively fuse global-to-local and local-to-global feature representations. We perform extensive experiments on three mainstream HAR benchmark datasets including PAMAP2, WISDM, and OPPORTUNITY, which verify that our method consistently outperforms several state-of-the-art baselines, reaching an F1-score of 92. 05%, 98. 17%, and 91. 55% respectively with fewer parameters and FLOPs. In addition, the practical execution time is validated on an embedded Raspberry Pi P3 system, which demonstrates that our approach is adequately efficient for real-time HAR implementations and deserves as a better alternative in ubiquitous HAR computing scenario. Our model code will be released soon.

YNIMG Journal 2022 Journal Article

Effective connectivity reveals distinctive patterns in response to others’ genuine affective experience of disgust

  • Yili Zhao
  • Lei Zhang
  • Markus Rütgen
  • Ronald Sladky
  • Claus Lamm

Empathy is significantly influenced by the identification of others' emotions. In a recent study, we have found increased activation in the anterior insular cortex (aIns) that could be attributed to affect sharing rather than perceptual saliency, when seeing another person genuinely experiencing pain as opposed to merely acting to be in pain. In that prior study, effective connectivity between aIns and the right supramarginal gyrus (rSMG) was revealed to represent what another person really feels. In the present study, we used a similar paradigm to investigate the corresponding neural signatures in the domain of empathy for disgust - with participants seeing others genuinely sniffing unpleasant odors as compared to pretending to smell something disgusting (in fact the disgust expressions in both conditions were acted for reasons of experimental control). Consistent with the previous findings on pain, we found stronger activations in aIns associated with affect sharing for genuine disgust (inferred) compared with pretended disgust. However, instead of rSMG we found engagement of the olfactory cortex. Using dynamic causal modeling (DCM), we estimated the neural dynamics of aIns and the olfactory cortex between the genuine and pretended conditions. This revealed an increased excitatory modulatory effect for genuine disgust compared to pretended disgust. For genuine disgust only, brain-to-behavior regression analyses highlighted a link between the observed modulatory effect and a few empathic traits. Altogether, the current findings complement and expand our previous work, by showing that perceptual saliency alone does not explain responses in the insular cortex. Moreover, it reveals that different brain networks are implicated in a modality-specific way when sharing the affective experiences associated with pain vs. disgust.

AAAI Conference 2022 Short Paper

From “Dynamics on Graphs” to “Dynamics of Graphs”: An Adaptive Echo-State Network Solution (Student Abstract)

  • Lei Zhang
  • Zhiqian Chen
  • Chang-Tien Lu
  • Liang Zhao

Many real-world networks evolve over time, which results in dynamic graphs such as human mobility networks and brain networks. Usually, the “dynamics on graphs” (e. g. , node attribute values evolving) are observable, and may be related to and indicative of the underlying “dynamics of graphs” (e. g. , evolving of the graph topology). Traditional RNN-based methods are not adaptive or scalable for learning the unknown mappings between two types of dynamic graph data. This study presents a AD-ESN, and adaptive echo state network that can automatically learn the best neural network architecture for certain data while keeping the efficiency advantage of echo state networks. We show that AD-ESN can successfully discover the underlying pre-defined mapping function and unknown nonlinear map-ping between time series and graphs.

AAAI Conference 2022 Conference Paper

Image-Adaptive YOLO for Object Detection in Adverse Weather Conditions

  • Wenyu Liu
  • Gaofeng Ren
  • Runsheng Yu
  • Shi Guo
  • Jianke Zhu
  • Lei Zhang

Though deep learning-based object detection methods have achieved promising results on the conventional datasets, it is still challenging to locate objects from the low-quality images captured in adverse weather conditions. The existing methods either have difficulties in balancing the tasks of image enhancement and object detection, or often ignore the latent information beneficial for detection. To alleviate this problem, we propose a novel Image-Adaptive YOLO (IA-YOLO) framework, where each image can be adaptively enhanced for better detection performance. Specifically, a differentiable image processing (DIP) module is presented to take into account the adverse weather conditions for YOLO detector, whose parameters are predicted by a small convolutional neural network (CNN-PP). We learn CNN-PP and YOLOv3 jointly in an end-to-end fashion, which ensures that CNN-PP can learn an appropriate DIP to enhance the image for detection in a weakly supervised manner. Our proposed IA-YOLO approach can adaptively process images in both normal and adverse weather conditions. The experimental results are very encouraging, demonstrating the effectiveness of our proposed IA- YOLO method in both foggy and low-light scenarios. The source code can be found at https: //github. com/wenyyu/Image- Adaptive-YOLO.

YNIMG Journal 2022 Journal Article

Lip movements enhance speech representations and effective connectivity in auditory dorsal stream

  • Lei Zhang
  • Yi Du

Viewing speaker's lip movements facilitates speech perception, especially under adverse listening conditions, but the neural mechanisms of this perceptual benefit at the phonemic and feature levels remain unclear. This fMRI study addressed this question by quantifying regional multivariate representation and network organization underlying audiovisual speech-in-noise perception. Behaviorally, valid lip movements improved recognition of place of articulation to aid phoneme identification. Meanwhile, lip movements enhanced neural representations of phonemes in left auditory dorsal stream regions, including frontal speech motor areas and supramarginal gyrus (SMG). Moreover, neural representations of place of articulation and voicing features were promoted differentially by lip movements in these regions, with voicing enhanced in Broca's area while place of articulation better encoded in left ventral premotor cortex and SMG. Next, dynamic causal modeling (DCM) analysis showed that such local changes were accompanied by strengthened effective connectivity along the dorsal stream. Moreover, the neurite orientation dispersion of the left arcuate fasciculus, the bearing skeleton of auditory dorsal stream, predicted the visual enhancements of neural representations and effective connectivity. Our findings provide novel insight to speech science that lip movements promote both local phonemic and feature encoding and network connectivity in the dorsal pathway and the functional enhancement is mediated by the microstructural architecture of the circuit.

JBHI Journal 2022 Journal Article

Multi-Level Attention Network for Retinal Vessel Segmentation

  • Yuchen Yuan
  • Lei Zhang
  • Lituan Wang
  • Haiying Huang

Automatic vessel segmentation in the fundus images plays an important role in the screening, diagnosis, treatment, and evaluation of various cardiovascular and ophthalmologic diseases. However, due to the limited well-annotated data, varying size of vessels, and intricate vessel structures, retinal vessel segmentation has become a long-standing challenge. In this paper, a novel deep learning model called AACA-MLA-D-UNet is proposed to fully utilize the low-level detailed information and the complementary information encoded in different layers to accurately distinguish the vessels from the background with low model complexity. The architecture of the proposed model is based on U-Net, and the dropout dense block is proposed to preserve maximum vessel information between convolution layers and mitigate the over-fitting problem. The adaptive atrous channel attention module is embedded in the contracting path to sort the importance of each feature channel automatically. After that, the multi-level attention module is proposed to integrate the multi-level features extracted from the expanding path, and use them to refine the features at each individual layer via attention mechanism. The proposed method has been validated on the three publicly available databases, i. e. the DRIVE, STARE, and CHASE $\_$ DB1. The experimental results demonstrate that the proposed method can achieve better or comparable performance on retinal vessel segmentation with lower model complexity. Furthermore, the proposed method can also deal with some challenging cases and has strong generalization ability.

AAAI Conference 2022 Conference Paper

Neighborhood-Adaptive Structure Augmented Metric Learning

  • Pandeng Li
  • Yan Li
  • Hongtao Xie
  • Lei Zhang

Most metric learning techniques typically focus on sample embedding learning, while implicitly assume a homogeneous local neighborhood around each sample, based on the metrics used in training (e. g. , hypersphere for Euclidean distance or unit hyperspherical crown for cosine distance). As realworld data often lies on a low-dimensional manifold curved in a high-dimensional space, it is unlikely that everywhere of the manifold shares the same local structures in the input space. Besides, considering the non-linearity of neural networks, the local structure in the output embedding space may not be as homogeneous as assumed. Therefore, representing each sample simply with its embedding while ignoring its individual neighborhood structure would have limitations in Embedding-Based Retrieval (EBR). By exploiting the heterogeneity of local structures in the embedding space, we propose a Neighborhood-Adaptive Structure Augmented metric learning framework (NASA), where the neighborhood structure is realized as a structure embedding, and learned along with the sample embedding in a self-supervised manner. In this way, without any modifications, most indexing techniques can be used to support large-scale EBR with NASA embeddings. Experiments on six standard benchmarks with two kinds of embeddings, i. e. , binary embeddings and real-valued embeddings, show that our method significantly improves and outperforms the state-of-the-art methods.

IJCAI Conference 2022 Conference Paper

Reconciling Cognitive Modeling with Knowledge Forgetting: A Continuous Time-aware Neural Network Approach

  • Haiping Ma
  • Jingyuan Wang
  • Hengshu Zhu
  • Xin Xia
  • Haifeng Zhang
  • Xingyi Zhang
  • Lei Zhang

As an emerging technology of computer-aided education, cognitive modeling aims at discovering the knowledge proficiency or learning ability of students, which can enable a wide range of intelligent educational applications. While considerable efforts have been made in this direction, a long-standing research challenge is how to naturally integrate the forgetting mechanism into the learning process of knowledge concepts. To this end, in this paper, we propose a novel Continuous Time based Neural Cognitive Modeling(CT-NCM) approach to integrate the dynamism and continuity of knowledge forgetting into students' learning process modeling in a realistic manner. To be specific, we first adapt the neural Hawkes process with a specially-designed learning event encoding method to model the relationship between knowledge learning and forgetting with continuous time. Then, we propose a learning function with extendable settings to jointly model the change of different knowledge states and their interactions with the exercises at each moment. In this way, CT-NCM can simultaneously predict the future knowledge state and exercise performance of students. Finally, we conduct extensive experiments on five real-world datasets with various benchmark methods. The experimental results clearly validate the effectiveness of CT-NCM and show its interpretability in terms of knowledge learning visualization.

AAAI Conference 2021 Conference Paper

Adversarial Pose Regression Network for Pose-Invariant Face Recognitions

  • Pengyu Li
  • Biao Wang
  • Lei Zhang

Face recognition has achieved significant progress in recent years. However, the large pose variation between face images remains a challenge in face recognition. We observe that the pose variation in the hidden feature maps is one of the most critical factors to hinder the representations from being pose-invariant. Based on the observation, we propose an Adversarial Pose Regression Network (APRN) to extract poseinvariant identity representations by disentangling their pose variation in hidden feature maps. To model the pose discriminator in APRN as a regression task in its 3D space, we also propose an Adversarial Regression Loss Function and extend the adversarial learning from classification problems to regression problems in this paper. Our APRN is a plug-andplay structure that can be embedded in other state-of-the-art face recognition algorithms to improve their performance additionally. The experiments show that the proposed APRN consistently and significantly boosts the performance of baseline networks without extra computational costs in the inference phase. APRN achieves comparable or even superior to the state-of-the-art on CFP, Multi-PIE, IJB-A and MegaFace datasets. The code will be released1, hoping to nourish our proposals to other computer vision fields.

AAAI Conference 2021 Conference Paper

Category Dictionary Guided Unsupervised Domain Adaptation for Object Detection

  • Shuai Li
  • Jianqiang Huang
  • Xian-Sheng Hua
  • Lei Zhang

Unsupervised domain adaption (UDA) is a promising solution to enhance the generalization ability of a model from a source domain to a target domain without manually annotating labels for the target data. Recent works in cross-domain object detection mostly resort to adversarial feature adaptation to match the marginal distributions of two domains. However, perfect feature alignment is hard to achieve and what’s more is likely to cause negative transfer due to the high complexity of object detection. In this paper, we take a different approach to reduce the domain gap by a selftraining paradigm, which regards the pseudo-labels as ground truth to fully exploit the unlabeled target data. In order to generate more informative pseudo labels, we further propose a category dictionary guided (CDG) UDA model for crossdomain object detection, which learns category-specific dictionaries from the source domain to represent the candidate boxes in target domain. The representation residual can be used for not only pseudo label assignment but also quality (e. g. , IoU) estimation of the candidate box. Compared with decision boundary based classifiers such as softmax, the proposed CDG scheme can select more informative and reliable pseudo-boxes. Experimental results on benchmark datasets show that the proposed CDG significantly exceeds the stateof-the-arts in cross-domain object detection.

NeurIPS Conference 2021 Conference Paper

Chasing Sparsity in Vision Transformers: An End-to-End Exploration

  • Tianlong Chen
  • Yu Cheng
  • Zhe Gan
  • Lu Yuan
  • Lei Zhang
  • Zhangyang Wang

Vision transformers (ViTs) have recently received explosive popularity, but their enormous model sizes and training costs remain daunting. Conventional post-training pruning often incurs higher training budgets. In contrast, this paper aims to trim down both the training memory overhead and the inference complexity, without sacrificing the achievable accuracy. We carry out the first-of-its-kind comprehensive exploration, on taking a unified approach of integrating sparsity in ViTs "from end to end''. Specifically, instead of training full ViTs, we dynamically extract and train sparse subnetworks, while sticking to a fixed small parameter budget. Our approach jointly optimizes model parameters and explores connectivity throughout training, ending up with one sparse network as the final output. The approach is seamlessly extended from unstructured to structured sparsity, the latter by considering to guide the prune-and-grow of self-attention heads inside ViTs. We further co-explore data and architecture sparsity for additional efficiency gains by plugging in a novel learnable token selector to adaptively determine the currently most vital patches. Extensive results on ImageNet with diverse ViT backbones validate the effectiveness of our proposals which obtain significantly reduced computational cost and almost unimpaired generalization. Perhaps most surprisingly, we find that the proposed sparse (co-)training can sometimes \textit{improve the ViT accuracy} rather than compromising it, making sparsity a tantalizing "free lunch''. For example, our sparsified DeiT-Small at ($5\%$, $50\%$) sparsity for (data, architecture), improves $\mathbf{0. 28\%}$ top-1 accuracy, and meanwhile enjoys $\mathbf{49. 32\%}$ FLOPs and $\mathbf{4. 40\%}$ running time savings. Our codes are available at https: //github. com/VITA-Group/SViTE.

AAAI Conference 2021 Conference Paper

Deep Metric Learning with Graph Consistency

  • Binghui Chen
  • Pengyu Li
  • Zhaoyi Yan
  • Biao Wang
  • Lei Zhang

Deep Metric Learning (DML) has been more attractive and widely applied in many computer vision tasks, in which a discriminative embedding is requested such that the image features belonging to the same class are gathered together and the ones belonging to different classes are pushed apart. Most existing works insist to learn this discriminative embedding by either devising powerful pair-based loss functions or hardsample mining strategies. However, in this paper, we start from another perspective and propose Deep Consistent Graph Metric Learning (CGML) framework to enhance the discrimination of the learned embedding. It is mainly achieved by rethinking the conventional distance constraints as a graph regularization and then introducing a Graph Consistency regularization term, which intends to optimize the feature distribution from a global graph perspective. Inspired by the characteristic of our defined ’Discriminative Graph’, which regards DML from another novel perspective, the Graph Consistency regularization term encourages the sub-graphs randomly sampled from the training set to be consistent. We show that our CGML indeed serves as an efficient technique for learning towards discriminative embedding and is applicable to various popular metric objectives, e. g. Triplet, N-Pair and Binomial losses. This paper empirically and experimentally demonstrates the effectiveness of our graph regularization idea, achieving competitive results on the popular CUB, CARS, Stanford Online Products and In-Shop datasets.

YNIMG Journal 2021 Journal Article

Effects of non-invasive brain stimulation on visual perspective taking: A meta-analytic study

  • Yuan-Wei Yao
  • Vivien Chopurian
  • Lei Zhang
  • Claus Lamm
  • Hauke R. Heekeren

Visual perspective taking (VPT) is a critical ability required by complex social interaction. Non-invasive brain stimulation (NIBS) has been increasingly used to examine the causal relationship between brain activity and VPT, yet with heterogeneous results. In the current study, we conducted two meta-analyses to examine the effects of NIBS of the right temporoparietal junction (rTPJ) or dorsomedial prefrontal cortex (dmPFC) on VPT, respectively. We performed a comprehensive literature search to identify qualified studies and computed the standardized effect size (ES) for each combination of VPT level (Level-1: visibility judgment; Level-2: mental rotation) and perspective (self and other). Thirteen studies (rTPJ: 12 studies, 23 ESs; dmPFC: 4 studies, 14 ESs) were included in the meta-analyses. Random-effects models were used to generate the overall effects. Subgroup analyses for distinct VPT conditions were also performed. We found that rTPJ stimulation significantly improved participants' visibility judgment from the allocentric perspective, whereas its effects on other VPT conditions are negligible. Stimulation of dmPFC appeared to influence Level-1 performance from the egocentric perspective, although this finding was only based on a small number of studies. Notably, contrary to some theoretical models, we did not find strong evidence that these regions are involved in Level-2 VPT with a higher requirement of mental rotation. These findings not only advance our understanding of the causal roles of the rTPJ and dmPFC in VPT, but also reveal that the efficacy of NIBS on VPT is relatively small. Additionally, researchers should also be cautious about the potential publication bias and selective reporting.

AAAI Conference 2021 Conference Paper

Question-Driven Span Labeling Model for Aspect–Opinion Pair Extraction

  • Lei Gao
  • Yulong Wang
  • Tongcun Liu
  • Jingyu Wang
  • Lei Zhang
  • Jianxin Liao

Aspect term extraction and opinion word extraction are two fundamental subtasks of aspect-based sentiment analysis. The internal relationship between aspect terms and opinion words is typically ignored, and information for the decisionmaking of buyers and sellers is insufficient. In this paper, we explore an aspect–opinion pair extraction (AOPE) task and propose a Question-Driven Span Labeling (QDSL) model to extract all the aspect–opinion pairs from user-generated reviews. Specifically, we divide the AOPE task into aspect term extraction (ATE) and aspect-specified opinion extraction (ASOE) subtasks; we first extract all the candidate aspect terms and then the corresponding opinion words given the aspect term. Unlike existing approaches that use the BIObased tagging scheme for extraction, the QDSL model adopts a span-based tagging scheme and builds a question–answerbased machine-reading comprehension task for an effective aspect–opinion pair extraction. Extensive experiments conducted on three tasks (ATE, ASOE, and AOPE) on four benchmark datasets demonstrate that the proposed method significantly outperforms state-of-the-art approaches.

JBHI Journal 2021 Journal Article

The Convolutional Neural Networks Training With Channel-Selectivity for Human Activity Recognition Based on Sensors

  • Wenbo Huang
  • Lei Zhang
  • Qi Teng
  • Chaoda Song
  • Jun He

Recently, the state-of-the-art performance in various sensor based human activity recognition (HAR) tasks have been acquired by deep learning, which can extract automatically features from raw data. In order to obtain the best accuracy, many static layers have been always used to train deep neural networks, and their weight connectivity in network remains unchanged. Pursuing the best accuracy in mobile platforms with a very limited computational budget at millions of FLOPs is impractical. In this paper, we make use of shallow convolutional neural networks (CNNs) with channel-selectivity for the use of HAR. As we have known, it is for the first time to adopt channel-selectivity CNN for sensor based HAR tasks. We perform extensive experiments on 5 public benchmark HAR datasets consisting of UCI-HAR dataset, OPPORTUNITY dataset, UniMib-SHAR dataset, WISDM dataset, and PAMAP2 dataset. As a result, the channel-selectivity can achieve lower test errors than static layers. The existing performance of deep HAR can be further improved by the CNN with channel-selectivity without any extra cost.

YNIMG Journal 2021 Journal Article

The effect of beta-amyloid and tau protein aggregations on magnetic susceptibility of anterior hippocampal laminae in Alzheimer's diseases

  • Zhiyong Zhao
  • Lei Zhang
  • Qingqing Wen
  • Wanrong Luo
  • Weihao Zheng
  • Tingting Liu
  • Yi Zhang
  • Keqing Zhu

Previous studies have reported the changes of magnetic susceptibility induced by iron deposition in hippocampus of Alzheimer's disease (AD) brains. It is well-known that hippocampus is divided into well-defined laminar architecture, which, however, is difficult to be resolved with in-vivo MRI due to the limited imaging resolution. The present study aims to investigate layer-specific magnetic susceptibility in the hippocampus of AD patients using high-resolution ex-vivo MRI, and elucidate its relationship with beta amyloid (Aβ) and tau protein histology. We performed quantitative susceptibility mapping (QSM) and T2* mapping on postmortem anterior hippocampus samples from four AD, four Primary Age-Related Tauopathy (PART), and three control brains. We manually segmented each sample into seven layers, including four layers in the cornu ammonis1 (CA1) and three layers in the dentate gyrus (DG), and then evaluated AD-related alterations of susceptibility and T2* values and their correlations with Aβ and tau in each hippocampal layer. Specifically, we found (1) layer-specific variations of susceptibility and T2* measurements in all samples; (2) the heterogeneity of susceptibility were higher in all layers of AD patients compared with the age- and gender-matched PART cases while the heterogeneity of T2* values were lower in four layers of CA1; and (3) voxel-wise MRI-histological correlation revealed both susceptibility and T2* values in the stratum molecular (SM) and stratum lacunosum (SL) layers were correlated with the Aβ content in AD, while the T2* values in the stratum radiatum (SR) layer were correlated with the tau content in the PART but not AD. These findings suggest a selective effect of the Aβ- and tau-pathology on the susceptibility and T2* values in the different layers of anterior hippocampus. Particularly, the alterations of magnetic susceptibility in the SM and SL layers may be associated with Aβ aggregation, while those in the SR layermay reflect the age-related tau protein aggregation.

AAAI Conference 2021 Conference Paper

VIVO: Visual Vocabulary Pre-Training for Novel Object Captioning

  • Xiaowei Hu
  • Xi Yin
  • Kevin Lin
  • Lei Zhang
  • Jianfeng Gao
  • Lijuan Wang
  • Zicheng Liu

It is highly desirable yet challenging to generate image captions that can describe novel objects which are unseen in caption-labeled training data, a capability that is evaluated in the novel object captioning challenge (nocaps). In this challenge, no additional image-caption training data, other than COCO Captions, is allowed for model training. Thus, conventional Vision-Language Pre-training (VLP) methods cannot be applied. This paper presents VIsual VOcabulary pretraining (VIVO) that performs pre-training in the absence of caption annotations. By breaking the dependency of paired image-caption training data in VLP, VIVO can leverage large amounts of paired image-tag data to learn a visual vocabulary. This is done by pre-training a multi-layer Transformer model that learns to align image-level tags with their corresponding image region features. To address the unordered nature of image tags, VIVO uses a Hungarian matching loss with masked tag prediction to conduct pre-training. We validate the effectiveness of VIVO by fine-tuning the pre-trained model for image captioning. In addition, we perform an analysis of the visual-text alignment inferred by our model. The results show that our model can not only generate fluent image captions that describe novel objects, but also identify the locations of these objects. Our single model has achieved new state-of-the-art results on nocaps and surpassed the human CIDEr score.

ICRA Conference 2021 Conference Paper

When Shall I Be Empathetic? The Utility of Empathetic Parameter Estimation in Multi-Agent Interactions

  • Yi Chen
  • Lei Zhang
  • Tanner Merry
  • Sunny Amatya
  • Wenlong Zhang
  • Yi Ren

Human-robot interactions (HRI) can be modeled as differential games with incomplete information, where each agent holds private reward parameters. Due to the open challenge in finding perfect Bayesian equilibria of such games, existing studies often decouple the belief and physical dynamics by iterating between belief update and motion planning. Importantly, the robot’s reward parameters are often assumed to be known to the humans, in order to simplify the computation. We show in this paper that under this simplification, the robot performs non-empathetic belief update about the humans’ parameters, which causes high safety risks in uncontrolled intersection scenarios. In contrast, we propose a model for empathetic belief update, where the agent updates the joint probabilities of all agents’ parameter combinations. The update uses a neural network that approximates the Nash equilibrial action-values of agents. We compare empathetic and non-empathetic belief update methods on a two-vehicle uncontrolled intersection case with short reaction time. Results show that when both agents are unknowingly aggressive (or non-aggressive), empathy is necessary for avoiding collisions when agents have false believes about each others’ parameters. This paper demonstrates the importance of acknowledging the incomplete-information nature of HRI.

AAAI Conference 2020 Conference Paper

A Multi-Unit Profit Competitive Mechanism for Cellular Traffic Offloading

  • Jun Wu
  • Yu Qiao
  • Lei Zhang
  • Chongjun Wang
  • Meilin Liu

Cellular traffic offloading is nowadays an important problem in mobile networking. We model it as a procurement problem where each agent sells multi-units of a homogeneous item with privately known capacity and unit cost, and the auctioneer’s demand valuation function is symmetric submodular. Based on the framework of random sampling and profit extraction, we aim to design a prior-free mechanism which guarantees a profit competitive to the omniscient single-price auction. However, the symmetric submodular demand valuation function and 2-parameter setting present new challenges. By adopting the highest feasible clear price, we successfully design a truthful profit extractor, and then we propose a mechanism which is proved to be truthful, individually rational and constant-factor competitive in a fixed market.

YNICL Journal 2020 Journal Article

Altered resting-state dynamic functional brain networks in major depressive disorder: Findings from the REST-meta-MDD consortium

  • Yicheng Long
  • Hengyi Cao
  • Chaogan Yan
  • Xiao Chen
  • Le Li
  • Francisco Xavier Castellanos
  • Tongjian Bai
  • Qijing Bo

BACKGROUND: Major depressive disorder (MDD) is known to be characterized by altered brain functional connectivity (FC) patterns. However, whether and how the features of dynamic FC would change in patients with MDD are unclear. In this study, we aimed to characterize dynamic FC in MDD using a large multi-site sample and a novel dynamic network-based approach. METHODS: Resting-state functional magnetic resonance imaging (fMRI) data were acquired from a total of 460 MDD patients and 473 healthy controls, as a part of the REST-meta-MDD consortium. Resting-state dynamic functional brain networks were constructed for each subject by a sliding-window approach. Multiple spatio-temporal features of dynamic brain networks, including temporal variability, temporal clustering and temporal efficiency, were then compared between patients and healthy subjects at both global and local levels. RESULTS: ). Corresponding local changes in MDD were mainly found in the default-mode, sensorimotor and subcortical areas. Measures of temporal variability and characteristic temporal path length were significantly correlated with depression severity in patients (corrected p < 0.05). Moreover, the observed between-group differences were robustly present in both first-episode, drug-naïve (FEDN) and non-FEDN patients. CONCLUSIONS: Our findings suggest that excessive temporal variations of brain FC, reflecting abnormal communications between large-scale bran networks over time, may underlie the neuropathology of MDD.

YNIMG Journal 2020 Journal Article

Evaluation of the diffusion MRI white matter tract integrity model using myelin histology and Monte-Carlo simulations

  • Zihan Zhou
  • Qiqi Tong
  • Lei Zhang
  • Qiuping Ding
  • Hui Lu
  • Laura E. Jonkman
  • Junye Yao
  • Hongjian He

Quantitative evaluation of brain myelination has drawn considerable attention. Conventional diffusion-based magnetic resonance imaging models, including diffusion tensor imaging and diffusion kurtosis imaging (DKI), 1 1 AD: Axial diffusivity; AK: Axial kurtosis; AVF: Axonal volume fraction; AWF: Axonal Water Fraction; DTI: Diffusion Tensor Imaging; DKI: Diffusion Kurtosis Imaging; dMRI: Diffusion magnetic resonance imaging; TE: Echo time; FOV: Field-of-view; FA: Fractional anisotropy; LFB: Luxor fast blue; MD: Mean diffusivity; MK: Mean kurtosis; MRI: Magnetic resonance imaging; MVF: Myelin volume fraction; PLP: Proteolipid protein; RESOLVE: Readout segmentation of long variable echo train; RD: Radial diffusivity; RK: Radial kurtosis; TR: Repetition time; ROI: Region-of-interest; WM: White matter; WMTI: White Matter Tract Integrity. have been used to infer the microstructure and its changes in neurological diseases. White matter tract integrity (WMTI) was proposed as a biophysical model to relate the DKI-derived metrics to the underlying microstructure. Although the model has been validated on ex vivo animal brains, it was not well evaluated with ex vivo human brains. In this study, histological samples (namely corpus callosum) from postmortem human brains have been investigated based on WMTI analyses on a clinical 3T scanner and comparisons with gold standard myelin staining in proteolipid protein and Luxol fast blue. In addition, Monte Carlo simulations were conducted to link changes from ex vivo to in vivo conditions based on the microscale parameters of water diffusivity and permeability. The results show that WMTI metrics, including axonal water fraction AWF, radial extra-axonal diffusivity D e ⊥, and intra-axonal diffusivity Da were needed to characterize myelin content alterations. Thus, WMTI model metrics are shown to be promising candidates as sensitive biomarkers of demyelination.

AIIM Journal 2020 Journal Article

Handling imbalanced medical image data: A deep-learning-based one-class classification approach

  • Long Gao
  • Lei Zhang
  • Chang Liu
  • Shandong Wu

In clinical settings, a lot of medical image datasets suffer from the imbalance problem which hampers the detection of outliers (rare health care events), as most classification methods assume an equal occurrence of classes. In this way, identifying outliers in imbalanced datasets has become a crucial issue. To help address this challenge, one-class classification, which focuses on learning a model using samples from only a single given class, has attracted increasing attention. Previous one-class modeling usually uses feature mapping or feature fitting to enforce the feature learning process. However, these methods are limited for medical images which usually have complex features. In this paper, a novel method is proposed to enable deep learning models to optimally learn single-class-relevant inherent imaging features by leveraging the concept of imaging complexity. We investigate and compare the effects of simple but effective perturbing operations applied to images to capture imaging complexity and to enhance feature learning. Extensive experiments are performed on four clinical datasets to show that the proposed method outperforms four state-of-the-art methods.

JBHI Journal 2020 Journal Article

Inaccurate Labels in Weakly-Supervised Deep Learning: Automatic Identification and Correction and Their Impact on Classification Performance

  • Degan Hao
  • Lei Zhang
  • Jules Sumkin
  • Aly Mohamed
  • Shandong Wu

In data-driven deep learning-based modeling, data quality may substantially influence classification performance. Correct data labeling for deep learning modeling is critical. In weakly-supervised learning, a challenge lies in dealing with potentially inaccurate or mislabeled training data. In this paper, we proposed an automated methodological framework to identify mislabeled data using two metric functions, namely, Cross-entropy Loss that indicates divergence between a prediction and ground truth, and Influence function that reflects the dependence of a model on data. After correcting the identified mislabels, we measured their impact on the classification performance. We also compared the mislabeling effects in three experiments on two different real-world clinical questions. A total of 10, 500 images were studied in the contexts of clinical breast density category classification and breast cancer malignancy diagnosis. We used intentionally flipped labels as mislabels to evaluate the proposed method at a varying proportion of mislabeled data included in model training. We also compared the effects of our method to two published schemes for breast density category classification. Experiment results show that when the dataset contains 10% of mislabeled data, our method can automatically identify up to 98% of these mislabeled data by examining/checking the top 30% of the full dataset. Furthermore, we show that correcting the identified mislabels leads to an improvement in the classification performance. Our method provides a feasible solution for weakly-supervised deep learning modeling in dealing with inaccurate labels.

AAAI Conference 2020 Conference Paper

Multi-Channel Reverse Dictionary Model

  • Lei Zhang
  • Fanchao Qi
  • Zhiyuan Liu
  • Yasheng Wang
  • Qun Liu
  • Maosong Sun

A reverse dictionary takes the description of a target word as input and outputs the target word together with other words that match the description. Existing reverse dictionary methods cannot deal with highly variable input queries and low-frequency target words successfully. Inspired by the description-to-word inference process of humans, we propose the multi-channel reverse dictionary model, which can mitigate the two problems simultaneously. Our model comprises a sentence encoder and multiple predictors. The predictors are expected to identify different characteristics of the target word from the input query. We evaluate our model on English and Chinese datasets including both dictionary definitions and human-written descriptions. Experimental results show that our model achieves the state-of-the-art performance, and even outperforms the most popular commercial reverse dictionary system on the human-written description dataset. We also conduct quantitative analyses and a case study to demonstrate the effectiveness and robustness of our model. All the code and data of this work can be obtained on https: //github. com/thunlp/MultiRD.

AAAI Conference 2020 Conference Paper

Pixel-Aware Deep Function-Mixture Network for Spectral Super-Resolution

  • Lei Zhang
  • Zhiqiang Lang
  • Peng Wang
  • Wei Wei
  • Shengcai Liao
  • Ling Shao
  • Yanning Zhang

Spectral super-resolution (SSR) aims at generating a hyperspectral image (HSI) from a given RGB image. Recently, a promising direction is to learn a complicated mapping function from the RGB image to the HSI counterpart using a deep convolutional neural network. This essentially involves mapping the RGB context within a size-specific receptive field centered at each pixel to its spectrum in the HSI. The focus thereon is to appropriately determine the receptive field size and establish the mapping function from RGB context to the corresponding spectrum. Due to their differences in category or spatial position, pixels in HSIs often require different-sized receptive fields and distinct mapping functions. However, few efforts have been invested to explicitly exploit this prior. To address this problem, we propose a pixel-aware deep function-mixture network for SSR, which is composed of a new class of modules, termed function-mixture (FM) blocks. Each FM block is equipped with some basis functions, i. e. , parallel subnets of different-sized receptive fields. Besides, it incorporates an extra subnet as a mixing function to generate pixel-wise weights, and then linearly mixes the outputs of all basis functions with those generated weights. This enables us to pixel-wisely determine the receptive field size and the mapping function. Moreover, we stack several such FM blocks to further increase the flexibility of the network in learning the pixel-wise mapping. To encourage feature reuse, intermediate features generated by the FM blocks are fused in late stage, which proves to be effective for boosting the SSR performance. Experimental results on three benchmark HSI datasets demonstrate the superiority of the proposed method.

IJCAI Conference 2020 Conference Paper

Self-adaptive Re-weighted Adversarial Domain Adaptation

  • Shanshan Wang
  • Lei Zhang

Existing adversarial domain adaptation methods mainly consider the marginal distribution and these methods may lead to either under transfer or negative transfer. To address this problem, we present a self-adaptive re-weighted adversarial domain adaptation approach, which tries to enhance domain alignment from the perspective of conditional distribution. In order to promote positive transfer and combat negative transfer, we reduce the weight of the adversarial loss for aligned features while increasing the adversarial force for those poorly aligned measured by the conditional entropy. Additionally, triplet loss leveraging source samples and pseudo-labeled target samples is employed on the confusing domain. Such metric loss ensures the distance of the intra-class sample pairs closer than the inter-class pairs to achieve the class-level alignment. In this way, the high accurate pseudolabeled target samples and semantic alignment can be captured simultaneously in the co-training process. Our method achieved low joint error of the ideal source and target hypothesis. The expected target error can then be upper bounded following Ben-David’s theorem. Empirical evidence demonstrates that the proposed model outperforms state of the arts on standard domain adaptation datasets.

YNICL Journal 2020 Journal Article

Spatiotemporal EEG microstate analysis in drug-free patients with Parkinson's disease

  • Chunguang Chu
  • Xing Wang
  • Lihui Cai
  • Lei Zhang
  • Jiang Wang
  • Chen Liu
  • Xiaodong Zhu

The clinical diagnosis of Parkinson's disease (PD) is very difficult, especially in the early stage of the disease, because there is no physiological indicator that can be referenced. Drug-free patients with early PD are characterized by clinical symptoms such as impaired motor function and cognitive decline, which was caused by the dysfunction of brain's dynamic activities. The indicators of brain dysfunction in patients with PD at an early unmedicated condition may provide a valuable basis for the diagnosis of early PD and later treatment. In order to find the spatiotemporal characteristic markers of brain dysfunction in PD, the resting-state EEG microstate analysis is used to explore the transient state of the whole brain of 23 drug-free patients with PD on the sub-second timescale compared to 23 healthy controls. EEG microstates reflect a transiently stable brain topological structure with spatiotemporal characteristics, and the spatial characteristic microstate classes and temporal parameters provide insight into the brain's functional activities in PD patients. The further exploration was to explore the relation between temporal microstate parameters and significant clinical symptoms to determine whether these parameters could be used as a basis for clinically assisted diagnosis. Therefore, we used a general linear model (GLM) to explore the relevance of microstate parameters to clinical scales and multiple patient attributes, and the Wilcoxon rank sum test was used to quantify the linear relation between influencing factors and microstate parameters. Results of microstate analysis revealed that there was an unique spatial microstate different from healthy controls in PD, and several other typical microstates had significant differences compared with the normal control group, and these differences were reflected in the microstate parameters, such as longer durations and more occurrences of one class of microstates in PD compared with healthy controls. Furthermore, correlation analysis showed that there was a significant correlation between multiple microstate classes' parameters and significant clinical symptoms, including impaired motor function and cognitive decline. These results indicate that we have found multiple quantifiable feature tags that reflect brain dysfunction in the early stage of PD. Importantly, such temporal dynamics in microstates are correlated with clinical scales which represent the motor function and recognize level. The obtained results may deepen our understanding of the brain dysfunction caused by PD, and obtain some quantifiable signatures to provide an auxiliary reference for the early diagnosis of PD.

AAAI Conference 2020 Conference Paper

Unified Vision-Language Pre-Training for Image Captioning and VQA

  • Luowei Zhou
  • Hamid Palangi
  • Lei Zhang
  • Houdong Hu
  • Jason Corso
  • Jianfeng Gao

This paper presents a unified Vision-Language Pre-training (VLP) model. The model is unified in that (1) it can be finetuned for either vision-language generation (e. g. , image captioning) or understanding (e. g. , visual question answering) tasks, and (2) it uses a shared multi-layer transformer network for both encoding and decoding, which differs from many existing methods where the encoder and decoder are implemented using separate models. The unified VLP model is pre-trained on a large amount of image-text pairs using the unsupervised learning objectives of two tasks: bidirectional and sequence-to-sequence (seq2seq) masked vision-language prediction. The two tasks differ solely in what context the prediction conditions on. This is controlled by utilizing specific self-attention masks for the shared transformer network. To the best of our knowledge, VLP is the first reported model that achieves state-of-the-art results on both vision-language generation and understanding tasks, as disparate as image captioning and visual question answering, across three challenging benchmark datasets: COCO Captions, Flickr30k Captions, and VQA 2. 0. The code and the pre-trained models are available at https: //github. com/LuoweiZhou/VLP.

TIST Journal 2020 Journal Article

WiSign

  • Lei Zhang
  • Yixiang Zhang
  • Xiaolong Zheng

In this article, we propose WiSign that recognizes the continuous sentences of American Sign Language (ASL) with existing WiFi infrastructure. Instead of identifying the individual ASL words from the manually segmented ASL sentence in existing works, WiSign can automatically segment the original channel state information (CSI) based on the power spectral density (PSD) segmentation method. WiSign constructs a five-layer Deep Belief Network (DBN) to automatically extract the features of isolated fragments, and then uses the Hidden Markov Model (HMM) with Gaussian mixture and Forward-Backward algorithm to recognize sign words. In order to further improve the accuracy, WiSign also integrates the language model N-gram, which uses the grammar rules of ASL to calibrate the recognized results of sign words. We implement a prototype of WiSign with commercial WiFi devices and evaluate its performance in real indoor environments. The results show that WiSign achieves satisfactory accuracy when recognizing ASL sentences that involve the movements of the head, arms, hands, and fingers.

AAAI Conference 2019 Conference Paper

An Efficient Compressive Convolutional Network for Unified Object Detection and Image Compression

  • Xichuan Zhou
  • Lang Xu
  • Shujun Liu
  • Yingcheng Lin
  • Lei Zhang
  • Cheng Zhuo

This paper addresses the challenge of designing efficient framework for real-time object detection and image compression. The proposed Compressive Convolutional Network (CCN) is basically a compressive-sensing-enabled convolutional neural network. Instead of designing different components for compressive sensing and object detection, the CCN optimizes and reuses the convolution operation for recoverable data embedding and image compression. Technically, the incoherence condition, which is the sufficient condition for recoverable data embedding, is incorporated in the first convolutional layer of the CCN model as regularization; Therefore, the CCN convolution kernels learned by training over the VOC and COCO image set can be used for data embedding and image compression. By reusing the convolution operation, no extra computational overhead is required for image compression. As a result, the CCN is 3. 1 to 5. 0 fold more efficient than the conventional approaches. In our experiments, the CCN achieved 78. 1 mAP for object detection and 3. 0 dB to 5. 2 dB higher PSNR for image compression than the examined compressive sensing approaches.

EAAI Journal 2019 Journal Article

An indexed set representation based multi-objective evolutionary approach for mining diversified top-k high utility patterns

  • Lei Zhang
  • Shangshang Yang
  • Xinpeng Wu
  • Fan Cheng
  • Ying Xie
  • Zhiting Lin

How to discover top-k patterns with the largest utility values, namely, mining top-k high utility patterns, is a hot topic in data mining. However, most of the existing works for mining top-k high utility patterns consider each pattern separately during the mining process, thus many mined patterns are highly similar and lack diversity. In this paper, we propose to mine top-k high utility patterns with high diversity for enhancing users’satisfaction in recommendation. Specifically, we first introduce a simple measure of coverage to quantify the diversity of the whole set, that is, the top-k patterns as a complete entity. Then we propose an i ndexed s et r epresentation based m ulti-o bjective e volutionary a pproach named ISR-MOEA to mine diversified top-k high utility patterns, due to the fact that the two measures utility and coverage are conflicting. In ISR-MOEA, an indexed set individual representation scheme is suggested for fast encoding and decoding the top-k pattern set. Experimental results on six real-world and two synthetic datasets demonstrate the effectiveness of the proposed approach. The proposed approach can obtain several groups of top-k pattern set with different trade-offs between utility and diversity in only one run, which would further enhance the satisfaction of users.

YNIMG Journal 2019 Journal Article

Effects of 3.5–23.0 T static magnetic fields on mice: A safety study

  • Xiaofei Tian
  • Dongmei Wang
  • Shuang Feng
  • Lei Zhang
  • Xinmiao Ji
  • Ze Wang
  • Qingyou Lu
  • Chuanying Xi

People are exposed to various magnetic fields, including the high static/steady magnetic field (SMF) of MRI, which has been increased to 9. 4 T in preclinical investigations. However, relevant safety studies about high SMF are deficient. Here we examined whether 3. 5–23. 0 T SMF exposure for 2 h has severe long-term effects on mice using 112 C57BL/6J mice. The food/water consumption, blood glucose levels, blood routine, blood biochemistry, as well as organ weight and HE stains were all examined. The food consumption and body weight were slightly decreased for 23. 0 T-exposed mice (14. 6%, P < 0. 01, and 1. 75–5. 57%, P < 0. 05, respectively), but not the other groups. While total bilirubin (TBIL), white blood cells, platelet and lymphocyte numbers were affected by some magnetic conditions, most of them were still within normal reference range. Although 13. 5 T magnetic fields with the highest gradient (117. 2 T/m) caused spleen weight increase, the blood count and biochemistry results were still within the control reference range. Moreover, the highest field 23. 0 T with no gradient did not cause organ weight or blood biochemistry abnormality, which indicates that field gradient is a key parameter. Collectively, these data suggest 3. 5–23. 0 T static magnetic field exposure for 2 h do not have severe long-term effects on mice.

AAAI Conference 2019 Conference Paper

Learning a Visual Tracker from a Single Movie without Annotation

  • Lingxiao Yang
  • David Zhang
  • Lei Zhang

The recent success of deep network in visual trackers learning largely relies on human labeled data, which are however expensive to annotate. Recently, some unsupervised methods have been proposed to explore the learning of visual trackers without labeled data, while their performance lags far behind the supervised methods. We identify the main bottleneck of these methods as inconsistent objectives between off-line training and online tracking stages. To address this problem, we propose a novel unsupervised learning pipeline which is based on the discriminative correlation filter network. Our method iteratively updates the tracker by alternating between target localization and network optimization. In particular, we propose to learn the network from a single movie, which could be easily obtained other than collecting thousands of video clips or millions of images. Extensive experiments demonstrate that our approach is insensitive to the employed movies, and the trained visual tracker achieves leading performance among existing unsupervised learning approaches. Even compared with the same network trained with human labeled bounding boxes, our tracker achieves similar results on many tracking benchmarks. Code is available at: https: //github. com/ZjjConan/UL-Tracker-AAAI2019.

AAMAS Conference 2019 Conference Paper

Multi-unit Budget Feasible Mechanisms for Cellular Traffic Offloading

  • Jun Wu
  • Yuan Zhang
  • Yu Qiao
  • Lei Zhang
  • Chongjun Wang
  • Junyuan Xie

Cellular traffic offloading is nowadays an important problem in mobile networking. Since the offloading resource owners (agents) are self-interested and have private costs, it is highly challenging to design procurement mechanisms that motivate agents to reveal their true costs and achieve guaranteed performance under the constraint of a strict budget. In this paper, we model cellular traffic offloading as a multi-unit budget feasible procurement auction design problem with diminishing return valuations. We design a novel greedy-based randomized mechanism, and prove it is budget-feasible, truthful, individually rational and a (3 + 2 ln 𝑁)-approximation, where 𝑁 is the total number of available resource units. We also propose a deterministic mechanism which achieves (2 + ln 𝑁 + √︀ 2 + 3 ln 𝑁 + ln2 𝑁) - approximation. We prove no budget-feasible and truthful mechanism can do better than ln 𝑁-approximation in our setting, thus our mechanism approaches the optimal to a constant factor. In addition to solving the cellular traffic offloading problem, our work successfully extends solvable valuation class of greedy-based multi-unit budget-feasible mechanism with performance guarantees from the concave-additive valuations to more general local diminishing return valuations.

AAAI Conference 2019 Conference Paper

Optimal Projection Guided Transfer Hashing for Image Retrieval

  • Ji Liu
  • Lei Zhang

Recently, learning to hash has been widely studied for image retrieval thanks to the computation and storage efficiency of binary codes. For most existing learning to hash methods, sufficient training images are required and used to learn precise hashing codes. However, in some real-world applications, there are not always sufficient training images in the domain of interest. In addition, some existing supervised approaches need a amount of labeled data, which is an expensive process in terms of time, labor and human expertise. To handle such problems, inspired by transfer learning, we propose a simple yet effective unsupervised hashing method named Optimal Projection Guided Transfer Hashing (GTH) where we borrow the images of other different but related domain i. e. , source domain to help learn precise hashing codes for the domain of interest i. e. , target domain. Besides, we propose to seek for the maximum likelihood estimation (MLE) solution of the hashing functions of target and source domains due to the domain gap. Furthermore, an alternating optimization method is adopted to obtain the two projections of target and source domains such that the domain hashing disparity is reduced gradually. Extensive experiments on various benchmark databases verify that our method outperforms many state-of-the-art learning to hash methods. The implementation details are available at https: //github. com/liuji93/GTH.

AAAI Conference 2019 Conference Paper

TDSNN: From Deep Neural Networks to Deep Spike Neural Networks with Temporal-Coding

  • Lei Zhang
  • Shengyuan Zhou
  • Tian Zhi
  • Zidong Du
  • Yunji Chen

Continuous-valued deep convolutional networks (DNNs) can be converted into accurate rate-coding based spike neural networks (SNNs). However, the substantial computational and energy costs, which is caused by multiple spikes, limit their use in mobile and embedded applications. And recent works have shown that the newly emerged temporal-coding based SNNs converted from DNNs can reduce the computational load effectively. In this paper, we propose a novel method to convert DNNs to temporal-coding SNNs, called TDSNN. Combined with the characteristic of the leaky integrate-andfire (LIF) neural model, we put forward a new coding principle Reverse Coding and design a novel Ticking Neuron mechanism. According to our evaluation, our proposed method achieves 42% total operations reduction on average in large networks comparing with DNNs with no more than 0. 5% accuracy loss. The evaluation shows that TDSNN may prove to be one of the key enablers to make the adoption of SNNs widespread.

NeurIPS Conference 2019 Conference Paper

Variational Denoising Network: Toward Blind Noise Modeling and Removal

  • Zongsheng Yue
  • Hongwei Yong
  • Qian Zhao
  • Deyu Meng
  • Lei Zhang

Blind image denoising is an important yet very challenging problem in computer vision due to the complicated acquisition process of real images. In this work we propose a new variational inference method, which integrates both noise estimation and image denoising into a unique Bayesian framework, for blind image denoising. Specifically, an approximate posterior, parameterized by deep neural networks, is presented by taking the intrinsic clean image and noise variances as latent variables conditioned on the input noisy image. This posterior provides explicit parametric forms for all its involved hyper-parameters, and thus can be easily implemented for blind image denoising with automatic noise estimation for the test noisy image. On one hand, as other data-driven deep learning methods, our method, namely variational denoising network (VDN), can perform denoising efficiently due to its explicit form of posterior expression. On the other hand, VDN inherits the advantages of traditional model-driven approaches, especially the good generalization capability of generative models. VDN has good interpretability and can be flexibly utilized to estimate and remove complicated non-i. i. d. noise collected in real scenarios. Comprehensive experiments are performed to substantiate the superiority of our method in blind image denoising.

AAAI Conference 2018 Conference Paper

A Probabilistic Hierarchical Model for Multi-View and Multi-Feature Classification

  • Jinxing Li
  • Hongwei Yong
  • Bob Zhang
  • Mu Li
  • Lei Zhang
  • David Zhang

Some recent works in classification show that the data obtained from various views with different sensors for an object contributes to achieving a remarkable performance. Actually, in many real-world applications, each view often contains multiple features, which means that this type of data has a hierarchical structure, while most of existing works do not take these features with multi-layer structure into consideration simultaneously. In this paper, a probabilistic hierarchical model is proposed to address this issue and applied for classi- fication. In our model, a latent variable is first learned to fuse the multiple features obtained from a same view, sensor or modality. Particularly, mapping matrices corresponding to a certain view are estimated to project the latent variable from a shared space to the multiple observations. Since this method is designed for the supervised purpose, we assume that the latent variables associated with different views are influenced by their ground-truth label. In order to effectively solve the proposed method, the Expectation-Maximization (EM) algorithm is applied to estimate the parameters and latent variables. Experimental results on the extensive synthetic and two real-world datasets substantiate the effectiveness and superiority of our approach as compared with state-of-the-art.

AAAI Conference 2018 Conference Paper

Learning a Wavelet-Like Auto-Encoder to Accelerate Deep Neural Networks

  • Tianshui Chen
  • Liang Lin
  • Wangmeng Zuo
  • Xiaonan Luo
  • Lei Zhang

Accelerating deep neural networks (DNNs) has been attracting increasing attention as it can benefit a wide range of applications, e. g. , enabling mobile systems with limited computing resources to own powerful visual recognition ability. A practical strategy to this goal usually relies on a two-stage process: operating on the trained DNNs (e. g. , approximating the convolutional filters with tensor decomposition) and finetuning the amended network, leading to difficulty in balancing the trade-off between acceleration and maintaining recognition performance. In this work, aiming at a general and comprehensive way for neural network acceleration, we develop a Wavelet-like Auto-Encoder (WAE) that decomposes the original input image into two low-resolution channels (sub-images) and incorporate the WAE into the classification neural networks for joint training. The two decomposed channels, in particular, are encoded to carry the low-frequency information (e. g. , image profiles) and high-frequency (e. g. , image details or noises), respectively, and enable reconstructing the original input image through the decoding process. Then, we feed the low-frequency channel into a standard classification network such as VGG or ResNet and employ a very lightweight network to fuse with the high-frequency channel to obtain the classification result. Compared to existing DNN acceleration solutions, our framework has the following advantages: i) it is tolerant to any existing convolutional neural networks for classification without amending their structures; ii) the WAE provides an interpretable way to preserve the main components of the input image for classification.

AAMAS Conference 2018 Conference Paper

Optimal Constraint Collection for Core-Selecting Path Mechanism

  • Hao Cheng
  • Lei Zhang
  • Yi Zhang
  • Jun Wu
  • Chongjun Wang

In path auctions, strategic bidders make bids for commodities. Each edge of the graph stands for a commodity and the weight on the edge represents the prime cost. Auctioneer needs to purchase a sequence of edges in order to get a path from one vertex to another at a low cost. Path auctions can be considered as a kind of combinatorial reverse-auctions. Computing prices in core-selecting combinatorial auctions is a computationally hard problem, the same is true in core-selecting path auctions. This problem can be solved by core constraint generation(CCG) algorithm. However, we find that there are many redundant constraints and the constraint collection can be conciser in core-selecting path mechanism. In this paper, 1) we put forward a new approach to get the constraint collection, and reduce the constraint number from exponential O(2n) to polynomial O(n2), where n is the network diameter; 2) we prove that the new constraint collection is not only equivalent to the original collection, but also has no redundant constraint in the worst case; 3) we validate our approach on real-world datasets and obtain excellent results. Furthermore, we provide new insights to think over the core-selecting mechanism in combinatorial auctions.

IJCAI Conference 2018 Conference Paper

Social Media based Simulation Models for Understanding Disease Dynamics

  • Ting Hua
  • Chandan K Reddy
  • Lei Zhang
  • Lijing Wang
  • Liang Zhao
  • Chang-Tien Lu
  • Naren Ramakrishnan

In this modern era, infectious diseases, such as H1N1, SARS, and Ebola, are spreading much faster than any time in history. Efficient approaches are therefore desired to monitor and track the diffusion of these deadly epidemics. Traditional computational epidemiology models are able to capture the disease spreading trends through contact network, however, one unable to provide timely updates via real-world data. In contrast, techniques focusing on emerging social media platforms can collect and monitor real-time disease data, but do not provide an understanding of the underlying dynamics of ailment propagation. To achieve efficient and accurate real-time disease prediction, the framework proposed in this paper combines the strength of social media mining and computational epidemiology. Specifically, individual health status is first learned from user's online posts through Bayesian inference, disease parameters are then extracted for the computational models at population-level, and the outputs of computational epidemiology model are inversely fed into social media data based models for further performance improvement. In various experiments, our proposed model outperforms current disease forecasting approaches with better accuracy and more stability.

NeurIPS Conference 2018 Conference Paper

Turbo Learning for CaptionBot and DrawingBot

  • Qiuyuan Huang
  • Pengchuan Zhang
  • Dapeng Wu
  • Lei Zhang

We study in this paper the problems of both image captioning and text-to-image generation, and present a novel turbo learning approach to jointly training an image-to-text generator (a. k. a. CaptionBot) and a text-to-image generator (a. k. a. DrawingBot). The key idea behind the joint training is that image-to-text generation and text-to-image generation as dual problems can form a closed loop to provide informative feedback to each other. Based on such feedback, we introduce a new loss metric by comparing the original input with the output produced by the closed loop. In addition to the old loss metrics used in CaptionBot and DrawingBot, this extra loss metric makes the jointly trained CaptionBot and DrawingBot better than the separately trained CaptionBot and DrawingBot. Furthermore, the turbo-learning approach enables semi-supervised learning since the closed loop can provide peudo-labels for unlabeled samples. Experimental results on the COCO dataset demonstrate that the proposed turbo learning can significantly improve the performance of both CaptionBot and DrawingBot by a large margin.

AIIM Journal 2017 Journal Article

DisTeam: A decision support tool for surgical team selection

  • Ashkan Ebadi
  • Patrick J. Tighe
  • Lei Zhang
  • Parisa Rashidi

Objective Surgical service providers play a crucial role in the healthcare system. Amongst all the influencing factors, surgical team selection might affect the patients’ outcome significantly. The performance of a surgical team not only can depend on the individual members, but it can also depend on the synergy among team members, and could possibly influence patient outcome such as surgical complications. In this paper, we propose a tool for facilitating decision making in surgical team selection based on considering history of the surgical team, as well as the specific characteristics of each patient. Methods DisTeam (a decision support tool for surgical team selection) is a metaheuristic framework for objective evaluation of surgical teams and finding the optimal team for a given patient, in terms of number of complications. It identifies a ranked list of surgical teams personalized for each patient, based on prior performance of the surgical teams. DisTeam takes into account the surgical complications associated with teams and their members, their teamwork history, as well as patient’s specific characteristics such as age, body mass index (BMI) and Charlson comorbidity index score. Results We tested DisTeam using intra-operative data from 6065 unique orthopedic surgery cases. Our results suggest high effectiveness of the proposed system in a health-care setting. The proposed framework converges quickly to the optimal solution and provides two sets of answers: a) The best surgical team over all the generations, and b) The best population which consists of different teams that can be used as an alternative solution. This increases the flexibility of the system as a complementary decision support tool. Conclusion DisTeam is a decision support tool for assisting in surgical team selection. It can facilitate the job of scheduling personnel in the hospital which involves an overwhelming number of factors pertaining to patients, individual team members, and team dynamics and can be used to compose patient-personalized surgical teams with minimum (potential) surgical complications.

AAMAS Conference 2017 Conference Paper

Mechanism Design for Social Law Synthesis under Incomplete Information

  • Jun Wu
  • Lei Zhang
  • Chongjun Wang
  • Junyuan Xie

For the social law synthesis problem, when the agents are rational in the sense of game theory and hold some information we need as private information, it naturally evolves into a setting that is perfectly addressed by the framework of algorithmic mechanism design. In this strategic setting, we are not only required to find out the feasible social law for the objective, but also required to formulate the right payment to the agents to induce incentive compatibility and individual rationality. We design a mechanism for this setting, prove that it satisfies all the required formal properties, and characterize the conditions for the existence of feasible mechanisms. Moreover, we show that the upper-bound of the total payment of the proposed mechanism is high.

AAMAS Conference 2017 Conference Paper

Synthesizing Optimal Social Laws for Strategical Agents via Bayesian Mechanism Design

  • Jun Wu
  • Lei Zhang
  • Chongjun Wang
  • Junyuan Xie

When rational behavior of the agents and private information are considered, the optimal social law synthesizing problem naturally evolves into a setting which can be handled by the framework of algorithmic mechanism design. We focus on the Bayesian case in this paper, that is, the probability distribution of each agent’s cost is known. It is easy to see that in this case our problem closely relates to path/spanning-tree auctions and Myerson’s optimal auction mechanism, but the optimization objective is new, that is, we focus on profit maximization instead of payment maximization. By studying this problem: we further extend the logic-based framework of social law optimization problem to the strategic case, and show that it becomes a new problem of algorithmic mechanism design; we find out a mechanism that is incentive compatible, individually rational and maximizes the expected profit for all input cost profiles; however, we can show that this mechanism is computational intractable; so, we finally find out a tractable constant-factor approximation mechanism. CCS Concepts •Computing methodologies → Multi-agent systems;

IJCAI Conference 2016 Conference Paper

A Self-Representation Induced Classifier

  • Pengfei Zhu
  • Lei Zhang
  • Wangmeng Zuo
  • Xiangchu Feng
  • Qinghua Hu

Almost all the existing representation based classifiers represent a query sample as a linear combination of training samples, and their time and memory cost will increase rapidly with the number of training samples. We investigate the representation based classification problem from a rather different perspective in this paper, that is, we learn how each feature (i. e. , each element) of a sample can be represented by the features of itself. Such a self-representation property of sample features can be readily employed for pattern classification and a novel self-representation induced classifier (SRIC) is proposed. SRIC learns a self-representation matrix for each class. Given a query sample, its self-representation residual can be computed by each of the learned self-representation matrices, and classification can then be performed by comparing these residuals. In light of the principle of SRIC, a discriminative SRIC (DSRIC) method is developed. For each class, a discriminative self-representation matrix is trained to minimize the self-representation residual of this class while representing little the features of other classes. Experimental results on different pattern recognition tasks show that DSRIC achieves comparable or superior recognition rate to state-of-the-art representation based classifiers, however, it is much more efficient and needs much less storage space.

YNIMG Journal 2014 Journal Article

Detection of optical neuronal signals in the visual cortex using continuous wave near-infrared spectroscopy

  • Bailei Sun
  • Lei Zhang
  • Hui Gong
  • Jinyan Sun
  • Qingming Luo

Near-infrared spectroscopy (NIRS) measures slow hemodynamic signals noninvasively to indirectly infer the neuronal activity in the brain. However, it remains a controversy on whether this optical measurement technique can detect the optical neuronal signal, which reflects the optical changes directly associated with neuronal activity, within the visual cortex of human and non-human primates. By carefully reviewing the important factors in the detection of optical neuronal signals, we aim to investigate the feasibility of performing NIRS measurements of optical neuronal signals within the visual cortex in humans. To ensure a strong optical neuronal response, a full-field circular black and white reversing checkerboard stimulus was presented, and the reversal frequency was carefully chosen. We used a homemade continuous wave (CW) NIRS system with high detection sensitivity (of the order of 0. 1pW) to record a large area of the visual cortex (approximately 6×14cm2). EEG was simultaneously acquired with the optical signal. Based on the mathematical morphology, we adapted the filter proposed by Gratton et al. to remove the influence of arterial pulsation and facilitate the detection and elimination of unknown artifacts from the data. We obtained reliable optical neuronal signals in 77% of the participants (10 out of 13). The amplitudes (latencies) of the obtained optical neuronal signals corresponding to the 785 and 850nm wavelengths were 0. 017±0. 003% (94. 7±8. 4ms) and 0. 025±0. 006% (99. 0±7. 7ms), respectively. There were no significant differences between the latencies of the N75 component of the visual evoked potential (VEP) and optical neuronal signals at either wavelength. This is the first study to report optical neuronal signals within the visual cortex in the intact human brain using a CW NIRS system. These results indicate the feasibility of measuring noninvasive optical neuronal signals using a CW NIRS system with high detection sensitivity.

NeurIPS Conference 2014 Conference Paper

Projective dictionary pair learning for pattern classification

  • Shuhang Gu
  • Lei Zhang
  • Wangmeng Zuo
  • Xiangchu Feng

Discriminative dictionary learning (DL) has been widely studied in various pattern classification problems. Most of the existing DL methods aim to learn a synthesis dictionary to represent the input signal while enforcing the representation coefficients and/or representation residual to be discriminative. However, the $\ell_0$ or $\ell_1$-norm sparsity constraint on the representation coefficients adopted in many DL methods makes the training and testing phases time consuming. We propose a new discriminative DL framework, namely projective dictionary pair learning (DPL), which learns a synthesis dictionary and an analysis dictionary jointly to achieve the goal of signal representation and discrimination. Compared with conventional DL methods, the proposed DPL method can not only greatly reduce the time complexity in the training and testing phases, but also lead to very competitive accuracies in a variety of visual classification tasks.

TCS Journal 2014 Journal Article

Sufficient conditions for k-restricted edge connected graphs

  • Shiying Wang
  • Lei Zhang

For a connected graph G = ( V, E ), an edge set S ⊆ E is a k-restricted edge cut if G − S is disconnected and every component of G − S has at least k vertices. The k-restricted edge connectivity of G, denoted by λ k ( G ), is defined as the cardinality of a minimum k-restricted edge cut. Let ξ k ( G ) = min ⁡ { | [ X, X ¯ ] |: | X | = k, G [ X ] is connected }, where X ¯ = V \ X. G is maximally k-restricted edge connected ( λ k -optimal for short) if λ k ( G ) = ξ k ( G ). The k-restricted edge connectivity is more refined network reliability indices than edge connectivity. In this paper, let k ≥ 2 be an integer, and let G be a graph of order ν ( G ) at least 2k satisfying | N ( u ) ∩ N ( v ) | ≥ 2 k − 2 for all pairs u, v of nonadjacent vertices. If for each triangle T there exists at least one vertex v ∈ V ( T ) such that d ( v ) ≥ ⌊ ν ( G ) 2 ⌋ + k − 1, then G is λ k -optimal.

AAAI Conference 2013 Conference Paper

A Cyclic Weighted Median Method for L1 Low-Rank Matrix Factorization with Missing Entries

  • Deyu Meng
  • Zongben Xu
  • Lei Zhang
  • Ji Zhao

A challenging problem in machine learning, information retrieval and computer vision research is how to recover a low-rank representation of the given data in the presence of outliers and missing entries. The L1-norm low-rank matrix factorization (LRMF) has been a popular approach to solving this problem. However, L1-norm LRMF is difficult to achieve due to its non-convexity and non-smoothness, and existing methods are often inefficient and fail to converge to a desired solution. In this paper we propose a novel cyclic weighted median (CWM) method, which is intrinsically a coordinate decent algorithm, for L1-norm LRMF. The CWM method minimizes the objective by solving a sequence of scalar minimization sub-problems, each of which is convex and can be easily solved by the weighted median filter. The extensive experimental results validate that the CWM method outperforms state-of-the-arts in terms of both accuracy and computational efficiency.

NeurIPS Conference 2012 Conference Paper

Generalization Bounds for Domain Adaptation

  • Chao Zhang
  • Lei Zhang
  • Jieping Ye

In this paper, we provide a new framework to study the generalization bound of the learning process for domain adaptation. Without loss of generality, we consider two kinds of representative domain adaptation settings: one is domain adaptation with multiple sources and the other is domain adaptation combining source and target data. In particular, we introduce two quantities that capture the inherent characteristics of domains. For either kind of domain adaptation, based on the two quantities, we then develop the specific Hoeffding-type deviation inequality and symmetrization inequality to achieve the corresponding generalization bound based on the uniform entropy number. By using the resultant generalization bound, we analyze the asymptotic convergence and the rate of convergence of the learning process for such kind of domain adaptation. Meanwhile, we discuss the factors that affect the asymptotic behavior of the learning process. The numerical experiments support our results.

AAAI Conference 2012 Conference Paper

MAXSAT Heuristics for Cost Optimal Planning

  • Lei Zhang
  • Fahiem Bacchus

The cost of an optimal delete relaxed plan, known as h+, is a powerful admissible heuristic but is in general intractable to compute. In this paper we examine the problem of computing h+ by encoding it as a MAXSAT problem. We develop a new encoding that utilizes constraint generation to supports the computation of a sequence of increasing lower bounds on h+. We show a close connection between the computations performed by a recent approach to solving MAXSAT and a hitting set approach recently proposed for computing h+. Using this connection we observe that our MAXSAT computation can be initialized with a set of landmarks computed via cheaper methods like LM-cut. By judicious use of MAXSAT solving along with a technique of lazy heuristic evaluation we obtain speedups for finding optimal plans over LM-cut on a number of domains. Our approach enables the exploitation of continued progress in MAXSAT solving, and also makes it possible to consider computing or approximating heuristics that are even more informed that h+ by, for example, adding some information about deletes back into the encoding.

AAAI Conference 2011 Conference Paper

Identifying Evaluative Sentences in Online Discussions

  • Zhongwu Zhai
  • Bing Liu
  • Lei Zhang
  • Hua Xu
  • Peifa Jia

Much of opinion mining research focuses on product reviews because reviews are opinion-rich and contain little irrelevant information. However, this cannot be said about online discussions and comments. In such postings, the discussions can get highly emotional and heated with many emotional statements, and even personal attacks. As a result, many of the postings and sentences do not express positive or negative opinions about the topic being discussed. To find people’s opinions on a topic and its different aspects, which we call evaluative opinions, those irrelevant sentences should be removed. The goal of this research is to identify evaluative opinion sentences. A novel unsupervised approach is proposed to solve the problem, and our experimental results show that it performs well.

TCS Journal 2010 Journal Article

Dykstra’s algorithm for constrained least-squares doubly symmetric matrix problems

  • Jiao-fen Li
  • Xi-yan Hu
  • Lei Zhang

In this work we apply Dykstra’s alternating projection algorithm for minimizing ‖ A X − B ‖ where ‖ ⋅ ‖ is the Frobenius norm and A ∈ R m × n, B ∈ R m × n and X ∈ R n × n are doubly symmetric positive definite matrices with entries within prescribed intervals. We first solve the constrained least-squares matrix problem by using the special structure properties of doubly symmetric matrices, and then use the singular value decomposition to transform the original problem into a simpler one that fits nicely with the algorithm originally developed by [R. Escalante, M. Raydan, Dykstra’s algorithm for a constrained least-squares matrix problem, Numer. Linear Algebra Appl. 3 (1996) 459–471].

YNIMG Journal 2007 Journal Article

The effect of practice on a sustained attention task in cocaine abusers

  • Rita Z. Goldstein
  • Dardo Tomasi
  • Nelly Alia-Klein
  • Lei Zhang
  • Frank Telang
  • Nora D. Volkow

Habituation enables the organism to attend selectively to novel stimuli by diminishing no-longer necessary responses to repeated stimuli. Because the prefrontal cortex (PFC) has a core role in monitoring attention and behavioral control especially under novelty, neural habituation responses may be modified in drug addiction, a psychopathology that entails PFC abnormalities in both structure and function. Sixteen cocaine abusers and 12 gender-, race-, education-, and intelligence-matched healthy control subjects performed an incentive sustained attention task twice, under novelty and after practice, during functional magnetic resonance imaging. For cocaine abusers practice effects were noted in the PFC (including anterior cingulate cortex/ventromedial rostral PFC, dorsolateral PFC, and medial frontal gyrus) and cerebellum (signal attenuations/decreases: return to baseline); activations in these regions were associated with craving, frequency of use, and length of abstinence. In the control subjects practice effects were instead restricted to posterior brain regions (precuneus and cuneus) (signal amplifications/increases: deactivation away from baseline). Also, only in the cocaine abusers, increased speed of behavioral performance between novelty to practice was associated with a respective attenuation of activation in the thalamus. Overall, we report for the first time a differential pattern of neural responses to repeated presentation of an incentive sustained attention task in cocaine addiction. Our results suggest a disruption in drug addiction of neural habituation to practice that possibly encompasses opponent anterior vs. posterior brain adaptation to the novelty of the experience: overly expeditious for the former but overly protracted for the latter. Overall, cocaine addicted individuals may be predisposed to an increased challenge when required to maintain alertness as a task progresses, not able to optimally utilize a prematurely habituating PFC to compensate with an increased attribution of salience to a desired reward.

IROS Conference 2006 Conference Paper

A Visual Tele-operation System for the Humanoid Robot BHR-02

  • Lei Zhang
  • Qiang Huang
  • Yuepin Lu
  • Tao Xiao
  • Jiapeng Yang
  • Muhammad Usman Keerio

This paper presents a method to reconstruct the virtual scene of the humanoid robot teleoperation to overcome the problem of the poor vision images feedback. A virtual robot interface is built to render the data of the real robot. It has the same DOF set and the same size scale as the real robot. The interface can render the multiple real-time feedback data from the robot. In the data-fusion module an algorithm is adopted to determine the position and attitude of the robot body. Some experiments are done to confirm the effectiveness of the virtual scene

YNIMG Journal 2006 Journal Article

Dissociation in the neural basis underlying Chinese tone and vowel production

  • Li Liu
  • Danling Peng
  • Guosheng Ding
  • Zhen Jin
  • Lei Zhang
  • Ke Li
  • Chuansheng Chen

Neuropsychologists have debated over whether the processing of segmental and suprasegmental units involves different neural mechanisms. Focusing on the production of Chinese lexical tones (suprasegmental units) and vowels (segmental units), this study used the adaptation paradigm to investigate a possible neural dissociation for tone and vowel production. Ten native Chinese speakers were asked to name Chinese characters and pinyin (Romanized phonetic system for Chinese language) that varied in terms of tones and vowels. fMRI results showed significant differences in the right inferior frontal gyrus between tone and vowel production (more activation for tones than for vowels). Brain asymmetry analysis further showed that tone production was less left-lateralized than vowel production, although both showed left-hemisphere dominance.

IROS Conference 2006 Conference Paper

Several Insights into Omnidirectional Static Walking of a Quadruped Robot on a slope

  • Lei Zhang
  • Shugen Ma
  • Kousuke Inoue

Several insights gained from observing quadruped robots in omnidirectional static walking experiments at high speed on a slope are presented here. In order to allow a robot to move as fast as possible, the height of center of gravity (COG) and three rotating axes, i. e. , roll, pitch, and yaw, were used to discuss the COG with the corresponding optimal body posture (COBP). The COBP is the posture in which the motion velocity is maximized according to the height of COG, stability, degree of slope, and direction. Successive gait transition with a minimum number of steps is achievable with the use of a common foot position before and after a gait transition. The time required to change gaits may be reduced by designing the foot positions during crawling and rotating while limiting reachable region of the foot on a slope. The robot, thus, walks in any direction fast and statically with COBP by dynamically changing the height of COG and body posture during gait transitions

NeurIPS Conference 2005 Conference Paper

Modeling Neuronal Interactivity using Dynamic Bayesian Networks

  • Lei Zhang
  • Dimitris Samaras
  • Nelly Alia-Klein
  • Nora Volkow
  • Rita Goldstein

Functional Magnetic Resonance Imaging (fMRI) has enabled scientists to look into the active brain. However, interactivity between functional brain regions, is still little studied. In this paper, we contribute a novel framework for modeling the interactions between multiple active brain regions, using Dynamic Bayesian Networks (DBNs) as generative mod- els for brain activation patterns. This framework is applied to modeling of neuronal circuits associated with reward. The novelty of our frame- work from a Machine Learning perspective lies in the use of DBNs to reveal the brain connectivity and interactivity. Such interactivity mod- els which are derived from fMRI data are then validated through a group classification task. We employ and compare four different types of DBNs: Parallel Hidden Markov Models, Coupled Hidden Markov Models, Fully-linked Hidden Markov Models and Dynamically Multi- Linked HMMs (DML-HMM). Moreover, we propose and compare two schemes of learning DML-HMMs. Experimental results show that by using DBNs, group classification can be performed even if the DBNs are constructed from as few as 5 brain regions. We also demonstrate that, by using the proposed learning algorithms, different DBN structures charac- terize drug addicted subjects vs. control subjects. This finding provides an independent test for the effect of psychopathology on brain function. In general, we demonstrate that incorporation of computer science prin- ciples into functional neuroimaging clinical studies provides a novel ap- proach for probing human brain function.

ICRA Conference 2005 Conference Paper

Omni-directional Walking of a Quadruped Robot with Optimal Body Postures on a Slope

  • Lei Zhang
  • Shugen Ma
  • Kousuke Inoue
  • Yoshinori Honda

In this paper, we discuss the optimal body postures of a quadruped robot to perform omni-directional static walking on a slope. The optimal body posture is the posture with the maximum possible moving speed w. r. t. slope and moving direction. The proposed method based on dynamically changing body posture during gait-transitions, is used to maintain high robot motion velocity on slope. The timing of changing body posture is designed by considering the stability during gait-transition. Using the proposed method, the robot can walk into any direction with the fastest moving speed on a slope. Through walking experiments by computer simulation, the validity of the proposed method has been verified.

ICRA Conference 2004 Conference Paper

Robust Neuro-fuzzy Navigation of Mobile Manipulator among Dynamic Obstacles

  • Jean Bosco Mbede
  • Shugen Ma
  • Lei Zhang
  • Youssoufi Touré
  • Volker Graefe

To fit well the needs of autonomous mobile manipulator, two robust adaptive Neuro-Fuzzy motion controllers are developed. The first controller, based on a computational efficient processing scheme for fuzzy reactive navigation, is used to generate the commands for the servo-systems of robot arm so that, locally, it may choose its way to its goal autonomously. The second fuzzy reactive navigation is implemented in mobile platform so that it maintains a permanent flexible path between two nodes in network generated by a probabilistic roadmap approach. In order to consider the compatibility of stabilisation, mobilisation and manipulation, we derive a coordinated fuzzy local planner algorithm so that the mobile manipulator can avoid stably unknown and/or dynamic obstacles The purpose of an integration of robust controller and Modified Elman Neural Network is to deal with unmodeled bounded disturbances and/or unstructured unmodeled dynamics.

v2026.09.13