Arrow Research search

Author name cluster

Lu Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

48 papers
2 author rows

Possible papers

48

TMLR Journal 2026 Journal Article

Algorithmic Recourse in Abnormal Multivariate Time Series

  • Xiao Han
  • Lu Zhang
  • Yongkai Wu
  • Shuhan Yuan

Algorithmic recourse provides actionable recommendations to alter unfavorable predictions of machine learning models, enhancing transparency through counterfactual explanations. While significant progress has been made in algorithmic recourse for static data, such as tabular and image data, limited research explores recourse for multivariate time series, particularly for reversing abnormal time series. This paper introduces Recourse in time series Anomaly Detection (RecAD), a framework for addressing anomalies in multivariate time series using backtracking counterfactual reasoning. By modeling the causes of anomalies as external interventions on exogenous variables, RecAD predicts recourse actions to restore normal status as counterfactual explanations, where the recourse function, responsible for generating actions based on observed data, is trained using an end-to-end approach. Experiments on synthetic and real-world datasets demonstrate its effectiveness.

AAAI Conference 2026 Conference Paper

Breaking Model Lock-in: Cost-Efficient Zero-Shot LLM Routing via a Universal Latent Space

  • Cheng Yan
  • Wuyang Zhang
  • Zhiyuan Ning
  • Fan Xu
  • Ziyang Tao
  • Lu Zhang
  • Bing Yin
  • Yanyong Zhang

The rapid proliferation of Large Language Models (LLMs) has led to a fragmented and inefficient ecosystem, a state of ``model lock-in'' where seamlessly integrating novel models remains a significant bottleneck. Current routing frameworks require exhaustive, costly retraining, hindering scalability and adaptability. We introduce ZeroRouter, a new paradigm for LLM routing that breaks this lock-in. Our approach is founded on a universal latent space, a model-agnostic representation of query difficulty that fundamentally decouples the characterization of a query from the profiling of a model. This allows for zero-shot onboarding of new models without full-scale retraining. ZeroRouter features a context-aware predictor that maps queries to this universal space and a dual-mode optimizer that balances accuracy, cost, and latency. Our framework consistently outperforms all baselines, delivering higher accuracy at lower cost and latency.

AAAI Conference 2026 Conference Paper

Deeply Seeking Boundary for Lunar Regolith Segmentation

  • Yifeng Wang
  • Lingxin Wang
  • Lu Zhang
  • Yang Li
  • Chao Xu
  • Weiwei Zhang
  • Junyue Tang
  • Yanhong Zheng

The sharp, intricate contours of lunar regolith particles hold critical clues to the Moon's geological evolution and inform engineering applications from habitat construction to spacecraft design, making their precise segmentation a task of significant scientific and engineering value. However, this task exposes a weakness in deep learning models known as spectral bias, an inherent tendency to learn smooth, low-frequency functions which causes them to systematically erase the very high-frequency boundary details that are of primary interest. To resolve this conflict, we propose a framework to deeply seek object boundaries. First, we propose High-Frequency Initialized LoRA (HiFi-LoRA) to counteract spectral bias. By initializing the LoRA adaptation matrices as the optimal low-rank approximation of a high-pass filter, it fundamentally enhances the model's high-frequency perception and injects a strong preference for edges. Second, we propose the Wavelet Energy Modulation (WEM) regularizer. It guides the model to learn the intrinsic correlation between contour complexity and mask area, forcing the model to build a geometric understanding of contour morphology upon its high-frequency perception, thereby enabling the generation of boundary details commensurate with the object's scale. Experimentally, we constructed the Lunar Regolith Segmentation Dataset (LRSD), the first large-scale benchmark with expert-annotated contours. Extensive experiments demonstrate that our method sets a new state of the art on this challenging benchmark, not only achieving top performance on regional metrics like mIoU and DSC but, more critically, drastically outperforming existing models on boundary accuracy. This work not only provides a powerful computational tool for lunar science but also offers a robust and synergistic design pattern for other fine-grained segmentation challenges.

AAAI Conference 2026 Conference Paper

MSAT-LDM: Toward Transferable High-Fidelity Watermarking for Latent Diffusion Model via Modular Self-Augmented Training

  • Lu Zhang
  • Liang Zeng

The rapid proliferation of AI-generated images necessitates effective watermarking techniques to protect intellectual property and detect fraudulent content. While existing training-based watermarking methods show promise, they often struggle with generalization across diverse prompts, introduce visible artifacts, and require substantial external data for retraining on new model variants. To this end, we propose Modular Self-Augmented Training for Latent Diffusion Models (MSAT-LDM), a novel and transferable watermarking framework. MSAT-LDM integrates two key components: (1) Self-Augmented Training (SAT) leverages an internally generated "free generation" distribution to train the watermark module, aligning the training and testing phases without relying on external data. We theoretically demonstrate that this design improves generalization by inducing a tighter generalization bound. (2) Modular watermark architecture is a plug-and-play module that can be independently fine-tuned, enabling efficient adaptation to various fine-tuned backbones or LoRA-enhanced variants with minimal overhead. Extensive experiments show that MSAT-LDM achieves robust watermarking, significantly improves the quality of watermarked images across diverse prompts, and exhibits strong transfer performance--all without the need for external training data.

AAAI Conference 2025 Conference Paper

Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding

  • Wenbo Zhang
  • Lu Zhang
  • Ping Hu
  • Liqian Ma
  • Yunzhi Zhuge
  • Huchuan Lu

Injecting semantics into 3D Gaussian Splatting (3DGS) has recently garnered significant attention. While current approaches typically distill 3D semantic features from 2D foundational models (e.g., CLIP and SAM) to facilitate novel view segmentation and semantic understanding, their heavy reliance on 2D supervision can undermine cross-view semantic consistency and necessitate complex data preparation processes, therefore hindering view-consistent scene understanding. In this work, we present FreeGS, an unsupervised semantic-embedded 3DGS framework that achieves view-consistent 3D scene understanding without the need for 2D labels. Instead of directly learning semantic features, we introduce the IDentity-coupled Semantic Field (IDSF) into 3DGS, which captures both semantic representations and view-consistent instance indices for each Gaussian. We optimize IDSF with a two-step alternating strategy: semantics help to extract coherent instances in 3D space, while the resulting instances regularize the injection of stable semantics from 2D space. Additionally, we adopt a 2D-3D joint contrastive loss to enhance the complementarity between view-consistent 3D geometry and rich semantics during the bootstrapping process, enabling FreeGS to uniformly perform tasks such as novel-view semantic segmentation, object selection, and 3D object detection. Extensive experiments on LERF-Mask, 3D-OVS, and ScanNet datasets demonstrate that FreeGS performs comparably to state-of-the-art methods while avoiding the complex data preprocessing workload.

YNIMG Journal 2025 Journal Article

Exploring the impact of APOE ɛ4 on functional connectivity in Alzheimer’s disease across cognitive impairment levels

  • Kangli Dong
  • Wei Liang
  • Ting Hou
  • Zhijie Lu
  • Yixuan Hao
  • Chenrui Li
  • Yue Qiu
  • Nan Kong

The apolipoprotein E (APOE) ɛ4 allele is a recognized genetic risk factor for Alzheimer's Disease (AD). Studies have shown that APOE ɛ4 mediates the modulation of intrinsic functional brain networks in cognitively normal individuals and significantly disrupts the whole-brain topological structure in AD patients. However, how APOE ɛ4 regulates brain functional connectivity (FC) and consequently affects the levels of cognitive impairment in AD patients remains unknown. In this study, we systematically analyzed functional magnetic resonance imaging (fMRI) data from two distinct cohorts: an In-house dataset includes 59 AD patients (73.37 ± 6.42 years), and the ADNI dataset includes 117 AD patients (74.91 ± 7.91 years). Experimental comparisons were conducted by grouping AD patients based on both APOE ɛ4 status and cognitive impairment levels of AD. Network-Based Statistic (NBS) method and the Graph Neural Network Explainer (GNN-Explainer) were combined to identify significant FC changes across different comparisons. Importantly, the GNN-Explainer method was introduced as an enhancement over the NBS method to better model complex high-order nonlinear characteristics for discovering FC features that significantly contribute to classification tasks. The results showed that APOE ɛ4 primarily influenced temporal lobe FCs, while it influenced different cognitive impairment levels of AD by adjusting prefrontal-parietal FCs. These findings were validated by p-values < 0.05 from NBS method, and 5-fold cross-validation along with ablation studies from the GNN-Explainer method. In conclusion, our findings provide new insights into the role of APOE ɛ4 in altering FC dynamics during the progression of AD, highlighting potential targets for early intervention.

NeurIPS Conference 2025 Conference Paper

FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning

  • Lu Zhang
  • Jiazuo Yu
  • Haomiao Xiong
  • Ping Hu
  • Yunzhi Zhuge
  • Huchuan Lu
  • You He

Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and localizing visual details in high-resolution images---particularly when dealing with extra-small objects embedded in cluttered contexts. To address this issue, we propose FineRS, a two-stage MLLM-based reinforcement learning framework for jointly reasoning and segmenting extremely small objects within high-resolution scenes. FineRS adopts a coarse-to-fine pipeline comprising Global Semantic Exploration (GSE) and Localized Perceptual Refinement (LPR). Specifically, GSE performs instruction-guided reasoning to generate a textural response and a coarse target region, while LPR refines this region to produce an accurate bounding box and segmentation mask. To couple the two stages, we introduce a locate-informed retrospective reward, where LPR's outputs are used to optimize GSE for more robust coarse region exploration. Additionally, we present FineRS-4k, a new dataset for evaluating MLLMs on attribute-level reasoning and pixel-level segmentation on subtle, small-scale targets in complex high-resolution scenes. Experimental results on FineRS-4k and public datasets demonstrate that our method consistently outperforms state-of-the-art MLLM-based approaches on both instruction-guided segmentation and visual reasoning tasks.

YNIMG Journal 2025 Journal Article

How spontaneous brain activity encodes the observation of grasping movements

  • Cristina Perciballi
  • Lorenzo Pini
  • Daniele Sili
  • Yara El Rassi
  • Lu Zhang
  • Giacomo Handjaras
  • Federico Giove
  • Emiliano Ricciardi

Spontaneous brain activity forms correlated networks resembling task-evoked activation patterns, yet its functional relevance remains debated. The representational hypothesis suggests that resting-state networks (RSNs) encode frequent behaviors, but whether these representations are motor-based or cognitive is unclear. Here, we used fMRI to examine RSNs activity during the observation of reach-to-grasp movements with either regular (common) or perturbed (uncommon) kinematics. We found that the dorsal attention network (DAN) exhibited greater similarity between rest and task patterns for common movements, whereas sensory networks showed no significant effects. While DAN is classically associated with attention mechanisms, these results suggest that it may also contribute to tracking the location or motion of the hand. Furthermore, uncommon movements elicited stronger activation in parietal and premotor areas, likely reflecting adaptive updating of internal models. Our findings support the role of spontaneous brain activity in maintaining cognitive representations of frequent behaviors, optimizing motor planning and perception.

ICRA Conference 2025 Conference Paper

MT-PCR: Leveraging Modality Transformation for Large-Scale Point Cloud Registration with Limited Overlap

  • Yilong Wu
  • Yifan Duan
  • Yuxi Chen
  • Xinran Zhang
  • Yedong Shen
  • Jianmin Ji
  • Yanyong Zhang
  • Lu Zhang

Large-scale scene point cloud registration with limited overlap is a challenging task due to computational load and constrained data acquisition. To tackle these issues, we propose a point cloud registration method, MT-PCR, based on Modality Transformation. MT-PCR leverages a Bird's Eye View (BEV) capturing the maximal overlap information to improve the accuracy and utilizes images to provide complementary spatial features. Specifically, MT-PCR converts 3D point clouds to BEV images and estimates correspondence by 2D image keypoints extraction and matching. Subsequently, the 2D correspondence estimates are then transformed back to 3D point clouds using inverse mapping. We have applied MT-PCR to Terrestrial Laser Scanning (TLS) and Aerial Laser Scanning (ALS) point cloud registration on the GrAco dataset, involving 8 low-overlap, square-kilometer scale registration scenarios. Experiments and comparisons with commonly used methods demonstrate that MT-PCR can achieve superior accuracy and robustness in large-scale scenes with limited overlap.

ICML Conference 2025 Conference Paper

Optimal Information Retention for Time-Series Explanations

  • Jinghang Yue
  • Jing Wang 0060
  • Lu Zhang
  • Shuo Zhang 0015
  • Da Li
  • Zhaoyang Ma
  • Youfang Lin

Explaining deep models for time-series data is crucial for identifying key patterns in sensitive domains, such as healthcare and finance. However, due to the lack of unified optimization criterion, existing explanation methods often suffer from redundancy and incompleteness, where irrelevant patterns are included or key patterns are missed in explanations. To address this challenge, we propose the Optimal Information Retention Principle, where conditional mutual information defines minimizing redundancy and maximizing completeness as optimization objectives. We then derive the corresponding objective function theoretically. As a practical framework, we introduce an explanation framework ORTE, learning a binary mask to eliminate redundant information while mining temporal patterns of explanations. We decouple the discrete mapping process to ensure the stability of gradient propagation, while employing contrastive learning to achieve precise filtering of explanatory patterns through the mask, thereby realizing a trade-off between low redundancy and high completeness. Extensive quantitative and qualitative experiments on synthetic and real-world datasets demonstrate that the proposed principle significantly improves the accuracy and completeness of explanations compared to baseline methods. The code is available at https: //github. com/moon2yue/ORTE_public.

EAAI Journal 2025 Journal Article

Possibilistic c-means clustering approach based on a novel weighted-kernel distance for imbalanced images with minority targets in sparsely distribution

  • Haiyan Yu
  • Yuting Wu
  • Haocong Zheng
  • Qianqian Luo
  • Lu Zhang

Possibilistic c-means clustering (PCM), a classic partition clustering algorithm, is an important data mining technique for unsupervised image segmentation in the field of artificial intelligence. The typicalities (memberships) of PCM have better description of local information and excellent noise resistance due to its absolute attribute. However, it still faces the challenges of segmenting color images with multiple characteristics, such as feature imbalance, cluster-size imbalance, noise attack, especially sparsely distribution of minority targets in the feature space. Therefore, this paper introduces a weighted-kernel distance-based PCM (WK-PCM) algorithm. Firstly, a weighted-kernel distance (WK-distance) is defined by combining the absolute attribute of typicalities and the Gaussian kernel function to enhance the intra-class compactness of sparse targets. Meanwhile, the feature weighting scheme in the WK-distance is expected to avoid center offset caused by imbalanced features. Secondly, to overcome the issue of center overlapping caused by insufficient interclass relationships of typicalities, the cutset theory is introduced based on the WK-distance to select partial objects. Then the typicalities of these selected objects are suppressed to increase inter-class separateness and avoid the issue of coincident clustering (also called center overlapping). Finally, an improved WK-PCM image segmentation algorithm (LWK-PCM) based on local spatial information acquired through bilateral filtering is proposed for imbalanced color images with noise corruption. Experiments conducted on synthetic datasets and imbalanced color images indicate that the proposed WK-PCM and LWK-PCM algorithms get excellent clustering performance compared to the relevant clustering algorithms.

ICLR Conference 2025 Conference Paper

Root Cause Analysis of Anomalies in Multivariate Time Series through Granger Causal Discovery

  • Xiao Han
  • Saima Absar
  • Lu Zhang
  • Shuhan Yuan

Identifying the root causes of anomalies in multivariate time series is challenging due to the complex dependencies among the series. In this paper, we propose a comprehensive approach called AERCA that inherently integrates Granger causal discovery with root cause analysis. By defining anomalies as interventions on the exogenous variables of time series, AERCA not only learns the Granger causality among time series but also explicitly models the distributions of exogenous variables under normal conditions. AERCA then identifies the root causes of anomalies by highlighting exogenous variables that significantly deviate from their normal states. Experiments on multiple synthetic and real-world datasets demonstrate that AERCA can accurately capture the causal relationships among time series and effectively identify the root causes of anomalies.

EAAI Journal 2024 Journal Article

Almost sure exponential synchronization analysis of stochastic strict-feedback systems with semi-Markov jump

  • Chang Gao
  • Lu Zhang
  • Haiying Zhang
  • Yu Xiao

In this article, we address the almost sure exponential synchronization issue of stochastic strict-feedback systems (SSFSs) with semi-Markov jump. It is the first time to consider semi-Markov jump into SSFSs. Different from previous works, the sojourn time of semi-Markov jump is influenced by the present and next states. Meanwhile, with the help of multiple mode-dependent Lyapunov-like functions, the relevant conditions for linear comparability become more relaxed. We propose a new controller design based on the idea of backstepping due to the introduction of a semi-Markov jump in the system. Then, we calculate differential inequalities for Lyapunov functions. Moreover, based on Lyapunov functions, differential inequality techniques, and the design controllers, some sufficient criteria are required to achieve almost sure exponential synchronization. In addition, the theoretical results are extended to SSFSs with Markov jump and some sufficient conditions are attained. Finally, a specific numerical example and an application about single-link robot manipulator are employed to demonstrate the effectiveness and feasibility of our results.

ICML Conference 2024 Conference Paper

Harnessing the Power of Neural Operators with Automatically Encoded Conservation Laws

  • Ning Liu 0019
  • Yiming Fan
  • Xianyi Zeng
  • Milan Klöwer
  • Lu Zhang
  • Yue Yu 0011

Neural operators (NOs) have emerged as effective tools for modeling complex physical systems in scientific machine learning. In NOs, a central characteristic is to learn the governing physical laws directly from data. In contrast to other machine learning applications, partial knowledge is often known a priori about the physical system at hand whereby quantities such as mass, energy and momentum are exactly conserved. Currently, NOs have to learn these conservation laws from data and can only approximately satisfy them due to finite training data and random noise. In this work, we introduce conservation law-encoded neural operators (clawNOs), a suite of NOs that endow inference with automatic satisfaction of such conservation laws. ClawNOs are built with a divergence-free prediction of the solution field, with which the continuity equation is automatically guaranteed. As a consequence, clawNOs are compliant with the most fundamental and ubiquitous conservation laws essential for correct physical consistency. As demonstrations, we consider a wide variety of scientific applications ranging from constitutive modeling of material deformation, incompressible fluid dynamics, to atmospheric simulation. ClawNOs significantly outperform the state-of-the-art NOs in learning efficacy, especially in small-data regimes.

NeurIPS Conference 2024 Conference Paper

LLMs Can Evolve Continually on Modality for $\mathbb{X}$-Modal Reasoning

  • Jiazuo Yu
  • Haomiao Xiong
  • Lu Zhang
  • Haiwen Diao
  • Yunzhi Zhuge
  • Lanqing Hong
  • Dong Wang
  • Huchuan Lu

Multimodal Large Language Models (MLLMs) have gained significant attention due to their impressive capabilities in multimodal understanding. However, existing methods rely heavily on extensive modal-specific pretraining and joint-modal tuning, leading to significant computational burdens when expanding to new modalities. In this paper, we propose \textbf{PathWeave}, a flexible and scalable framework with modal-\textbf{path} s\textbf{w}itching and \textbf{e}xp\textbf{a}nsion abilities that enables MLLMs to continually \textbf{ev}olve on modalities for $\mathbb{X}$-modal reasoning. We leverage the concept of Continual Learning and develop an incremental training strategy atop pre-trained MLLMs, enabling their expansion to new modalities using uni-modal data, without executing joint-modal pretraining. In detail, a novel Adapter-in-Adapter (AnA) framework is introduced, in which uni-modal and cross-modal adapters are seamlessly integrated to facilitate efficient modality alignment and collaboration. Additionally, an MoE-based gating module is applied between two types of adapters to further enhance the multimodal interaction. To investigate the proposed method, we establish a challenging benchmark called \textbf{C}ontinual \textbf{L}earning of \textbf{M}odality (MCL), which consists of high-quality QA data from five distinct modalities: image, video, \textcolor{black}{audio, depth} and point cloud. Extensive experiments demonstrate the effectiveness of the proposed AnA framework on learning plasticity and memory stability during continual learning. Furthermore, PathWeave performs comparably to state-of-the-art MLLMs while concurrently reducing parameter training burdens by 98. 73\%. Our code locates at \url{https: //github. com/JiazuoYu/PathWeave}.

AAAI Conference 2024 Conference Paper

Long-Term Fair Decision Making through Deep Generative Models

  • Yaowei Hu
  • Yongkai Wu
  • Lu Zhang

This paper studies long-term fair machine learning which aims to mitigate group disparity over the long term in sequential decision-making systems. To define long-term fairness, we leverage the temporal causal graph and use the 1-Wasserstein distance between the interventional distributions of different demographic groups at a sufficiently large time step as the quantitative metric. Then, we propose a three-phase learning framework where the decision model is trained on high-fidelity data generated by a deep generative model. We formulate the optimization problem as a performative risk minimization and adopt the repeated gradient descent algorithm for learning. The empirical evaluation shows the efficacy of the proposed method using both synthetic and semi-synthetic datasets.

YNIMG Journal 2024 Journal Article

Meso-scale reorganization of local–global brain networks under mild sedation of propofol anesthesia

  • Kangli Dong
  • Lu Zhang
  • Yuming Zhong
  • Tao Xu
  • Yue Zhao
  • Siya Chen
  • Seedahmed S. Mahmoud
  • Qiang Fang

The fragmentation of the functional brain network has been identified through the functional connectivity (FC) analysis in studies investigating anesthesia-induced loss of consciousness (LOC). However, it remains unclear whether mild sedation of anesthesia can cause similar effects. This paper aims to explore the changes in local-global brain network topology during mild anesthesia, to better understand the macroscopic neural mechanism underlying anesthesia sedation. We analyzed high-density EEG from 20 participants undergoing mild and moderate sedation of propofol anesthesia. By employing a local-global brain parcellation in EEG source analysis, we established binary functional brain networks for each participant. Furthermore, we investigated the global-scale properties of brain networks by estimating global efficiency and modularity, and examined the changes in meso-scale properties of brain networks by quantifying the distribution of high-degree and high-betweenness hubs and their corresponding rich-club coefficients. It is evident from the results that the mild sedation of anesthesia does not cause a significant change in the global-scale properties of brain networks. However, network components centered on SomMot L show a significant decrease, while those centered on Default L, Vis L and Limbic L exhibit a significant increase during the transition from wakefulness to mild sedation (p<0.05). Compared to the baseline state, mild sedation almost doubled the number of high-degree hubs in Vis L, DorsAttn L, Limbic L, Cont L, and reduced by half the number of high-degree hubs in SomMot R, DorsAttn R, SalVentAttn R. Further, mild sedation almost doubled the number of high-betweenness hubs in Vis L, Vis R, Limbic R, Cont R, and reduced by half the number of high-betweenness hubs in SomMot L, SalVentAttn L, Default L, and SomMot R. Our results indicate that mild anesthesia cannot affect the global integration and segregation of brain networks, but influence meso-scale function for integrating different resting-state systems involved in various segregation processes. Our findings suggest that the meso-scale brain network reorganization, situated between global integration and local segregation, could reflect the autonomic compensation of the brain for drug effects. As a direct response and adjustment of the brain network system to drug administration, this spontaneous reorganization of the brain network aims at maintaining consciousness in the case of sedation.

ICRA Conference 2024 Conference Paper

PathRL: An End-to-End Path Generation Method for Collision Avoidance via Deep Reinforcement Learning

  • Wenhao Yu 0010
  • Jie Peng 0002
  • Quecheng Qiu
  • Hanyu Wang
  • Lu Zhang
  • Jianmin Ji

Robot navigation using deep reinforcement learning (DRL) has shown great potential in improving the performance of mobile robots. Nevertheless, most existing DRL-based navigation methods primarily focus on training a policy that directly commands the robot with low-level controls, like linear and angular velocities, which leads to unstable speeds and unsmooth trajectories of the robot during the long-term execution. An alternative method is to train a DRL policy that outputs the navigation path directly. Then the robot can follow the generated path smoothly using sophisticated velocity-planning and path-following controllers, whose parameters are specified according to the hardware platform. However, two roadblocks arise for training a DRL policy that outputs paths: (1) The action space for potential paths often involves higher dimensions comparing to low-level commands, which increases the difficulties of training; (2) It takes multiple time steps to track a path instead of a single time step, which requires the path to predicate the interactions of the robot w. r. t. the dynamic environment in multiple time steps. This, in turn, amplifies the challenges associated with training. In response to these challenges, we propose PathRL, a novel DRL method that trains the policy to generate the navigation path for the robot. Specifically, we employ specific action space discretization techniques and tailored state space representation methods to address the associated challenges. Curriculum learning is employed to expedite the training process, while the reward function also takes into account the smooth transition between adjacent paths. In our experiments, PathRL achieves better success rates and reduces angular rotation variability compared to other DRL navigation methods, facilitating stable and smooth robot movement. We demonstrate the competitive edge of PathRL in both real-world scenarios and multiple challenging simulation environments.

AAAI Conference 2024 Conference Paper

Weakly Supervised Few-Shot Object Detection with DETR

  • Chenbo Zhang
  • Yinglu Zhang
  • Lu Zhang
  • Jiajia Zhao
  • Jihong Guan
  • Shuigeng Zhou

In recent years, Few-shot Object Detection (FSOD) has become an increasingly important research topic in computer vision. However, existing FSOD methods require strong annotations including category labels and bounding boxes, and their performance is heavily dependent on the quality of box annotations. However, acquiring strong annotations is both expensive and time-consuming. This inspires the study on weakly supervised FSOD (WS-FSOD in short), which realizes FSOD with only image-level annotations, i.e., category labels. In this paper, we propose a new and effective weakly supervised FSOD method named WFS-DETR. By a well-designed pretraining process, WFS-DETR first acquires general object localization and integrity judgment capabilities on large-scale pretraining data. Then, it introduces object integrity into multiple-instance learning to solve the common local optimum problem by comprehensively exploiting both semantic and visual information. Finally, with simple fine-tuning, it transfers the knowledge learned from the base classes to the novel classes, which enables accurate detection of novel objects. Benefiting from this ``pretraining-refinement'' mechanism, WSF-DETR can achieve good generalization on different datasets. Extensive experiments also show that the proposed method clearly outperforms the existing counterparts in the WS-FSOD task.

YNIMG Journal 2023 Journal Article

Cascaded Multi-Modal Mixing Transformers for Alzheimer’s Disease Classification with Incomplete Data

  • Linfeng Liu
  • Siyu Liu
  • Lu Zhang
  • Xuan Vinh To
  • Fatima Nasrallah
  • Shekhar S. Chandra

Accurate medical classification requires a large number of multi-modal data, and in many cases, different feature types. Previous studies have shown promising results when using multi-modal data, outperforming single-modality models when classifying diseases such as Alzheimer's Disease (AD). However, those models are usually not flexible enough to handle missing modalities. Currently, the most common workaround is discarding samples with missing modalities which leads to considerable data under-utilisation. Adding to the fact that labelled medical images are already scarce, the performance of data-driven methods like deep learning can be severely hampered. Therefore, a multi-modal method that can handle missing data in various clinical settings is highly desirable. In this paper, we present Multi-Modal Mixing Transformer (3MT), a disease classification transformer that not only leverages multi-modal data but also handles missing data scenarios. In this work, we test 3MT for AD and Cognitively normal (CN) classification and mild cognitive impairment (MCI) conversion prediction to progressive MCI (pMCI) or stable MCI (sMCI) using clinical and neuroimaging data. The model uses a novel Cascaded Modality Transformers architecture with cross-attention to incorporate multi-modal information for more informed predictions. We propose a novel modality dropout mechanism to ensure an unprecedented level of modality independence and robustness to handle missing data scenarios. The result is a versatile network that enables the mixing of arbitrary numbers of modalities with different feature types and also ensures full data utilization in missing data scenarios. The model is trained and evaluated on the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset with the state-of-the-art performance and further evaluated with The Australian Imaging Biomarker & Lifestyle Flagship Study of Ageing (AIBL) dataset with missing data.

AAAI Conference 2023 Conference Paper

DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery – a Focus on Affinity Prediction Problems with Noise Annotations

  • Yuanfeng Ji
  • Lu Zhang
  • Jiaxiang Wu
  • Bingzhe Wu
  • Lanqing Li
  • Long-Kai Huang
  • Tingyang Xu
  • Yu Rong

AI-aided drug discovery (AIDD) is gaining popularity due to its potential to make the search for new pharmaceuticals faster, less expensive, and more effective. Despite its extensive use in numerous fields (e.g., ADMET prediction, virtual screening), little research has been conducted on the out-of-distribution (OOD) learning problem with noise. We present DrugOOD, a systematic OOD dataset curator and benchmark for AIDD. Particularly, we focus on the drug-target binding affinity prediction problem, which involves both macromolecule (protein target) and small-molecule (drug compound). DrugOOD offers an automated dataset curator with user-friendly customization scripts, rich domain annotations aligned with biochemistry knowledge, realistic noise level annotations, and rigorous benchmarking of SOTA OOD algorithms, as opposed to only providing fixed datasets. Since the molecular data is often modeled as irregular graphs using graph neural network (GNN) backbones, DrugOOD also serves as a valuable testbed for graph OOD learning problems. Extensive empirical studies have revealed a significant performance gap between in-distribution and out-of-distribution experiments, emphasizing the need for the development of more effective schemes that permit OOD generalization under noise for AIDD.

TIST Journal 2023 Journal Article

Recent Few-shot Object Detection Algorithms: A Survey with Performance Comparison

  • Tianying Liu
  • Lu Zhang
  • Yang Wang
  • Jihong Guan
  • Yanwei Fu
  • Jiajia Zhao
  • Shuigeng Zhou

The generic object detection (GOD) task has been successfully tackled by recent deep neural networks, trained by an avalanche of annotated training samples from some common classes. However, it is still non-trivial to generalize these object detectors to the novel long-tailed object classes, which have only few labeled training samples. To this end, the Few-Shot Object Detection (FSOD) has been topical recently, as it mimics the humans’ ability of learning to learn and intelligently transfers the learned generic object knowledge from the common heavy-tailed to the novel long-tailed object classes. Especially, the research in this emerging field has been flourishing in recent years with various benchmarks, backbones, and methodologies proposed. To review these FSOD works, there are several insightful FSOD survey articles [ 58, 59, 74, 78 ] that systematically study and compare them as the groups of fine-tuning/transfer learning and meta-learning methods. In contrast, we review the existing FSOD algorithms from a new perspective under a new taxonomy based on their contributions, i.e., data-oriented, model-oriented, and algorithm-oriented. Thus, a comprehensive survey with performance comparison is conducted on recent achievements of FSOD. Furthermore, we also analyze the technical challenges, the merits and demerits of these methods, and envision the future directions of FSOD. Specifically, we give an overview of FSOD, including the problem definition, common datasets, and evaluation protocols. The taxonomy is then proposed that groups FSOD methods into three types. Following this taxonomy, we provide a systematic review of the advances in FSOD. Finally, further discussions on performance, challenges, and future directions are presented.

IJCAI Conference 2023 Conference Paper

Video Diffusion Models with Local-Global Context Guidance

  • Siyuan Yang
  • Lu Zhang
  • Yu Liu
  • Zhizhuo Jiang
  • You He

Diffusion models have emerged as a powerful paradigm in video synthesis tasks including prediction, generation, and interpolation. Due to the limitation of the computational budget, existing methods usually implement conditional diffusion models with an autoregressive inference pipeline, in which the future fragment is predicted based on the distribution of adjacent past frames. However, only the conditions from a few previous frames can't capture the global temporal coherence, leading to inconsistent or even outrageous results in long-term video prediction. In this paper, we propose a Local-Global Context guided Video Diffusion model (LGC-VD) to capture multi-perception conditions for producing high-quality videos in both conditional/unconditional settings. In LGC-VD, the UNet is implemented with stacked residual blocks with self-attention units, avoiding the undesirable computational cost in 3D Conv. We construct a local-global context guidance strategy to capture the multi-perceptual embedding of the past fragment to boost the consistency of future prediction. Furthermore, we propose a two-stage training strategy to alleviate the effect of noisy frames for more stable predictions. Our experiments demonstrate that the proposed method achieves favorable performance on video prediction, interpolation, and unconditional video generation. We release code at https: //github. com/exisas/LGC-VD.

AAAI Conference 2022 Conference Paper

Achieving Counterfactual Fairness for Causal Bandit

  • Wen Huang
  • Lu Zhang
  • Xintao Wu

In online recommendation, customers arrive in a sequential and stochastic manner from an underlying distribution and the online decision model recommends a chosen item for each arriving individual based on some strategy. We study how to recommend an item at each step to maximize the expected reward while achieving user-side fairness for customers, i. e. , customers who share similar profiles will receive a similar reward regardless of their sensitive attributes and items being recommended. By incorporating causal inference into bandits and adopting soft intervention to model the arm selection strategy, we first propose the d-separation based UCB algorithm (D-UCB) to explore the utilization of the d-separation set in reducing the amount of exploration needed to achieve low cumulative regret. Based on that, we then propose the fair causal bandit (F-UCB) for achieving the counterfactual individual fairness. Both theoretical analysis and empirical evaluation demonstrate effectiveness of our algorithms.

AAAI Conference 2022 Conference Paper

Achieving Long-Term Fairness in Sequential Decision Making

  • Yaowei Hu
  • Lu Zhang

In this paper, we propose a framework for achieving longterm fair sequential decision making. By conducting both the hard and soft interventions, we propose to take path-specific effects on the time-lagged causal graph as a quantitative tool for measuring long-term fairness. The problem of fair sequential decision making is then formulated as a constrained optimization problem with the utility as the objective and the long-term and short-term fairness as constraints. We show that such an optimization problem can be converted to a performative risk optimization. Finally, repeated risk minimization (RRM) is used for model training, and the convergence of RRM is theoretically analyzed. The empirical evaluation shows the effectiveness of the proposed algorithm on synthetic and semi-synthetic temporal datasets.

AAAI Conference 2022 Conference Paper

Generalized Equivariance and Preferential Labeling for GNN Node Classification

  • Zeyu Sun
  • Wenjie Zhang
  • Lili Mou
  • Qihao Zhu
  • Yingfei Xiong
  • Lu Zhang

Existing graph neural networks (GNNs) largely rely on node embeddings, which represent a node as a vector by its identity, type, or content. However, graphs with unattributed nodes widely exist in real-world applications (e. g. , anonymized social networks). Previous GNNs either assign random labels to nodes (which introduces artefacts to the GNN) or assign one embedding to all nodes (which fails to explicitly distinguish one node from another). Further, when these GNNs are applied to unattributed node classification problems, they have an undesired equivariance property, which are fundamentally unable to address the data with multiple possible outputs. In this paper, we analyze the limitation of existing approaches to node classification problems. Inspired by our analysis, we propose a generalized equivariance property and a Preferential Labeling technique that satisfies the desired property asymptotically. Experimental results show that we achieve high performance in several unattributed node classification tasks.

IJCAI Conference 2022 Conference Paper

Grape: Grammar-Preserving Rule Embedding

  • Qihao Zhu
  • Zeyu Sun
  • Wenjie Zhang
  • Yingfei Xiong
  • Lu Zhang

Word embedding has been widely used in various areas to boost the performance of the neural models. However, when processing context-free languages, embedding grammar rules with word embedding loses two types of information. One is the structural relationship between the grammar rules, and the other one is the content information of the rule definition. In this paper, we make the first attempt to learn a grammar-preserving rule embedding. We first introduce a novel graph structure to represent the context-free grammar. Then, we apply a Graph Neural Network (GNN) to extract the structural information and use a gating layer to integrate content information. We conducted experiments on six widely-used benchmarks containing four context-free languages. The results show that our approach improves the accuracy of the base model by 0. 8 to 6. 4 percentage points. Furthermore, Grape also achieves 1. 6 F1 score improvement on the method naming task which shows the generality of our approach.

IJCAI Conference 2022 Conference Paper

Lyra: A Benchmark for Turducken-Style Code Generation

  • Qingyuan Liang
  • Zeyu Sun
  • Qihao Zhu
  • Wenjie Zhang
  • Lian Yu
  • Yingfei Xiong
  • Lu Zhang

Recently, neural techniques have been used to generate source code automatically. While promising for declarative languages, these approaches achieve much poorer performance on datasets for imperative languages. Since a declarative language is typically embedded in an imperative language (i. e. , the turducken-style programming) in real-world software development, the promising results on declarative languages can hardly lead to significant reduction of manual software development efforts. In this paper, we define a new code generation task: given a natural language comment, this task aims to generate a program in a base imperative language with an embedded declarative language. To our knowledge, this is the first turducken-style code generation task. For this task, we present Lyra: a dataset in Python with embedded SQL. This dataset contains 2, 000 carefully annotated database manipulation programs from real usage projects. Each program is paired with both a Chinese comment and an English comment. In our experiment, we adopted Transformer, BERT-style, and GPT-style models as baselines. In the best setting, GPT-style model can achieve 24% and 25. 5% AST exact matching accuracy using Chinese and English comments, respectively. Therefore, we believe that Lyra provides a new challenge for code generation. Yet, overcoming this challenge may significantly boost the applicability of code generation techniques for real-world software development.

IJCAI Conference 2022 Conference Paper

Next Point-of-Interest Recommendation with Inferring Multi-step Future Preferences

  • Lu Zhang
  • Zhu Sun
  • Ziqing Wu
  • Jie Zhang
  • Yew Soon Ong
  • Xinghua Qu

Existing studies on next point-of-interest (POI) recommendation mainly attempt to learn user preference from the past and current sequential behaviors. They, however, completely ignore the impact of future behaviors on the decision-making, thus hindering the quality of user preference learning. Intuitively, users' next POI visits may also be affected by their multi-step future behaviors, as users may often have activity planning in mind. To fill this gap, we propose a novel Context-aware Future Preference inference Recommender (CFPRec) to help infer user future preference in a self-ensembling manner. In particular, it delicately derives multi-step future preferences from the learned past preference thanks to the periodic property of users' daily check-ins, so as to implicitly mimic user’s activity planning before her next visit. The inferred future preferences are then seamlessly integrated with the current preference for more expressive user preference learning. Extensive experiments on three datasets demonstrate the superiority of CFPRec against state-of-the-arts.

JMLR Journal 2022 Journal Article

Pathfinder: Parallel quasi-Newton variational inference

  • Lu Zhang
  • Bob Carpenter
  • Andrew Gelman
  • Aki Vehtari

We propose Pathfinder, a variational method for approximately sampling from differentiable probability densities. Starting from a random initialization, Pathfinder locates normal approximations to the target density along a quasi-Newton optimization path, with local covariance estimated using the inverse Hessian estimates produced by the optimizer. Pathfinder returns draws from the approximation with the lowest estimated Kullback-Leibler (KL) divergence to the target distribution. We evaluate Pathfinder on a wide range of posterior distributions, demonstrating that its approximate draws are better than those from automatic differentiation variational inference (ADVI) and comparable to those produced by short chains of dynamic Hamiltonian Monte Carlo (HMC), as measured by 1-Wasserstein distance. Compared to ADVI and short dynamic HMC runs, Pathfinder requires one to two orders of magnitude fewer log density and gradient evaluations, with greater reductions for more challenging posteriors. Importance resampling over multiple runs of Pathfinder improves the diversity of approximate draws, reducing 1-Wasserstein distance further and providing a measure of robustness to optimization failures on plateaus, saddle points, or in minor modes. The Monte Carlo KL divergence estimates are embarrassingly parallelizable in the core Pathfinder algorithm, as are multiple runs in the resampling version, further increasing Pathfinder's speed advantage with multiple cores. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2022. ( edit, beta )

IROS Conference 2022 Conference Paper

Towards edible drones for rescue missions: design and flight of nutritional wings

  • Bokeon Kwak
  • Jun Shintake
  • Lu Zhang
  • Dario Floreano

Drones have shown to be useful aerial vehicles for unmanned transport missions such as food and medical supply delivery. This can be leveraged to deliver life-saving nutrition and medicine for people in emergency situations. However, commercial drones can generally only carry 10 %–30 % of their own mass as payload, which limits the amount of food delivery in a single flight. One novel solution to noticeably increase the food-carrying ratio of a drone, is recreating some structures of a drone, such as the wings, with edible materials. We thus propose a drone, which is no longer only a food-transporting aircraft, but itself is partially edible, increasing its food-carrying mass ratio to 50 %, owing to its edible wings. Furthermore, should the edible drone be left behind in the environment after performing its task in an emergency situation, it will be more biodegradable than its non-edible counterpart, leaving less waste in the environment. Here we describe the choice of materials and scalable design of edible wings, and validate the method in a flight-capable prototype that can provide 300 kcal and carry a payload of 80 g of water.

AAAI Conference 2022 Conference Paper

You Only Infer Once: Cross-Modal Meta-Transfer for Referring Video Object Segmentation

  • Dezhuang Li
  • Ruoqi Li
  • Lijun Wang
  • Yifan Wang
  • Jinqing Qi
  • Lu Zhang
  • Ting Liu
  • Qingquan Xu

We present YOFO (You Only inFer Once), a new paradigm for referring video object segmentation (RVOS) that operates in an one-stage manner. Our key insight is that the language descriptor should serve as target-specific guidance to identify the target object, while a direct feature fusion of image and language can increase feature complexity and thus may be sub-optimal for RVOS. To this end, we propose a metatransfer module, which is trained in a learning-to-learn fashion and aims to transfer the target-specific information from the language domain to the image domain, while discarding the uncorrelated complex variations of language description. To bridge the gap between the image and language domains, we develop a multi-scale cross-modal feature mining block that aggregates all the essential features required by RVOS from both domains and generates regression labels for the meta-transfer module. The whole system can be trained in an end-to-end manner and shows competitive performance against state-of-the-art two-stage approaches.

AAAI Conference 2021 Conference Paper

A Generative Adversarial Framework for Bounding Confounded Causal Effects

  • Yaowei Hu
  • Yongkai Wu
  • Lu Zhang
  • Xintao Wu

Causal inference from observational data is receiving wide applications in many fields. However, unidentifiable situations, where causal effects cannot be uniquely computed from observational data, pose critical barriers to applying causal inference to complicated real applications. In this paper, we develop a bounding method for estimating the average causal effect (ACE) under unidentifiable situations due to hidden confounding based on Pearl’s structural causal model. We propose to parameterize the unknown exogenous random variables and structural equations of a causal model using neural networks and implicit generative models. Then, using an adversarial learning framework, we search the parameter space to explicitly traverse causal models that agree with the given observational distribution, and find those that minimize or maximize the ACE to obtain its lower and upper bounds. The proposed method does not make assumption about the type of structural equations and variables. Experiments using both synthetic and real-world datasets are conducted.

IJCAI Conference 2021 Conference Paper

Conditional Self-Supervised Learning for Few-Shot Classification

  • Yuexuan An
  • Hui Xue
  • Xingyu Zhao
  • Lu Zhang

How to learn a transferable feature representation from limited examples is a key challenge for few-shot classification. Self-supervision as an auxiliary task to the main supervised few-shot task is considered to be a conceivable way to solve the problem since self-supervision can provide additional structural information easily ignored by the main task. However, learning a good representation by traditional self-supervised methods is usually dependent on large training samples. In few-shot scenarios, due to the lack of sufficient samples, these self-supervised methods might learn a biased representation, which more likely leads to the wrong guidance for the main tasks and finally causes the performance degradation. In this paper, we propose conditional self-supervised learning (CSS) to use auxiliary information to guide the representation learning of self-supervised tasks. Specifically, CSS leverages supervised information as prior knowledge to shape and improve the learning feature manifold of self-supervision without auxiliary unlabeled data, so as to reduce representation bias and mine more effective semantic information. Moreover, CSS exploits more meaningful information through supervised and the improved self-supervised learning respectively and integrates the information into a unified distribution, which can further enrich and broaden the original representation. Extensive experiments demonstrate that our proposed method without any fine-tuning can achieve a significant accuracy improvement on the few-shot classification scenarios compared to the state-of-the-art few-shot learning methods.

NeurIPS Conference 2021 Conference Paper

DP-SSL: Towards Robust Semi-supervised Learning with A Few Labeled Samples

  • Yi Xu
  • Jiandong Ding
  • Lu Zhang
  • Shuigeng Zhou

The scarcity of labeled data is a critical obstacle to deep learning. Semi-supervised learning (SSL) provides a promising way to leverage unlabeled data by pseudo labels. However, when the size of labeled data is very small (say a few labeled samples per class), SSL performs poorly and unstably, possibly due to the low quality of learned pseudo labels. In this paper, we propose a new SSL method called DP-SSL that adopts an innovative data programming (DP) scheme to generate probabilistic labels for unlabeled data. Different from existing DP methods that rely on human experts to provide initial labeling functions (LFs), we develop a multiple-choice learning~(MCL) based approach to automatically generate LFs from scratch in SSL style. With the noisy labels produced by the LFs, we design a label model to resolve the conflict and overlap among the noisy labels, and finally infer probabilistic labels for unlabeled samples. Extensive experiments on four standard SSL benchmarks show that DP-SSL can provide reliable labels for unlabeled data and achieve better classification performance on test sets than existing SSL methods, especially when only a small number of labeled samples are available. Concretely, for CIFAR-10 with only 40 labeled samples, DP-SSL achieves 93. 82% annotation accuracy on unlabeled data and 93. 46% classification accuracy on test data, which are higher than the SOTA results.

YNIMG Journal 2021 Journal Article

Sensory, somatomotor and internal mentation networks emerge dynamically in the resting brain with internal mentation predominating in older age

  • Lu Zhang
  • Jiajia Zhao
  • Qunjie Zhou
  • ZhaoWen Liu
  • Yi Zhang
  • Wei Cheng
  • Weikang Gong
  • Xiaoping Hu

Age-related changes in the brain are associated with a decline in functional flexibility. Intrinsic functional flexibility is evident in the brain's dynamic ability to switch between alternative spatiotemporal states during resting state. However, the relationship between brain connectivity states, associated psychological functions during resting state, and the changes in normal aging remain poorly understood. In this study, we analyzed resting-state functional magnetic resonance imaging (rsfMRI) data from the Human Connectome Project (HCP; N = 812) and the UK Biobank (UKB; N = 6,716). Using signed community clustering to identify distinct states of dynamic functional connectivity, and text-mining of a large existing literature for functional annotation of each state, our findings from the HCP dataset indicated that the resting brain spontaneously transitions between three functionally specialized states: sensory, somatomotor, and internal mentation networks. The occurrence, transition-rate, and persistence-time parameters for each state were correlated with behavioural scores using canonical correlation analysis. We estimated the same brain states and parameters in the UKB dataset, subdivided into three distinct age ranges: 50-55, 56-67, and 68-78 years. We found that the internal mentation network was more frequently expressed in people aged 71 and older, whereas people younger than 55 more frequently expressed sensory and somatomotor networks. Furthermore, analysis of the functional entropy - a measure of uncertainty of functional connectivity - also supported this finding across the three age ranges. Our study demonstrates that dynamic functional connectivity analysis can expose the time-varying patterns of transition between functionally specialized brain states, which are strongly tied to increasing age.

IJCAI Conference 2020 Conference Paper

An Interactive Multi-Task Learning Framework for Next POI Recommendation with Uncertain Check-ins

  • Lu Zhang
  • Zhu Sun
  • Jie Zhang
  • Yu Lei
  • Chen Li
  • Ziqing Wu
  • Horst Kloeden
  • Felix Klanner

Studies on next point-of-interest (POI) recommendation mainly seek to learn users' transition patterns with certain historical check-ins. However, in reality, users' movements are typically uncertain (i. e. , fuzzy and incomplete) where most existing methods suffer from the transition pattern vanishing issue. To ease this issue, we propose a novel interactive multi-task learning (iMTL) framework to better exploit the interplay between activity and location preference. Specifically, iMTL introduces: (1) temporal-aware activity encoder equipped with fuzzy characterization over uncertain check-ins to unveil the latent activity transition patterns; (2) spatial-aware location preference encoder to capture the latent location transition patterns; and (3) task-specific decoder to make use of the learned latent transition patterns and enhance both activity and location prediction tasks in an interactive manner. Extensive experiments on three real-world datasets show the superiority of iMTL.

NeurIPS Conference 2020 Conference Paper

Fair Multiple Decision Making Through Soft Interventions

  • Yaowei Hu
  • Yongkai Wu
  • Lu Zhang
  • Xintao Wu

Previous research in fair classification mostly focuses on a single decision model. In reality, there usually exist multiple decision models within a system and all of which may contain a certain amount of discrimination. Such realistic scenarios introduce new challenges to fair classification: since discrimination may be transmitted from upstream models to downstream models, building decision models separately without taking upstream models into consideration cannot guarantee to achieve fairness. In this paper, we propose an approach that learns multiple classifiers and achieves fairness for all of them simultaneously, by treating each decision model as a soft intervention and inferring the post-intervention distributions to formulate the loss function as well as the fairness constraints. We adopt surrogate functions to smooth the loss function and constraints, and theoretically show that the excess risk of the proposed loss function can be bounded in a form that is the same as that for traditional surrogated loss functions. Experiments using both synthetic and real-world datasets show the effectiveness of our approach.

IJCAI Conference 2020 Conference Paper

NLocalSAT: Boosting Local Search with Solution Prediction

  • Wenjie Zhang
  • Zeyu Sun
  • Qihao Zhu
  • Ge Li
  • Shaowei Cai
  • Yingfei Xiong
  • Lu Zhang

The Boolean satisfiability problem (SAT) is a famous NP-complete problem in computer science. An effective way for solving a satisfiable SAT problem is the stochastic local search (SLS). However, in this method, the initialization is assigned in a random manner, which impacts the effectiveness of SLS solvers. To address this problem, we propose NLocalSAT. NLocalSAT combines SLS with a solution prediction model, which boosts SLS by changing initialization assignments with a neural network. We evaluated NLocalSAT on five SLS solvers (CCAnr, Sparrow, CPSparrow, YalSAT, and probSAT) with instances in the random track of SAT Competition 2018. The experimental results show that solvers with NLocalSAT achieve 27% ~ 62% improvement over the original SLS solvers.

AAAI Conference 2020 Conference Paper

TreeGen: A Tree-Based Transformer Architecture for Code Generation

  • Zeyu Sun
  • Qihao Zhu
  • Yingfei Xiong
  • Yican Sun
  • Lili Mou
  • Lu Zhang

A code generation system generates programming language code based on an input natural language description. State-ofthe-art approaches rely on neural networks for code generation. However, these code generators suffer from two problems. One is the long dependency problem, where a code element often depends on another far-away code element. A variable reference, for example, depends on its definition, which may appear quite a few lines before. The other problem is structure modeling, as programs contain rich structural information. In this paper, we propose a novel tree-based neural architecture, TreeGen, for code generation. TreeGen uses the attention mechanism of Transformers to alleviate the longdependency problem, and introduces a novel AST reader (encoder) to incorporate grammar rules and AST structures into the network. We evaluated TreeGen on a Python benchmark, HearthStone, and two semantic parsing benchmarks, ATIS and GEO. TreeGen outperformed the previous state-of-theart approach by 4. 5 percentage points on HearthStone, and achieved the best accuracy among neural network-based approaches on ATIS (89. 1%) and GEO (89. 6%). We also conducted an ablation test to better understand each component of our model.

AAAI Conference 2019 Conference Paper

A Grammar-Based Structural CNN Decoder for Code Generation

  • Zeyu Sun
  • Qihao Zhu
  • Lili Mou
  • Yingfei Xiong
  • Ge Li
  • Lu Zhang

Code generation maps a program description to executable source code in a programming language. Existing approaches mainly rely on a recurrent neural network (RNN) as the decoder. However, we find that a program contains significantly more tokens than a natural language sentence, and thus it may be inappropriate for RNN to capture such a long sequence. In this paper, we propose a grammar-based structural convolutional neural network (CNN) for code generation. Our model generates a program by predicting the grammar rules of the programming language; we design several CNN modules, including the tree-based convolution and pre-order convolution, whose information is further aggregated by dedicated attentive pooling layers. Experimental results on the HearthStone benchmark dataset show that our CNN code generator significantly outperforms the previous state-of-the-art method by 5 percentage points; additional experiments on several semantic parsing tasks demonstrate the robustness of our model. We also conduct in-depth ablation test to better understand each component of our model.

IJCAI Conference 2019 Conference Paper

Achieving Causal Fairness through Generative Adversarial Networks

  • Depeng Xu
  • Yongkai Wu
  • Shuhan Yuan
  • Lu Zhang
  • Xintao Wu

Achieving fairness in learning models is currently an imperative task in machine learning. Meanwhile, recent research showed that fairness should be studied from the causal perspective, and proposed a number of fairness criteria based on Pearl's causal modeling framework. In this paper, we investigate the problem of building causal fairness-aware generative adversarial networks (CFGAN), which can learn a close distribution from a given dataset, while also ensuring various causal fairness criteria based on a given causal graph. CFGAN adopts two generators, whose structures are purposefully designed to reflect the structures of causal graph and interventional graph. Therefore, the two generators can respectively simulate the underlying causal model that generates the real data, as well as the causal model after the intervention. On the other hand, two discriminators are used for producing a close-to-real distribution, as well as for achieving various fairness criteria based on causal quantities simulated by generators. Experiments on a real-world dataset show that CFGAN can generate high quality fair data.

IJCAI Conference 2019 Conference Paper

Counterfactual Fairness: Unidentification, Bound and Algorithm

  • Yongkai Wu
  • Lu Zhang
  • Xintao Wu

Fairness-aware learning studies the problem of building machine learning models that are subject to fairness requirements. Counterfactual fairness is a notion of fairness derived from Pearl's causal model, which considers a model is fair if for a particular individual or group its prediction in the real world is the same as that in the counterfactual world where the individual(s) had belonged to a different demographic group. However, an inherent limitation of counterfactual fairness is that it cannot be uniquely quantified from the observational data in certain situations, due to the unidentifiability of the counterfactual quantity. In this paper, we address this limitation by mathematically bounding the unidentifiable counterfactual quantity, and develop a theoretically sound algorithm for constructing counterfactually fair classifiers. We evaluate our method in the experiments using both synthetic and real-world datasets, as well as compare with existing methods. The results validate our theory and show the effectiveness of our method.

NeurIPS Conference 2019 Conference Paper

PC-Fairness: A Unified Framework for Measuring Causality-based Fairness

  • Yongkai Wu
  • Lu Zhang
  • Xintao Wu
  • Hanghang Tong

A recent trend of fair machine learning is to define fairness as causality-based notions which concern the causal connection between protected attributes and decisions. However, one common challenge of all causality-based fairness notions is identifiability, i. e. , whether they can be uniquely measured from observational data, which is a critical barrier to applying these notions to real-world situations. In this paper, we develop a framework for measuring different causality-based fairness. We propose a unified definition that covers most of previous causality-based fairness notions, namely the path-specific counterfactual fairness (PC fairness). Based on that, we propose a general method in the form of a constrained optimization problem for bounding the path-specific counterfactual fairness under all unidentifiable situations. Experiments on synthetic and real-world datasets show the correctness and effectiveness of our method.

IJCAI Conference 2018 Conference Paper

Achieving Non-Discrimination in Prediction

  • Lu Zhang
  • Yongkai Wu
  • Xintao Wu

In discrimination-aware classification, the pre-process methods for constructing a discrimination-free classifier first remove discrimination from the training data, and then learn the classifier from the cleaned data. However, they lack a theoretical guarantee for the potential discrimination when the classifier is deployed for prediction. In this paper, we fill this gap by mathematically bounding the discrimination in prediction. We adopt the causal model for modeling the data generation mechanism, and formally defining discrimination in population, in a dataset, and in prediction. We obtain two important theoretical results: (1) the discrimination in prediction can still exist even if the discrimination in the training data is completely removed; and (2) not all pre-process methods can ensure non-discrimination in prediction even though they can achieve non-discrimination in the modified training data. Based on the results, we develop a two-phase framework for constructing a discrimination-free classifier with a theoretical guarantee. The experiments demonstrate the theoretical results and show the effectiveness of our two-phase framework.

IJCAI Conference 2017 Conference Paper

A Causal Framework for Discovering and Removing Direct and Indirect Discrimination

  • Lu Zhang
  • Yongkai Wu
  • Xintao Wu

In this paper, we investigate the problem of discovering both direct and indirect discrimination from the historical data, and removing the discriminatory effects before the data is used for predictive analysis (e. g. , building classifiers). The main drawback of existing methods is that they cannot distinguish the part of influence that is really caused by discrimination from all correlated influences. In our approach, we make use of the causal network to capture the causal structure of the data. Then we model direct and indirect discrimination as the path-specific effects, which accurately identify the two types of discrimination as the causal effects transmitted along different paths in the network. Based on that, we propose an effective algorithm for discovering direct and indirect discrimination, as well as an algorithm for precisely removing both types of discrimination while retaining good data utility. Experiments using the real dataset show the effectiveness of our approaches.

AAAI Conference 2016 Conference Paper

Convolutional Neural Networks over Tree Structures for Programming Language Processing

  • Lili Mou
  • Ge Li
  • Lu Zhang
  • Tao Wang
  • Zhi Jin

Programming language processing (similar to natural language processing) is a hot research topic in the field of software engineering; it has also aroused growing interest in the artificial intelligence community. However, different from a natural language sentence, a program contains rich, explicit, and complicated structural information. Hence, traditional NLP models may be inappropriate for programs. In this paper, we propose a novel tree-based convolutional neural network (TBCNN) for programming language processing, in which a convolution kernel is designed over programs’ abstract syntax trees to capture structural information. TBCNN is a generic architecture for programming language processing; our experiments show its effectiveness in two different program analysis tasks: classifying programs according to functionality, and detecting code snippets of certain patterns. TBCNN outperforms baseline methods, including several neural models for NLP.

IJCAI Conference 2016 Conference Paper

Situation Testing-Based Discrimination Discovery: A Causal Inference Approach

  • Lu Zhang
  • Yongkai Wu
  • Xintao Wu

Discrimination discovery is to unveil discrimination against a specific individual by analyzing the historical dataset. In this paper, we develop a general technique to capture discrimination based on the legally grounded situation testing methodology. For any individual, we find pairs of tuples from the dataset with similar characteristics apart from belonging or not to the protected-by-law group and assign them in two groups. The individual is considered as discriminated if significant difference is observed between the decisions from the two groups. To find similar tuples, we make use of the Causal Bayesian Networks and the associated causal inference as a guideline. The causal structure of the dataset and the causal effect of each attribute on the decision are used to facilitate the similarity measurement. Through empirical assessments on a real dataset, our approach shows good efficacy both in accuracy and efficiency.

v2026.09.13