Arrow Research search

Author name cluster

Yan Zeng

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

13 papers
2 author rows

Possible papers

13

AAAI Conference 2026 Conference Paper

Class-Aware Active Annotation in Federated Semi-Supervised Learning for Medical Image Classification

  • Meiting Xue
  • Miaoqi Li
  • Yukun Shi
  • Yan Zeng
  • Jilin Zhang
  • Jing Ma

In medical image classification, data privacy constraints and the high cost of expert annotations pose significant challenges to building generalizable models. Federated semi-supervised learning (FSSL), which combines the privacy-preserving nature of federated learning with the label efficiency of semi-supervised learning, offers a promising direction. However, in real-world deployments, client data often exhibits highly non-independent and identically distributed (Non-IID) characteristics. This distributional heterogeneity undermines the reliability of pseudo-labels generated by global models, ultimately limiting model generalization. A key limitation of existing FSSL approaches lies in their reliance on a static labeled set fixed prior to training. Such strategies lack the ability to adaptively correct pseudo-label noise or address class imbalance throughout training, particularly under Non-IID settings. To address this, we propose FSSAL, a novel framework that introduces an active learning component into the FSSL pipeline. By continuously identifying informative and representative samples during training, our method adaptively refines the labeled set and enhances the model’s robustness to distribution shifts. FSSAL employs client-private models for pseudo-label generation to reduce global bias, applies a class-aware dynamic thresholding mechanism to ensure more reliable and balanced label selection, and incorporates a sample selection strategy guided by both feature diversity and model uncertainty. Extensive experiments on four public medical image classification datasets demonstrate that FSSAL consistently outperforms competitive FSSL methods in accuracy and F1-score, especially under highly Non-IID conditions, highlighting its robustness and practical potential.

EAAI Journal 2026 Journal Article

Target-aware proposal-level fusion for multi-modal three-dimensional detection

  • Zilong Zhao
  • Baofu Wu
  • Yuyu Yin
  • Jilin Zhang
  • Youhuizi Li
  • Yan Zeng
  • Honghao Gao

Three-dimensional detection technology based on cameras and light detection and ranging is crucial for safe autonomous driving, and proposal-level fusion detection schemes have received increasing attention due to their efficiency The accuracy depends on proposal extraction, feature sampling, and multi-modal fusion. However, current methods suffer from poor proposal quality, inaccurate sampling, and limited multi-modal integration, diminishing their effectiveness. Addressing these challenges, this paper introduces the target-aware proposal-level fusion for multi-modal three-dimensional detection framework, which leverages the complementarity of point clouds and multi-view images to enhance the perception accuracy of autonomous driving systems. The framework comprises three key modules: the channel-spatial selective proposal extraction module improves proposal quality and reduces background interference through dual-dimensional filtering; the target-aware feature sampling module, which refines deformable attention mechanisms for superior adaptability to complex multi-modal data across varying environments, strengthening the robustness of feature sampling; and the adaptive feature fusion module, which dynamically integrates multi-modal features via learnable weights, optimizing cross-modal information integration. Extensive evaluations on the nutonomy scenes dataset highlight the superiority of the proposed method, with the model achieving 72. 1 mean average precision and 74. 0 nutonomy scenes detection score on the test set, affirming its advanced capabilities in multi-modal data processing. These performance enhancements translate to more accurate and dependable autonomous driving systems, capable of better navigating and responding to dynamic and challenging real-world scenarios.

NeurIPS Conference 2025 Conference Paper

Learning Counterfactual Outcomes Under Rank Preservation

  • Peng Wu
  • Haoxuan Li
  • Chunyuan Zheng
  • Yan Zeng
  • Jiawei Chen
  • Yang Liu
  • Ruocheng Guo
  • Kun Zhang

Counterfactual inference aims to estimate the counterfactual outcome at the individual level given knowledge of an observed treatment and the factual outcome, with broad applications in fields such as epidemiology, econometrics, and management science. Previous methods rely on a known structural causal model (SCM) or assume the homogeneity of the exogenous variable and strict monotonicity between the outcome and exogenous variable. In this paper, we propose a principled approach for identifying and estimating the counterfactual outcome. We first introduce a simple and intuitive rank preservation assumption to identify the counterfactual outcome without relying on a known structural causal model. Building on this, we propose a novel ideal loss for theoretically unbiased learning of the counterfactual outcome and further develop a kernel-based estimator for its empirical estimation. Our theoretical analysis shows that the rank preservation assumption is not stronger than the homogeneity and strict monotonicity assumptions, and shows that the proposed ideal loss is convex, and the proposed estimator is unbiased. Extensive semi-synthetic and real-world experiments are conducted to demonstrate the effectiveness of the proposed method.

NeurIPS Conference 2025 Conference Paper

Local Learning for Covariate Selection in Nonparametric Causal Effect Estimation with Latent Variables

  • Zheng Li
  • Xichen Guo
  • Feng Xie
  • Yan Zeng
  • Hao Zhang
  • Zhi Geng

Estimating causal effects from nonexperimental data is a fundamental problem in many fields of science. A key component of this task is selecting an appropriate set of covariates for confounding adjustment to avoid bias. Most existing methods for covariate selection often assume the absence of latent variables and rely on learning the global causal structure among variables. However, identifying the global structure can be unnecessary and inefficient, especially when our primary interest lies in estimating the effect of a treatment variable on an outcome variable. To address this limitation, we propose a novel local learning approach for covariate selection in nonparametric causal effect estimation, which accounts for the presence of latent variables. Our approach leverages testable independence and dependence relationships among observed variables to identify a valid adjustment set for a target causal relationship, ensuring both soundness and completeness under standard assumptions. We validate the effectiveness of our algorithm through extensive experiments on both synthetic and real-world data.

ICML Conference 2024 Conference Paper

Boximator: Generating Rich and Controllable Motions for Video Synthesis

  • Jiawei Wang
  • Yuchen Zhang
  • Jiaxin Zou
  • Yan Zeng
  • Guoqiang Wei
  • Liping Yuan
  • Hang Li

Generating rich and controllable motion is a pivotal challenge in video synthesis. We propose Boximator, a new approach for fine-grained motion control. Boximator introduces two constraint types: hard box and soft box. Users select objects in the conditional frame using hard boxes and then use either type of boxes to roughly or rigorously define the object’s position, shape, or motion path in future frames. Boximator functions as a plug-in for existing video diffusion models. Its training process preserves the base model’s knowledge by freezing the original weights and training only the control module. To address training challenges, we introduce a novel self-tracking technique that greatly simplifies the learning of box-object correlations. Empirically, Boximator achieves state-of-the-art video quality (FVD) scores, improving on two base models, and further enhanced after incorporating box constraints. Its robust motion controllability is validated by drastic increases in the bounding box alignment metric. Human evaluation also shows that users favor Boximator generation results over the base model.

NeurIPS Conference 2024 Conference Paper

CryoGEM: Physics-Informed Generative Cryo-Electron Microscopy

  • Jiakai Zhang
  • Qihe Chen
  • Yan Zeng
  • Wenyuan Gao
  • Xuming He
  • Zhijie Liu
  • Jingyi Yu

In the past decade, deep conditional generative models have revolutionized the generation of realistic images, extending their application from entertainment to scientific domains. Single-particle cryo-electron microscopy (cryo-EM) is crucial in resolving near-atomic resolution 3D structures of proteins, such as the SARS-COV-2 spike protein. To achieve high-resolution reconstruction, a comprehensive data processing pipeline has been adopted. However, its performance is still limited as it lacks high-quality annotated datasets for training. To address this, we introduce physics-informed generative cryo-electron microscopy (CryoGEM), which for the first time integrates physics-based cryo-EM simulation with a generative unpaired noise translation to generate physically correct synthetic cryo-EM datasets with realistic noises. Initially, CryoGEM simulates the cryo-EM imaging process based on a virtual specimen. To generate realistic noises, we leverage an unpaired noise translation via contrastive learning with a novel mask-guided sampling scheme. Extensive experiments show that CryoGEM is capable of generating authentic cryo-EM images. The generated dataset can be used as training data for particle picking and pose estimation models, eventually improving the reconstruction resolution.

NeurIPS Conference 2024 Conference Paper

DRACO: A Denoising-Reconstruction Autoencoder for Cryo-EM

  • Yingjun Shen
  • Haizhao Dai
  • Qihe Chen
  • Yan Zeng
  • Jiakai Zhang
  • Yuan Pei
  • Jingyi Yu

Foundation models in computer vision have demonstrated exceptional performance in zero-shot and few-shot tasks by extracting multi-purpose features from large-scale datasets through self-supervised pre-training methods. However, these models often overlook the severe corruption in cryogenic electron microscopy (cryo-EM) images by high-level noises. We introduce DRACO, a Denoising-Reconstruction Autoencoder for CryO-EM, inspired by the Noise2Noise (N2N) approach. By processing cryo-EM movies into odd and even images and treating them as independent noisy observations, we apply a denoising-reconstruction hybrid training scheme. We mask both images to create denoising and reconstruction tasks. For DRACO's pre-training, the quality of the dataset is essential, we hence build a high-quality, diverse dataset from an uncurated public database, including over 270, 000 movies or micrographs. After pre-training, DRACO naturally serves as a generalizable cryo-EM image denoiser and a foundation model for various cryo-EM downstream tasks. DRACO demonstrates the best performance in denoising, micrograph curation, and particle picking tasks compared to state-of-the-art baselines.

AAAI Conference 2024 Conference Paper

eTag: Class-Incremental Learning via Embedding Distillation and Task-Oriented Generation

  • Libo Huang
  • Yan Zeng
  • Chuanguang Yang
  • Zhulin An
  • Boyu Diao
  • Yongjun Xu

Class incremental learning (CIL) aims to solve the notorious forgetting problem, which refers to the fact that once the network is updated on a new task, its performance on previously-learned tasks degenerates catastrophically. Most successful CIL methods store exemplars (samples of learned tasks) to train a feature extractor incrementally, or store prototypes (features of learned tasks) to estimate the incremental feature distribution. However, the stored exemplars would violate the data privacy concerns, while the fixed prototypes might not reasonably be consistent with the incremental feature distribution, hindering the exploration of real-world CIL applications. In this paper, we propose a data-free CIL method with embedding distillation and Task-oriented generation (eTag), which requires neither exemplar nor prototype. Embedding distillation prevents the feature extractor from forgetting by distilling the outputs from the networks' intermediate blocks. Task-oriented generation enables a lightweight generator to produce dynamic features, fitting the needs of the top incremental classifier. Experimental results confirm that the proposed eTag considerably outperforms state-of-the-art methods on several benchmark datasets.

NeurIPS Conference 2024 Conference Paper

Identification and Estimation of the Bi-Directional MR with Some Invalid Instruments

  • Feng Xie
  • Zhen Yao
  • Lin Xie
  • Yan Zeng
  • Zhi Geng

We consider the challenging problem of estimating causal effects from purely observational data in the bi-directional Mendelian randomization (MR), where some invalid instruments, as well as unmeasured confounding, usually exist. To address this problem, most existing methods attempt to find proper valid instrumental variables (IVs) for the target causal effect by expert knowledge or by assuming that the causal model is a one-directional MR model. As such, in this paper, we first theoretically investigate the identification of the bi-directional MR from observational data. In particular, we provide necessary and sufficient conditions under which valid IV sets are correctly identified such that the bi-directional MR model is identifiable, including the causal directions of a pair of phenotypes (i. e. , the treatment and outcome). Moreover, based on the identification theory, we develop a cluster fusion-like method to discover valid IV sets and estimate the causal effects of interest. We theoretically demonstrate the correctness of the proposed algorithm. Experimental results show the effectiveness of our method for estimating causal effects in both one-directional and bi-directional MR models.

NeurIPS Conference 2024 Conference Paper

Learning the Optimal Policy for Balancing Short-Term and Long-Term Rewards

  • Qinwei Yang
  • Xueqing Liu
  • Yan Zeng
  • Ruocheng Guo
  • Yang Liu
  • Peng Wu

Learning the optimal policy to balance multiple short-term and long-term rewards has extensive applications across various domains. Yet, there is a noticeable scarcity of research addressing policy learning strategies in this context. In this paper, we aim to learn the optimal policy capable of effectively balancing multiple short-term and long-term rewards, especially in scenarios where the long-term outcomes are often missing due to data collection challenges over extended periods. Towards this goal, the conventional linear weighting method, which aggregates multiple rewards into a single surrogate reward through weighted summation, can only achieve sub-optimal policies when multiple rewards are related. Motivated by this, we propose a novel decomposition-based policy learning (DPPL) method that converts the whole problem into subproblems. The DPPL method is capable of obtaining optimal policies even when multiple rewards are interrelated. Nevertheless, the DPPL method requires a set of preference vectors specified in advance, posing challenges in practical applications where selecting suitable preferences is non-trivial. To mitigate this, we further theoretically transform the optimization problem in DPPL into an $\varepsilon$-constraint problem, where $\varepsilon$ represents the minimum acceptable levels of other rewards while maximizing one reward. This transformation provides intuitive into the selection of preference vectors. Extensive experiments are conducted on the proposed method and the results validate the effectiveness of the method.

JMLR Journal 2023 Journal Article

Python package for causal discovery based on LiNGAM

  • Takashi Ikeuchi
  • Mayumi Ide
  • Yan Zeng
  • Takashi Nicholas Maeda
  • Shohei Shimizu

Causal discovery is a methodology for learning causal graphs from data, and LiNGAM is a well-known model for causal discovery. This paper describes an open-source Python package for causal discovery based on LiNGAM. The package implements various LiNGAM methods under different settings like time series cases, multiple-group cases, mixed data cases, and hidden common cause cases, in addition to evaluation of statistical reliability and model assumptions. The source code is freely available under the MIT license at https://github.com/cdt15/lingam. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2023. ( edit, beta )

CLeaR Conference 2022 Conference Paper

Causal Discovery for Linear Mixed Data

  • Yan Zeng
  • Shohei Shimizu
  • Hidetoshi Matsui
  • Fuchun Sun

Discovery of causal relationships from observational data, especially from mixed data that consist of both continuous and discrete variables, is a fundamental yet challenging problem. Traditional methods focus on polishing the data type processing policy, which may lose data information. Compared with such methods, the constraint-based and score-based methods for mixed data derive certain conditional independence tests or score functions from the data’s characteristics. However, they may return the Markov equivalence class due to the lack of identifiability guarantees, which may limit their applicability or hinder their interpretability of causal graphs. Thus, in this paper, based on the structural causal models of continuous and discrete variables, we provide sufficient identifiability conditions in bivariate as well as multivariate cases. We show that if the data follow our proposed restricted Linear Mixed causal model (LiM), such a model is identifiable. In addition, we proposed a two-step hybrid method to discover the causal structure for mixed data. Experiments on both synthetic and real-world data empirically demonstrate the identifiability and efficacy of our proposed LiM model.

IJCAI Conference 2021 Conference Paper

Causal Discovery with Multi-Domain LiNGAM for Latent Factors

  • Yan Zeng
  • Shohei Shimizu
  • Ruichu Cai
  • Feng Xie
  • Michio Yamamoto
  • Zhifeng Hao

Discovering causal structures among latent factors from observed data is a particularly challenging problem. Despite some efforts for this problem, existing methods focus on the single-domain data only. In this paper, we propose Multi-Domain Linear Non-Gaussian Acyclic Models for LAtent Factors (MD-LiNA), where the causal structure among latent factors of interest is shared for all domains, and we provide its identification results. The model enriches the causal representation for multi-domain data. We propose an integrated two-phase algorithm to estimate the model. In particular, we first locate the latent factors and estimate the factor loading matrix. Then to uncover the causal structure among shared latent factors of interest, we derive a score function based on the characterization of independence relations between external influences and the dependence relations between multi-domain latent factors and latent factors of interest. We show that the proposed method provides locally consistent estimators. Experimental results on both synthetic and real-world data demonstrate the efficacy and robustness of our approach.

v2026.09.13