Arrow Research search

Author name cluster

Zihao Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

35 papers
2 author rows

Possible papers

35

AAAI Conference 2026 Conference Paper

CyPortQA: Benchmarking Multimodal Large Language Models for Cyclone Preparedness in Port Operation

  • Chenchen Kuai
  • Chenhao Wu
  • Yang Zhou
  • Bruce Wang
  • Tianbao Yang
  • Zhengzhong Tu
  • Zihao Li
  • Yunlong Zhang

As tropical cyclones intensify and track forecasts become increasingly uncertain, U.S. ports face heightened supply-chain risk under extreme weather conditions. Port operators need to rapidly synthesize diverse multimodal forecast products, such as probabilistic wind maps, track cones, and official advisories, into clear, actionable guidance as cyclones approach. Multimodal large language models (MLLMs) offer a powerful means to integrate these heterogeneous data sources alongside broader contextual knowledge, yet their accuracy and reliability in the specific context of port cyclone preparedness have not been rigorously evaluated. To fill this gap, we introduce CyPortQA, the first multimodal benchmark tailored to port operations under cyclone threat. CyPortQA assembles 2,917 real-world disruption scenarios from 2015 through 2023, spanning 145 U.S. principal ports and 90 named storms. Each scenario fuses multi-source data (i.e., tropical cyclone products, port operational impact records, and port condition bulletins) and is expanded through an automated pipeline into 117,178 structured question–answer pairs. Using this benchmark, we conduct extensive experiments on diverse MLLMs, including both open-source and proprietary model. MLLMs demonstrate great potential in situation understanding but still face considerable challenges in reasoning tasks, including potential impact estimation and decision reasoning.

I&C Journal 2026 Journal Article

Fair division with prioritized agents

  • Xiaolin Bu
  • Zihao Li
  • Shengxin Liu
  • Jiaxin Song
  • Biaoshuai Tao
  • Ziqi Yu

We study the fair division of indivisible items. Since an envy-free allocation may not exist, a standard relaxation is envy-freeness up to one item (EF1), where any envy can be eliminated by removing a single item from the envied agent's bundle. In many applications, however, it is desirable to designate a subset of prioritized agents for whom strict envy-freeness toward the remaining agents must be guaranteed, while the overall allocation remains EF1. Such agents may correspond to those who were envious in a previous EF1 allocation or to members of underrepresented groups. Motivated by this, we propose a new fairness notion named envy-freeness with prioritized agents EFprior, and study the existence and the algorithmic aspects of computing an EFprior allocation. For additive valuations, the simple round-robin algorithm suffices to compute an EFprior allocation. In this paper, we mainly focus on general valuations. In particular, we present a polynomial-time algorithm that computes an EFprior allocation with most of the items allocated. When all the items need to be allocated, we also present polynomial-time algorithms for several well-motivated special cases. We finally extend the setting to a general prioritized ordering case, where we are given a full ordering of agents and each agent with a higher priority cannot envy an agent with a lower priority. We propose a generalized fairness notion named envy-freeness under rank r EF r and present a polynomial-time algorithm with most of the items allocated.

EAAI Journal 2026 Journal Article

Physics-enhanced simulation-to-measurement translation for rolling bearing fault diagnosis under limited samples

  • Zhen Ming
  • Baoping Tang
  • Lei Deng
  • Xiaolong Zhang
  • Zihao Li

The application of deep learning in fault diagnosis is constrained by limited samples. While dynamic simulation data have been widely utilized to augment datasets, significant distribution discrepancies between simulation and measurement data remain, leading to suboptimal diagnostic performance. To bridge this gap, this paper proposes a Physics-Enhanced Generative Adversarial Network (PEGAN), which incorporates physics-informed learning to achieve high-fidelity simulation-to-measurement translation. The framework includes a bearing dynamics simulation model that generates abundant simulation data, providing comprehensive prior knowledge for PEGAN. To enhance feature alignment, an adaptive noise injection module is introduced to embed real-world noise characteristics into simulation data. In addition, a time-frequency bidomain-aware adjustment module is designed to perform joint time-frequency alignment of simulation data, thereby enhancing the network's feature adjustment capability. Furthermore, a physics-embedded loss function is introduced to optimize the network by minimizing discrepancies in key physical features between generated and measurement data. The augmented dataset generated by PEGAN is then used to train diagnostic classifiers. Comprehensive experimental validations on two test rig datasets and one real-world engineering case demonstrate that PEGAN significantly improves fault diagnosis accuracy under limited sample conditions, offering an innovative physics-guided paradigm for industrial fault diagnosis.

FM Conference 2026 Conference Paper

QSeqSim: A Symbolic Simulator for Qiskit While Loops Using Sequential Quantum Circuits (Long Tool Paper)

  • Zihao Li
  • Ji Guan
  • Mingsheng Ying

Abstract We present a tool QSeqSim, a Qiskit-integrated symbolic backend that fills the current gap of having no Qiskit-native support for simulating -loop quantum programs and their induced sequential quantum circuits. QSeqSim takes Qiskit objects, translates them into OpenQASM 3 code, and organises the resulting program into a combination of combinational, dynamic, and sequential circuits, thereby assigning -loops a precise sequential circuit semantics with explicit internal and external qubits. Building on this semantics, QSeqSim adopts a Binary Decision Diagram (BDD)-based symbolic representation and integrates weighted model counting to compute measurement probabilities efficiently by exploiting sharing in structured and sparse BDDs. On top of this Boolean backbone, it introduces dedicated symbolic operators for quantum state composition and state retention, thereby enabling efficient symbolic execution of sequential quantum circuits. Our experiments demonstrate that QSeqSim scales to substantial -induced sequential circuits; in particular, in the quantum random walk benchmark we successfully simulate circuits with over 1000 qubits for more than 10 loop iterations. QSeqSim is available at https: //github. com/Veri-Q/QSeqSim.

EAAI Journal 2026 Journal Article

Vibration mechanism driven discrete wavelet hybrid attention weighted transfer network for partial domain fault diagnosis of gearboxes under data scarcity

  • Peng Zhu
  • Baoping Tang
  • Lei Deng
  • Jing Wei
  • Zihao Li
  • Qikang Li

Data-driven intelligent fault diagnosis methods play a vital role in maintaining safe and reliable operation of gearboxes. Currently, some simulation data-driven transfer learning methods have been developed to achieve gearboxes fault diagnosis under the scarcity of high-quality labeled data. However, these methods still have problems such as insufficient domain-invariant feature extraction capabilities and negative transfer caused by large discrepancies between simulation and measured data, which poses challenges to their deployment in real-world industrial environments. To address these challenges, a vibration mechanism driven discrete wavelet hybrid attention weighted transfer network (DWAWTN) is proposed for partial domain fault diagnosis of gearboxes. Firstly, a vibration signal model that can reflect the gear failure vibration response mechanism is established, which can generate simulation data of different fault types. Secondly, a hybrid attention-guided multi-scale discrete wavelet convolutional network is proposed to extract domain-invariant features of simulation and measured data, which dynamically focuses on the multi-resolution fault features after wavelet decomposition from the channel and spatial dimensions. Then, a domain adaptation method combining the maximum likelihood weight estimation strategy and pseudo-label self-learning technology is constructed, which can not only suppress the negative transfer effect of outlier classes data in the source domain, but also reduce the uncertainty of the prediction of measured target data. Finally, comprehensive experimental verification and analysis are carried out on two gearbox datasets. Experimental results show that DWAWTN outperforms other compared methods in diagnostic performance.

NeurIPS Conference 2025 Conference Paper

Chain-of-Model Learning for Language Model

  • Xiaohua Wang
  • Kaitao Song
  • Xu Tan
  • Huiqiang Jiang
  • Chengruidong Zhang
  • Yongliang Shen
  • Cen Lu
  • Zihao Li

In this paper, we propose a novel learning paradigm, termed Chain-of-Model (CoM), which incorporates the causal relationship into the hidden states of each layer as a chain style. thereby introducing great scaling efficiency in model training and inference flexibility in deployment. We introduce the concept of Chain-of-Representation (CoR), which formulates the hidden states at each layer as a combination of multiple sub-representations (i. e. , chains). In each layer, each chain from the output representations can only view all of its preceding chains in the input representations. Consequently, the model built upon CoM framework can progressively scale up the model size by increasing the chains based on the previous models (i. e. , chains), and offer multiple sub-models at varying sizes for elastic inference by using different chain numbers. Based on this principle, we devise Chain-of-Language-Model (CoLM), which incorporates the idea of CoM into each layer of Transformer architecture. Based on CoLM, we further introduce CoLM-Air by introducing a KV sharing mechanism, that computes all keys and values within the first chain and then shares across all chains. This design demonstrates additional extensibility, such as enabling seamless LM switching, prefilling acceleration and so on. Experimental results demonstrate our CoLM family can achieve comparable performance to the standard Transformer, while simultaneously enabling greater flexiblity, such as progressive scaling to improve training efficiency and offer multiple varying model sizes for elastic inference, paving a a new way toward building language models.

NeurIPS Conference 2025 Conference Paper

CLIMB: Class-imbalanced Learning Benchmark on Tabular Data

  • Zhining Liu
  • Zihao Li
  • Ze Yang
  • Tianxin Wei
  • Jian Kang
  • Yada Zhu
  • Hendrik Hamann
  • Jingrui He

Class-imbalanced learning (CIL) on tabular data is important in many real-world applications where the minority class holds the critical but rare outcomes. In this paper, we present CLIMB, a comprehensive benchmark for class-imbalanced learning on tabular data. CLIMB includes 73 real-world datasets across diverse domains and imbalance levels, along with unified implementations of 29 representative CIL algorithms. Built on a high-quality open-source Python package with unified API designs, detailed documentation, and rigorous code quality controls, CLIMB supports easy implementation and comparison between different CIL algorithms. Through extensive experiments, we provide practical insights on method accuracy and efficiency, highlighting the limitations of naive rebalancing, the effectiveness of ensembles, and the importance of data quality. Our code, documentation, and examples are available at https: //github. com/ZhiningLiu1998/imbalanced-ensemble.

EAAI Journal 2025 Journal Article

Discriminative feature learning using class-aware and selective transfer adversarial network for partial cross-domain fault diagnosis of gearboxes

  • Peng Zhu
  • Baoping Tang
  • Lei Deng
  • Qikang Li
  • Zihao Li

Recently, partial domain adaptation (PDA) fault diagnosis methods have gained significant attention, as they assume that the target domain (TD) data class labels are only a subset of the source domain (SD) data class labels, making them more closely aligned with actual industrial scenarios. Most existing PDA methods adopt reweighting strategies to suppress the influence of outlier classes in the SD. However, these methods still face two major challenges: involving all SD samples in network training causes premature domain discrepancy, hindering the improvement of model transfer performance in the later stage of training; the relationship between target class samples is ignored, which makes the network classifier insufficient in extracting TD discriminable features. To solve these challenges, this study proposes a class-aware and selective transfer adversarial network (CSTAN) for partial cross-domain fault diagnosis of gearboxes. Firstly, the source data selector (SDS) in the CSTAN network is proposed to automatically filter out the source outlier data to avoid negative transfer; Secondly, the residual multi-channel attention mechanism feature extractor network is constructed to extract the features from the source data filtered by SDS and the target data, and the entropy-enhanced domain adversarial metric is adopted to promote positive transfer; Thirdly, to improve the model's ability to extract discriminable features of the TD, a target class-aware classification mechanism based on pseudo-label self-training technology is designed; Finally, the superiority and effectiveness of the CSTAN method are verified on the datasets of drivetrain diagnostics simulator test bench and wind turbine gearboxes.

JBHI Journal 2025 Journal Article

Graph Attention Fusion With Kolmogorov-Arnold Network for Drug-Gene Interaction Prediction

  • Xinguo Lu
  • Zihao Li
  • Ping Liu
  • Anqi Tang
  • Xing Liu
  • Hongrui Liu

Deep learning-based computational methods have emerged as powerful tools for predicting novel drug-gene interactions. It is essential to parse the joint influence of diverse attention focuses in large complex datasets in the model's decision-making process. Here, we propose graph attention fusion with Kolmogorov-Arnold network (KAN) for drug-gene interaction prediction (dgKAN). This approach parses the mutual influence of heterogeneous attention in drug-gene relationships by constructing an interpretable KAN network. Specifically, we use dynamic neighbor selection module by dynamic attention sampling to construct subgraphs and generate embedding representations for drugs and genes within these subgraphs. Then, we utilize a module consisted of Transformer and GNN architectures (TransGNN) to fuse the mechanism of global attention and local attention. Finally, we develop an interpretable KAN network with spline functions to model and analyze the cross-domain information flow between drugs and genes, enabling the prediction of drug-gene interactions. We conducted comprehensive experiments on various datasets, and the results demonstrate that dgKAN outperforms other baseline methods. Meanwhile, results illustrate that dgKAN captures the implicit characteristics by parsing heterogeneous attention in drug-gene relationships. The predicted drug-gene interactions have the potential to significantly aid in drug development for disease treatment.

TMLR Journal 2025 Journal Article

Hypergraphs as Weighted Directed Self-Looped Graphs: Spectral Properties, Clustering, Cheeger Inequality

  • Zihao Li
  • Dongqi Fu
  • Hengyu Liu
  • Jingrui He

Hypergraphs naturally arise when studying group relations and have been widely used in the field of machine learning. To the best of our knowledge, the recently proposed edge-dependent vertex weights (EDVW) modeling is one of the most generalized modeling methods of hypergraphs, i.e., most existing hypergraph conceptual modeling methods can be generalized as EDVW hypergraphs without information loss. However, the relevant algorithmic developments on EDVW hypergraphs remain nascent: compared to the spectral theories for graphs, its formulations are incomplete, the spectral clustering algorithms are not well-developed, and the hypergraph Cheeger Inequality is not well-defined. To this end, deriving a unified random walk-based formulation, we propose our definitions of hypergraph Rayleigh Quotient, NCut, boundary/cut, volume, and conductance, which are consistent with the corresponding definitions on graphs. Then, we prove that the normalized hypergraph Laplacian is associated with the NCut value, which inspires our proposed HyperClus-G algorithm for spectral clustering on EDVW hypergraphs. Finally, we prove that HyperClus-G can always find an approximately linearly optimal partitioning in terms of both NCut and conductance. Additionally, we provide extensive experiments to validate our theoretical findings from an empirical perspective.

AAAI Conference 2025 Conference Paper

Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages

  • Zihao Li
  • Yucheng Shi
  • Zirui Liu
  • Fan Yang
  • Ali Payani
  • Ninghao Liu
  • Mengnan Du

The development of Large Language Models (LLMs) relies on extensive text corpora, which are often unevenly distributed across languages. This imbalance results in LLMs performing significantly better on high-resource languages like English, German, and French, while their capabilities in low-resource languages remain inadequate. Currently, there is a lack of quantitative methods to evaluate the performance of LLMs in these low-resource languages. To address this gap, we propose the Language Ranker, an intrinsic metric designed to benchmark and rank languages based on LLM performance using internal representations. By comparing the LLM's internal representation of various languages against a baseline derived from English, we can assess the model's multilingual capabilities in a robust and language-agnostic manner. Our analysis reveals that high-resource languages exhibit higher similarity scores with English, demonstrating superior performance, while low-resource languages show lower similarity scores, underscoring the effectiveness of our metric in assessing language-specific capabilities. Besides, the experiments show that there is a strong correlation between the LLM’s performance in different languages and the proportion of those languages in its pre-training corpus. These insights underscore the efficacy of the Language Ranker as a tool for evaluating LLM performance across different languages, particularly those with limited resources.

IROS Conference 2025 Conference Paper

Lip Geometry-Constrained Smooth Sliding Path Planning for Robotic Negative Pressure Therapy on Extremities

  • Zihao Li
  • Zhenguo Nie
  • Qi Shao
  • Huichan Zhao
  • Xin-Jun Liu

Negative pressure (NP) therapy with sliding suction is an effective method for limb lymphedema. Due to the caregiver shortage and the patients increase, the robotic NP therapeutic system with a variable-sized suction head can be used to help the lymphedema therapy. However, the varying complexity of different limb regions can affect the accuracy of the suction path. Moreover, the moving suction path should maintain smoothness to ensure therapeutic efficacy. Therefore, finding a smooth sliding path with highly accurate suction poses on the unstructured limb surface poses a significant challenge for robotic therapy. In this paper, a smooth sliding path planning method is proposed for robotic continuous suction in limb lymphedema therapy. The easily-sealed region is identified by comparing point normals to the lip’s suction angle, simplifying path planning to a 2D plane due to lip and limb flexibility. The conjugate gradient method optimizes the path with centroid distance and smoothness constraints. Finally, after the generation of suction poses under the constraints of the lip shape, a smooth sliding path along with lip pressure commands, is obtained to regulate the robot in performing continuous suction therapy. In the experiment, the manipulator with a variable-sized head has been used to finish 10 sliding suctions from different planning path. From the result, the robot could complete 6 times sliding suctions on the phantom arm.

ICML Conference 2025 Conference Paper

MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations

  • Kaixuan Huang
  • Jiacheng Guo
  • Zihao Li
  • Xiang Ji
  • Jiawei Ge 0003
  • Wenzhe Li
  • Yingqing Guo
  • Tianle Cai

Large language models have demonstrated impressive performance on challenging mathematical reasoning tasks, which has triggered the discussion of whether the performance is achieved by true reasoning capability or memorization. To investigate this question, prior work has constructed mathematical benchmarks when questions undergo simple perturbations – modifications that still preserve the underlying reasoning patterns of the solutions. However, no work has explored hard perturbations, which fundamentally change the nature of the problem so that the original solution steps do not apply. To bridge the gap, we construct MATH-P-Simple and MATH-P-Hard via simple perturbation and hard perturbation, respectively. Each consists of 279 perturbed math problems derived from level-5 (hardest) problems in the MATH dataset (Hendrycks et al. , 2021). We observe significant performance drops on MATH-P-Hard across various models, including o1-mini (-16. 49%) and gemini-2. 0-flash-thinking (-12. 9%). We also raise concerns about a novel form of memorization where models blindly apply learned problem-solving skills without assessing their applicability to modified contexts. This issue is amplified when using original problems for in-context learning. We call for research efforts to address this challenge, which is critical for developing more robust and reliable reasoning models. The project is available at https: //math-perturb. github. io/.

NeurIPS Conference 2025 Conference Paper

Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning

  • Jiaru Zou
  • Yikun Ban
  • Zihao Li
  • Yunzhe Qi
  • Ruizhong Qiu
  • Ling Yang
  • Jingrui He

Large language models are typically adapted to downstream tasks through supervised fine-tuning on domain-specific data. While standard fine-tuning focuses on minimizing generation loss to optimize model parameters, we take a deeper step by retaining and leveraging the model’s own learning signals, analogous to how human learners reflect on past mistakes to improve future performance. We first introduce the concept of Mistake Log to systematically track the model’s learning behavior and recurring errors throughout fine-tuning. Treating the original transformer-based model as the Pilot, we correspondingly design a Copilot model to refine the Pilot’s inference performance via logits rectification. We name the overall Pilot-Copilot framework the Transformer Copilot, which introduces (i) a novel Copilot model design, (ii) a joint training paradigm where the Copilot continuously learns from the evolving Mistake Log alongside the Pilot, and (iii) a fused inference paradigm where the Copilot rectifies the Pilot’s logits for enhanced generation. We provide both theoretical and empirical analyses on our new learning framework. Experiments on 12 benchmarks spanning commonsense, arithmetic, and recommendation tasks demonstrate that Transformer Copilot consistently improves performance by up to 34. 5%, while introducing marginal computational overhead to Pilot models and exhibiting strong scalability and transferability. Our code is released at https: //github. com/jiaruzouu/TransformerCopilot.

AAMAS Conference 2024 Conference Paper

A Complete Landscape for the Price of Envy-Freeness

  • Zihao Li
  • Shengxin Liu
  • Xinhang Lu
  • Biaoshuai Tao
  • Yichen Tao

We study the efficiency of fair allocations using the well-studied price of fairness concept, which quantitatively measures the worstcase efficiency loss when imposing fairness constraints. Previous works provided partial results on the price of fairness with wellknown fairness notions such as envy-freeness up to one good (EF1) and envy-freeness up to any good (EFX). In this paper, we give a complete characterization for the price of envy-freeness in various settings. In particular, we first consider the two-agent case under the indivisible-goods setting and present tight ratios for the price of EF1 (for scaled utility) and EFX (for unscaled utility), which resolve questions left open in the literature. Next, we consider the mixed goods setting which concerns a mixture of both divisible and indivisible goods. We focus on envy-freeness for mixed goods (EFM), which generalizes both envy-freeness and EF1, as well as its strengthening called envy-freeness up to any good for mixed goods (EFXM), which generalizes envy-freeness and EFX. To this end, we settle the price of EFM and EFXM by providing a complete picture of tight bounds for two agents and asymptotically tight bounds for 𝑛 agents, for both scaled and unscaled utilities.

IJCAI Conference 2024 Conference Paper

Allocating Mixed Goods with Customized Fairness and Indivisibility Ratio

  • Bo Li
  • Zihao Li
  • Shengxin Liu
  • Zekai Wu

We consider the problem of fairly allocating a combination of divisible and indivisible goods. While fairness criteria like envy-freeness (EF) and proportionality (PROP) can always be achieved for divisible goods, only their relaxed versions, such as the “up to one” relaxations EF1 and PROP1, can be satisfied when the goods are indivisible. The “up to one” relaxations require the fairness conditions to be satisfied provided that one good can be completely eliminated or added in the comparison. In this work, we bridge the gap between the two extremes and propose “up to a fraction” relaxations for the allocation of mixed divisible and indivisible goods. The fraction is determined based on the proportion of indivisible goods, which we call the indivisibility ratio. The new concepts also introduce asymmetric conditions that are customized for individuals with varying indivisibility ratios. We provide both upper and lower bounds on the fractions of the modified item in order to satisfy the fairness criterion. Our results are tight up to a constant for EF and asymptotically tight for PROP.

IROS Conference 2024 Conference Paper

Efficient Path Planning for Modular Reconfigurable Robots

  • Matthias Mayer
  • Zihao Li
  • Matthias Althoff

Industrial robots are essential for modern production but often struggle to adapt to new tasks. Modular (reconfigurable) robots can overcome this challenge by eliminating the need to replace the whole robot. However, finding the optimal assembly for a task remains difficult because a valid path has to be computed for each generated assembly – consuming a significant fraction of the computation time. Similar to online path planning, where previous approaches adapt known paths to a changing environment, we show that transferring paths from previously considered module assemblies accelerates path planning for the next assemblies. On average, our method reduces the planning time for single-goal tasks by 50%. The usefulness of our method is evaluated by integrating it in a genetic algorithm (GA) for optimizing assemblies and evaluating it on our benchmark suite CoBRA. Within the optimization loop for modular robots, the time used to check a single assembly is shortened by up to 50%.

EAAI Journal 2024 Journal Article

Exploring explicit and implicit graph learning for multivariate time series imputation

  • Yakun Chen
  • Ruotong Hu
  • Zihao Li
  • Chao Yang
  • Xianzhi Wang
  • Guodong Long
  • Guandong Xu

Multivariate time series inherently contain missing values due to various issues, including incorrect data entry, broken equipment, and package loss during data transferring. The successful completion of time series data analysis tasks heavily relies on the essential task of imputing missing values. Inter-variable relationships in time series are typically overlooked by missing value imputation techniques. Although some graph-based algorithms can capture these relationships, the design of graph structures is commonly handcrafted and dataset-centric. We introduce a novel Explicit and Implicit Graph Recurrent Network (EIGRN) for multivariate time series imputation that integrates graph and recurrent neural networks to capture variable and time dependencies together. This proficiency is achieved by effectively integrating external data sources such as domain knowledge and the implicit relationships among nodes. In order to make our approach more applicable to datasets with larger numbers of missing values, we additionally discuss the model’s performance for various missing value ratios. Our comprehensive experiments on real-world datasets show that our model outperforms state-of-the-art baselines in different industrial fields.

NeurIPS Conference 2024 Conference Paper

Global Convergence in Training Large-Scale Transformers

  • Cheng Gao
  • Yuan Cao
  • Zihao Li
  • Yihan He
  • Mengdi Wang
  • Han Liu
  • Jason M. Klusowski
  • Jianqing Fan

Despite the widespread success of Transformers across various domains, their optimization guarantees in large-scale model settings are not well-understood. This paper rigorously analyzes the convergence properties of gradient flow in training Transformers with weight decay regularization. First, we construct the mean-field limit of large-scale Transformers, showing that as the model width and depth go to infinity, gradient flow converges to the Wasserstein gradient flow, which is represented by a partial differential equation. Then, we demonstrate that the gradient flow reaches a global minimum consistent with the PDE solution when the weight decay regularization parameter is sufficiently small. Our analysis is based on a series of novel mean-field techniques that adapt to Transformers. Compared with existing tools for deep networks (Lu et al. , 2020) that demand homogeneity and global Lipschitz smoothness, we utilize a refined analysis assuming only $\textit{partial homogeneity}$ and $\textit{local Lipschitz smoothness}$. These new techniques may be of independent interest.

EAAI Journal 2024 Journal Article

Intermittent fault diagnosis of analog circuit based on enhanced one-dimensional vision transformer and transfer learning strategy

  • Shengdong Wang
  • Zhenbao Liu
  • Zhen Jia
  • Wen Zhao
  • Zihao Li
  • Luyao Wang

As the major cause of false alarms in built-in test (BIT) system, intermittent faults of analog circuits may trigger abnormal equipment shutdown and lead to catastrophic accidents. With complete randomness and great non-repeatability, intermittent faults are arduous to be detected. To enhance the reliability and safety of electronic systems, an end-to-end approach based on enhanced one-dimensional Vision Transformer (1DViT) is proposed to realize intelligent diagnosis for intermittent faults of analog circuits. The signal anomaly caused by intermittent faults can be regarded as a kind of random anomaly from global perspective, and there are also rich local feature information in the fault interval. Completely composed of self-attention mechanism, Vision Transformer possesses prominent performance on extracting global features and modelling global representations, thus can be applied to identify intermittent faults. Meanwhile, to further enrich the feature representation, one multi-scale convolution fusion module (MSC) incorporating a series of convolution operations is designed and combined with 1DViT to extract and fuse the valuable local information. However, in practical test, due to the complex operation process, it is cumbersome to collect sufficient fault data to guarantee the effective training of the proposed model. To cope with this problem, transfer learning strategy is introduced. The model will be first pre-trained with adequate simulation data which is easily accessible, and then fine-tuned with a relatively small amount of actual fault data to help match the practical feature distribution. Experiments on two typical circuits demonstrate that the proposed method could achieve excellent diagnostic result in practical test.

ICML Conference 2024 Conference Paper

Meta-Reinforcement Learning Robust to Distributional Shift Via Performing Lifelong In-Context Learning

  • Tengye Xu
  • Zihao Li
  • Qinyuan Ren

A key challenge in Meta-Reinforcement Learning (meta-RL) is the task distribution shift, since the generalization ability of most current meta-RL methods is limited to tasks sampled from the training distribution. In this paper, we propose Posterior Sampling Bayesian Lifelong In-Context Reinforcement Learning (PSBL), which is robust to task distribution shift. PSBL meta-trains a variant of transformer to directly perform amortized inference about the Predictive Posterior Distribution (PPD) of the optimal policy. Once trained, the network can infer the PPD online with frozen parameters. The agent then samples actions from the approximate PPD to perform online exploration, which progressively reduces uncertainty and enhances performance in the interaction with the environment. This property is known as in-context learning. Experimental results demonstrate that PSBL significantly outperforms standard Meta RL methods both in tasks with sparse rewards and dense rewards when the test task distribution is strictly shifted from the training distribution.

NeurIPS Conference 2024 Conference Paper

One-Layer Transformer Provably Learns One-Nearest Neighbor In Context

  • Zihao Li
  • Yuan Cao
  • Cheng Gao
  • Yihan He
  • Han Liu
  • Jason M. Klusowski
  • Jianqing Fan
  • Mengdi Wang

Transformers have achieved great success in recent years. Interestingly, transformers have shown particularly strong in-context learning capability -- even without fine-tuning, they are still able to solve unseen tasks well purely based on task-specific prompts. In this paper, we study the capability of one-layer transformers in learning the one-nearest neighbor prediction rule. Under a theoretical framework where the prompt contains a sequence of labeled training data and unlabeled test data, we show that, although the loss function is nonconvex, when trained with gradient descent, a single softmax attention layer can successfully learn to behave like a one-nearest neighbor classifier. Our result gives a concrete example on how transformers can be trained to implement nonparametric machine learning algorithms, and sheds light on the role of softmax attention in transformer models.

NeurIPS Conference 2024 Conference Paper

PageRank Bandits for Link Prediction

  • Yikun Ban
  • Jiaru Zou
  • Zihao Li
  • Yunzhe Qi
  • Dongqi Fu
  • Jian Kang
  • Hanghang Tong
  • Jingrui He

Link prediction is a critical problem in graph learning with broad applications such as recommender systems and knowledge graph completion. Numerous research efforts have been directed at solving this problem, including approaches based on similarity metrics and Graph Neural Networks (GNN). However, most existing solutions are still rooted in conventional supervised learning, which makes it challenging to adapt over time to changing customer interests and to address the inherent dilemma of exploitation versus exploration in link prediction. To tackle these challenges, this paper reformulates link prediction as a sequential decision-making process, where each link prediction interaction occurs sequentially. We propose a novel fusion algorithm, PRB (PageRank Bandits), which is the first to combine contextual bandits with PageRank for collaborative exploitation and exploration. We also introduce a new reward formulation and provide a theoretical performance guarantee for PRB. Finally, we extensively evaluate PRB in both online and offline settings, comparing it with bandit-based and graph-based methods. The empirical success of PRB demonstrates the value of the proposed fusion approach. Our code is released at https: //github. com/jiaruzouu/PRB.

ICML Conference 2024 Conference Paper

Theoretical insights for diffusion guidance: A case study for Gaussian mixture models

  • Yuchen Wu
  • Minshuo Chen
  • Zihao Li
  • Mengdi Wang 0001
  • Yuting Wei 0001

Diffusion models benefit from instillation of task-specific information into the score function to steer the sample generation towards desired properties. Such information is coined as guidance. For example, in text-to-image synthesis, text input is encoded as guidance to generate semantically aligned images. Proper guidance inputs are closely tied with the performance of diffusion models. A common observation is that strong guidance promotes a tight alignment to the task-specific information, while reduces the diversity of the generated samples. In this paper, we provide the first theoretical study towards the influence of guidance on diffusion models in the context of Gaussian mixture models. Under mild conditions, we prove that incorporating diffusion guidance not only boosts prediction confidence but also diminishes distribution diversity, leading to a reduction in the differential entropy of the output distribution. Our analysis covers the widely used DDPM and DDIM sampling schemes, and leverages comparison inequalities in differential equations as well as the Fokker-Planck equation that characterizes the evolution of probability density function, which may be of independent theoretical interest.

JMLR Journal 2023 Journal Article

Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning

  • Zihao Li
  • Boyi Liu
  • Zhuoran Yang
  • Zhaoran Wang
  • Mengdi Wang

We study the Constrained Convex Markov Decision Process (MDP), where the goal is to minimize a convex functional of the visitation measure, subject to a convex constraint. Designing algorithms for a constrained convex MDP faces several challenges, including (1) handling the large state space, (2) managing the exploration/exploitation tradeoff, and (3) solving the constrained optimization where the objective and the constraint are both nonlinear functions of the visitation measure. In this work, we present a model-based algorithm, Variational Primal-Dual Policy Optimization (VPDPO), in which Lagrangian and Fenchel duality are implemented to reformulate the original constrained problem into an unconstrained primal-dual optimization. The primal variables are updated by model-based value iteration following the principle of Optimism in the Face of Uncertainty (OFU), while the dual variables are updated by gradient ascent. Moreover, by embedding the visitation measure into a finite-dimensional space, we can handle large state spaces by incorporating function approximation. Two notable examples are (1) Kernelized Nonlinear Regulators and (2) Low-rank MDPs. We prove that with an optimistic planning oracle, our algorithm achieves sublinear regret and constraint violation in both cases and can attain the globally optimal policy of the original constrained problem. [abs] [ pdf ][ bib ] &copy JMLR 2023. ( edit, beta )

AAAI Conference 2023 Conference Paper

Fair Division with Prioritized Agents

  • Xiaolin Bu
  • Zihao Li
  • Shengxin Liu
  • Jiaxin Song
  • Biaoshuai Tao

We consider the fair division problem of indivisible items. It is well-known that an envy-free allocation may not exist, and a relaxed version of envy-freeness, envy-freeness up to one item (EF1), has been widely considered. In an EF1 allocation, an agent may envy others' allocated shares, but only up to one item. In many applications, we may wish to specify a subset of prioritized agents where strict envy-freeness needs to be guaranteed from these agents to the remaining agents, while ensuring the whole allocation is still EF1. Prioritized agents may be those agents who are envious in a previous EF1 allocation, those agents who belong to underrepresented groups, etc. Motivated by this, we propose a new fairness notion named envy-freeness with prioritized agents EFprior, and study the existence and the algorithmic aspects for the problem of computing an EFprior allocation. With additive valuations, the simple round-robin algorithm is able to compute an EFprior allocation. In this paper, we mainly focus on general valuations. In particular, we present a polynomial-time algorithm that outputs an EFprior allocation with most of the items allocated. When all the items need to be allocated, we also present polynomial-time algorithms for some well-motivated special cases.

AAAI Conference 2023 Conference Paper

Fully Online Matching with Stochastic Arrivals and Departures

  • Zihao Li
  • Hao Wang
  • Zhenzhen Yan

We study a fully online matching problem with stochastic arrivals and departures. In this model, each online arrival follows a known identical and independent distribution over a fixed set of agent types. Its sojourn time is unknown in advance and follows type-specific distributions with known expectations. The goal is to maximize the weighted reward from successful matches. To solve this problem, we first propose a linear program (LP)-based algorithm whose competitive ratio is lower bounded by 0.155 under mild conditions. We further achieve better ratios in some special cases. To demonstrate the challenges of the problem, we further establish several hardness results. In particular, we show that no online algorithm can achieve a competitive ratio better than 2/3 in this model and there is no LP-based algorithm (with respect to our proposed LP) with a competitive ratio better than 1/3. Finally, we demonstrate the effectiveness and efficiency of our algorithm numerically.

EAAI Journal 2023 Journal Article

Incipient fault diagnosis of analog circuit with ensemble HKELM based on fused multi-channel and multi-scale features

  • Shengdong Wang
  • Zhenbao Liu
  • Zhen Jia
  • Zihao Li

As an essential part in electronics-rich system, the failure of analog circuits will severely affect the system reliability and security. Incipient fault of analog circuit refers to the early stage of degradation fault where the fault characteristics are generally weak and almost indistinguishable. In order to enhance the reliability of electronic systems, it is necessary to diagnose incipient faults of analog circuits promptly and effectively. Existing approaches generally capture fault characteristics only from single signal, ignoring the valuable information inherent in different domains and scales. To address this problem, a novel diagnostic strategy based on multi-scale feature extraction and multi-channel feature fusion is designed to guarantee the completeness and richness of fault information. In this study, a deep extreme learning machine denoising auto-encoder (DELM-DAE) based method is proposed to conduct unsupervised multi-scale and multi-channel feature fusion to extract distinguishable features for incipient faults. The proposed method has higher learning efficiency and overcomes the common problem of low efficiency in deep learning model training. Meanwhile, in order to improve the ability to distinguish high-resolution features, an ensemble hybrid kernel extreme learning machine with novel roulette selection and weighted voting scheme is proposed to enhance the recognition performance and stability. In the verification experiment, the diagnosis accuracy on four typical circuits all reaches above 98%, which demonstrates that the proposed incipient fault diagnosis method for analog circuits has more conspicuous performance than other state-of-the-art methods.

ICML Conference 2023 Conference Paper

Provably Efficient Representation Learning with Tractable Planning in Low-Rank POMDP

  • Jiacheng Guo
  • Zihao Li
  • Huazheng Wang
  • Mengdi Wang 0001
  • Zhuoran Yang
  • Xuezhou Zhang

In this paper, we study representation learning in partially observable Markov Decision Processes (POMDPs), where the agent learns a decoder function that maps a series of high-dimensional raw observations to a compact representation and uses it for more efficient exploration and planning. We focus our attention on the sub-classes of $\gamma$-observable and decodable POMDPs, for which it has been shown that statistically tractable learning is possible, but there has not been any computationally efficient algorithm. We first present an algorithm for decodable PMMDPs that combines maximum likelihood estimation (MLE) and optimism in the face of uncertainty (OFU) to perform representation learning and achieve efficient sample complexity, while only calling supervised learning computational oracles. We then show how to adapt this algorithm to also work in the broader class of $\gamma$-observable POMDPs.

AAAI Conference 2023 Conference Paper

Trusted Fine-Grained Image Classification through Hierarchical Evidence Fusion

  • Zhikang Xu
  • Xiaodong Yue
  • Ying Lv
  • Wei Liu
  • Zihao Li

Fine-Grained Image Classification (FGIC) aims to classify images into specific subordinate classes of a superclass. Due to insufficient training data and confusing data samples, FGIC may produce uncertain classification results that are untrusted for data applications. In fact, FGIC can be viewed as a hierarchical classification process and the multilayer information facilitates to reduce uncertainty and improve the reliability of FGIC. In this paper, we adopt the evidence theory to measure uncertainty and confidence in hierarchical classification process and propose a trusted FGIC method through fusing multilayer classification evidence. Comparing with the traditional approaches, the trusted FGIC method not only generates accurate classification results but also reduces the uncertainty of fine-grained classification. Specifically, we construct an evidence extractor at each classification layer to extract multilayer (multi-grained) evidence for image classification. To fuse the extracted multi-grained evidence from coarse to fine, we formulate evidence fusion with the Dirichlet hyper probability distribution and thereby hierarchically decompose the evidence of coarse-grained classes into fine-grained classes to enhance the classification performances. The ablation experiments validate that the hierarchical evidence fusion can improve the precision and also reduce the uncertainty of fine-grained classification. The comparison with state-of-the-art FGIC methods shows that our proposed method achieves competitive performances.

IJCAI Conference 2023 Conference Paper

Truthful Fair Mechanisms for Allocating Mixed Divisible and Indivisible Goods

  • Zihao Li
  • Shengxin Liu
  • Xinhang Lu
  • Biaoshuai Tao

We study the problem of designing truthful and fair mechanisms when allocating a mixture of divisible and indivisible goods. We first show that there does not exist an EFM (envy-free for mixed goods) and truthful mechanism in general. This impossibility result holds even if there is only one indivisible good and one divisible good and there are only two agents. Thus, we focus on some more restricted settings. Under the setting where agents have binary valuations on indivisible goods and identical valuations on a single divisible good (e. g. , money), we design an EFM and truthful mechanism. When agents have binary valuations over both divisible and indivisible goods, we first show there exist EFM and truthful mechanisms when there are only two agents or when there is a single divisible good. On the other hand, we show that the mechanism maximizing Nash welfare cannot ensure EFM and truthfulness simultaneously.

JBHI Journal 2023 Journal Article

Variable Augmented Network for Invertible Modality Synthesis and Fusion

  • Yuhao Wang
  • Ruirui Liu
  • Zihao Li
  • Shanshan Wang
  • Cailian Yang
  • Qiegen Liu

As an effective way to integrate the information contained in multiple medical images under different modalities, medical image synthesis and fusion have emerged in various clinical applications such as disease diagnosis and treatment planning. In this paper, an invertible and variable augmented network (iVAN) is proposed for medical image synthesis and fusion. In iVAN, the channel number of the network input and output is the same through variable augmentation technology, and data relevance is enhanced, which is conducive to the generation of characterization information. Meanwhile, the invertible network is used to achieve the bidirectional inference processes. Empowered by the invertible and variable augmentation schemes, iVAN not only be applied to the mappings of multi-input to one-output and multi-input to multi-output, but also to the case of one-input to multi-output. Experimental results demonstrated superior performance and potential task flexibility of the proposed method, compared with existing synthesis and fusion methods.

AIJ Journal 2021 Journal Article

Fair division of mixed divisible and indivisible goods

  • Xiaohui Bei
  • Zihao Li
  • Jinyan Liu
  • Shengxin Liu
  • Xinhang Lu

We study the problem of fair division when the set of resources contains both divisible and indivisible goods. Classic fairness notions such as envy-freeness (EF) and envy-freeness up to one good (EF1) cannot be directly applied to this mixed goods setting. In this work, we propose a new fairness notion, envy-freeness for mixed goods (EFM), which is a direct generalization of both EF and EF1 to the mixed goods setting. We prove that an EFM allocation always exists for any number of agents with additive valuations. We also propose efficient algorithms to compute an EFM allocation for two agents with general additive valuations and for n agents with piecewise linear valuations over the divisible goods. Finally, we relax the envy-freeness requirement, instead asking for ϵ-envy-freeness for mixed goods (ϵ-EFM), and present an efficient algorithm that finds an ϵ-EFM allocation.

ICML Conference 2021 Conference Paper

Fast Algorithms for Stackelberg Prediction Game with Least Squares Loss

  • Jiali Wang
  • He Chen
  • Rujun Jiang
  • Xudong Li 0004
  • Zihao Li

The Stackelberg prediction game (SPG) has been extensively used to model the interactions between the learner and data provider in the training process of various machine learning algorithms. Particularly, SPGs played prominent roles in cybersecurity applications, such as intrusion detection, banking fraud detection, spam filtering, and malware detection. Often formulated as NP-hard bi-level optimization problems, it is generally computationally intractable to find global solutions to SPGs. As an interesting progress in this area, a special class of SPGs with the least squares loss (SPG-LS) have recently been shown polynomially solvable by a bisection method. However, in each iteration of this method, a semidefinite program (SDP) needs to be solved. The resulted high computational costs prevent its applications for large-scale problems. In contrast, we propose a novel approach that reformulates a SPG-LS as a single SDP of a similar form and the same dimension as those solved in the bisection method. Our SDP reformulation is, evidenced by our numerical experiments, orders of magnitude faster than the existing bisection method. We further show that the obtained SDP can be reduced to a second order cone program (SOCP). This allows us to provide real-time response to large-scale SPG-LS problems. Numerical results on both synthetic and real world datasets indicate that the proposed SOCP method is up to 20, 000+ times faster than the state of the art.

AAAI Conference 2020 Conference Paper

Fair Division of Mixed Divisible and Indivisible Goods

  • Xiaohui Bei
  • Zihao Li
  • Jinyan Liu
  • Shengxin Liu
  • Xinhang Lu

We study the problem of fair division when the resources contain both divisible and indivisible goods. Classic fairness notions such as envy-freeness (EF) and envy-freeness up to one good (EF1) cannot be directly applied to the mixed goods setting. In this work, we propose a new fairness notion envyfreeness for mixed goods (EFM), which is a direct generalization of both EF and EF1 to the mixed goods setting. We prove that an EFM allocation always exists for any number of agents. We also propose efficient algorithms to compute an EFM allocation for two agents and for n agents with piecewise linear valuations over the divisible goods. Finally, we relax the envy-free requirement, instead asking for -envyfreeness for mixed goods ( -EFM), and present an algorithm that finds an -EFM allocation in time polynomial in the number of agents, the number of indivisible goods, and 1/.

v2026.09.13