Arrow Research search

Author name cluster

Yao Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

49 papers
2 author rows

Possible papers

49

AAAI Conference 2026 Conference Paper

AUVIC: Adversarial Unlearning of Visual Concepts for Multi-modal Large Language Models

  • Haokun Chen
  • Jianing Li
  • Yao Zhang
  • Jinhe Bi
  • Yan Xia
  • Jindong Gu
  • Volker Tresp

Multimodal Large Language Models (MLLMs) achieve impressive performance once optimized on massive datasets. Such datasets often contain sensitive or copyrighted content, raising significant data privacy concerns. Regulatory frameworks mandating the 'right to be forgotten' drive the need for machine unlearning. This technique allows for the removal of target data without resource-consuming retraining. However, while well-studied for text, visual concept unlearning in MLLMs remains underexplored. A primary challenge is precisely removing a target visual concept without disrupting model performance on related entities. To address this, we introduce AUVIC, a novel visual concept unlearning framework for MLLMs. AUVIC applies adversarial perturbations to enable precise forgetting. This approach effectively isolates the target concept while avoiding unintended effects on similar entities. To evaluate our method, we construct VCUBench. It is the first benchmark designed to assess visual concept unlearning in group contexts. Experimental results demonstrate that AUVIC achieves state-of-the-art target forgetting rates while incurs minimal performance degradation on non-target concepts.

AAMAS Conference 2026 Conference Paper

CentaurMD: Confidence-Aware Human-AI Decision Fusion for Multi-Label Disease Diagnosis via Label-Specific MoE

  • Youcheng Zhang
  • Hui Wang
  • Jiaqi Liu
  • Yao Zhang
  • Zhiwen Yu
  • Bin Guo

Multi-label disease diagnosis is prevalent in clinical applications, such as chest X-rays that may indicate multiple coexisting diseases. Despite advances in AI, current models remain insufficient for reliably addressing such complexity. Human–AI synergy thus emerges as both a necessary and promising approach, motivating our focus on effective decision fusion for multi-label disease diagnosis. There are two challenges. Confidence, a key factor in decision fusion, is often unrecorded in human annotations, making its estimation nontrivial. Moreover, label-specific variations in human and model expertise must be considered to achieve effective fusion. To address these challenges, we propose CentaurMD, a confidence-aware human–AI decision fusion framework based on label-specific Mixture-of-Experts (MoE). We first present a novel multi-label confusion matrix construction method that employs maximum entropy modeling to capture label correlations, enabling more accurate confidence estimation and weight allocation. Then, we develop a label-specific MoE module with dedicated gating networks and thresholds, which dynamically adjust expert weights using information extracted from the confusion matrix via a Transformer encoder. Extensive experiments on three real-world clinical datasets demonstrate that our method reduces Hamming loss by 39. 14% and improves MMR (missed-misdiagnosis reduction) by 17. 38%, achieving substantial diagnostic improvements.

JBHI Journal 2026 Journal Article

EAP-LSTM: A Bi-LSTM-Based Deep Learning Framework for Quantitatively Predicting Enhancer Activity in Drosophila and Human Cell Lines

  • Yao Zhang
  • Lichang Dai
  • Yu Dou
  • Xin Li
  • Chang Lu
  • Hao Wu

Enhancer activity plays a critical role in gene regulation, influencing various biological processes such as development and disease progression. Accurate prediction of enhancer activity is essential for understanding the mechanisms underlying gene regulation and enhancer function. This study introduces a novel deep learning framework, EAP-LSTM (Enhancer Activity Prediction based on Bi-LSTM), to quantitatively predict enhancer activity across different species and cell lines. The model integrates multiple feature modules, including Word2Vec-based representations of DNA sequences, reverse complement k-mer, mismatch k-mer features, and epigenomic data. Evaluated on six cell lines, including five human cell lines (A549, HCT116, HepG2, K562, and MCF-7) and one Drosophila cell line (S2), EAP-LSTM consistently outperforms state-of-the-art models, such as DeepSTARR and HEAP, in all datasets. For example, on the K562 dataset, EAP-LSTM achieves a Pearson correlation coefficient (PCC) of 0. 7944, outperforming DeepSTARR and HEAP by 13. 65% and 2. 73%, respectively. In addition, EAP-LSTM demonstrates strong performance in small-sample learning scenarios, showing clear improvements compared with baseline models. Furthermore, the study investigates the role of transcription factor binding sites (TFBSs) within enhancer regions, identifying critical motifs associated with enhancer activity. These findings not only improve enhancer prediction accuracy but also provide valuable insights into the molecular mechanisms underlying enhancer function.

AAAI Conference 2026 Conference Paper

Fair Diffusion Auctions

  • Zixin Gu
  • Yaoxin Ge
  • Yao Zhang
  • Dengji Zhao

Diffusion auction design is a new trend in mechanism design which extends the original incentive compatibility property to include buyers' private connection report. Reporting connections is equivalent to inviting their neighbors to join the auction in practice. Then, the social welfare is collectively accumulated by all participants: reporting high valuations or inviting high-valuation neighbors. Hence, we can measure each participant's contribution by the marginal social welfare increase due to her participation. Therefore, in this paper, we introduce a new property called Shapley fairness to capture participants' social welfare contribution and use it as a benchmark to guide our auction design for a fairer utility allocation. Not surprisingly, none of the existing diffusion auctions has ever approximated the fairness, because Shapley fairness depends on each buyer's own valuation and this dependence can easily violate incentive compatibility. Thus, we combat this challenge by proposing a new diffusion auction called Permutation Diffusion Auction (PDA) for selling k homogeneous items, which is the first diffusion auction satisfying 1/(k+1)-Shapley fairness, incentive compatibility and individual rationality. Moreover, PDA can be extended to the general combinatorial auction setting where the literature did not discover meaningful diffusion auctions yet.

AAAI Conference 2026 Conference Paper

Fair Incentives for Early Arrival in 0-1 Cooperative Games

  • Yaoxin Ge
  • Yao Zhang
  • Dengji Zhao

Incentives for early arrival (I4EA) was recently proposed for studying online cooperative games. In an online cooperative game, players arrive in an unknown order, and the value increase after each player arrived should be distributed immediately among all the arrived players. Although there is only one arriving order in the game, we also hope that the value distribution is equal to their Shapley value in expectation. To achieve these goals, the early solutions ignored the fairness in each single arriving order. More specifically, an important player may receive nothing in a game, which seems unfair in reality. To combat this, we propose refined fairness in this paper and design new solutions in 0-1 value games. Specifically, we compute the distance of the distribution in each order to the Shapley value and aim to minimize it. We propose a new mechanism called Egalitarian Value-Sharing (EVS) to do so. We also show that the mechanism can maximize the egalitarian welfare among all the players who made contributions.

AAMAS Conference 2026 Conference Paper

Feature-based Uncertainty Model for School Choice

  • Yao Zhang
  • Makoto Yokoo

In this work, we consider a school choice scenario where a student does not exactly know which college is better for her. Although it is hard for a student to obtain an exact preference, she can usually compare specific features of colleges, such as reputation, location, andcampusfacilities. Motivatedbythis, weproposeafeature-based uncertainty model for school choice where a student’s preference is based on a linear combination of her utilities over different features, and the coefficients of the combination are treated as random variables. Our main goal is to achieve a higher probability of stability (ProS) and incentive compatibility (IC) for students. Unfortunately, thesetwogoalsareincompatibleingeneral. Weshowthatastudentproposing deferred acceptance (DA) that prioritizes colleges with higher expected ranking can achieve a worst-case approximation ratio of (1/𝑛)𝑛 on ProS, while a DA with a carefully defined iterated comparison vector can guarantee the strongest achievable form of IC. Finally, we provide additional results for some specific restrictions on the model.

AIJ Journal 2026 Journal Article

Incentives for early arrival in online cooperative games

  • Dengji Zhao
  • Yaoxin Ge
  • Yao Zhang
  • Zhihao Gavin Tang
  • Hu Fu
  • Pinyan Lu

• We formalize a new concept of Incentives for Early Arrival (I4EA) in online cooperative games where players join sequentially. • We propose the very first mechanism called Rewarding First Critical players (RFC) to satisfy I4EA in 0–1 valued monotone games. • We propose an online decomposition of monotone games into 0–1 valued games, generalizing RFC while preserving properties. • The conference version won the best paper award at AAMAS 2024. This extended version adds an efficient algorithm for RFC and expands future directions. We study cooperative games where players join sequentially, and the value generated by those who have joined at any point must be irrevocably divided among these players. We introduce two desiderata for the value division mechanism: that the players should have incentives to join as early as possible, and that the division should be considered fair. For the latter, we require that each player’s expected share in the mechanism should equal her Shapley value if the players’ arrival order is uniformly at random. When the value generation function is submodular, allocating the marginal value to the player satisfies these properties. This is no longer true for more general functions. Our main technical contribution is a complete characterization of 0–1 value games for which desired mechanisms exist. We show that a natural mechanism, Rewarding First Critical Player (RFC), is complete, in that a 0–1 value function admits a mechanism with the properties above if and only if RFC satisfies them; we analytically characterize all such value functions. Moreover, we give an algorithm that decomposes, in an online fashion, any value function into 0–1 value functions, on each of which RFC can be run. In this way, we design an extension of RFC for general monotone games, and the properties are proved to be maintained.

AAAI Conference 2026 Conference Paper

On the Evaluation of Capability Estimation Methods for Large Language Models

  • Qiang Hu
  • Jin Wen
  • Yao Zhang
  • Maxime Cordy
  • Yongqiang Lyu

The emergence of large language models (LLMs) marks a transformative era in artificial intelligence~(AI). However, systematically evaluating the capability of LLMs is challenging due to the necessity of a large number of labeled test data. To tackle this problem, in the conventional AI field, AutoEval has been proposed to estimate the capability of AI models without data labeling effort. Unfortunately, even though multiple AutoEval methods have been proposed, most are constructed for classification tasks and evaluated only on image datasets. As a result, their effectiveness for LLMs is unclear, as LLMs often target generation tasks. In this work, we introduce the first AutoEval benchmark specifically designed to estimate the capability of LLMs using unlabeled test data, AEBench. Besides existing AutoEval methods, AEBench also supports our designed method, which utilizes the correlation between data uncertainty and model ability for the capability estimation. In total, AEBench covers 12 AutoEval methods and 120 method combinations. Based on AEBench, we conducted a comprehensive study to explore the usefulness of AutoEval on LLMs. Experimental results on 10 datasets demonstrated that our designed uncertainty features-based methods perform the best in achieving the lowest estimation errors.

EAAI Journal 2026 Journal Article

Progressive category-aware anti-distillation

  • Yao Zhang
  • Yang Li
  • Zhisong Pan

The widespread use of knowledge distillation has intensified the risk of model theft, revealing the inadequacy of traditional protection techniques against such threats. Anti-distillation has emerged as a promising defense paradigm by disrupting knowledge transfer to prevent unauthorized extraction of dark knowledge while preserving the teacher model’s performance. However, existing anti-distillation methods largely overlook the pivotal role of inter-class relationships in the distillation process. To address this limitation, we propose a progressive category-aware anti-distillation method. Our approach first constructs a relationship matrix between class prototypes to accurately model inter-class relationships, then reconstructs the output distribution to eliminate inter-class correlation information in output. To enhance stability and maintain distributional symmetry, we replace the standard Kullback–Leibler divergence with a symmetric Jensen–Shannon divergence. Moreover, we implement a curriculum learning mechanism to progressively adjust the intensity of inter-class correlation information removal. Extensive experiments on Cifar-100 and ImageNet demonstrate that our approach consistently surpasses existing anti-distillation methods, achieving strong robustness across various architectures — including Convolutional Neural Networks and Transformers — and different distillation settings such as logits-based, feature-based, and data-free distillation, with less than 2. 2% degradation in teacher performance.

AAAI Conference 2026 Conference Paper

T-SKM-Net: Trainable Neural Network Framework for Linear Constraint Satisfaction via Sampling Kaczmarz-Motzkin Method

  • Haoyu Zhu
  • Yao Zhang
  • Jiashen Ren
  • Qingchun Hou

Neural network constraint satisfaction is crucial for safety-critical applications such as power system optimization, robotic path planning, and autonomous driving. However, existing constraint satisfaction methods face efficiency-applicability trade-offs, with hard constraint methods suffering from either high computational complexity or restrictive assumptions on constraint structures. The Sampling Kaczmarz-Motzkin (SKM) method is a randomized iterative algorithm for solving large-scale linear inequality systems with favorable convergence properties, but its argmax operations introduce non-differentiability, posing challenges for neural network applications. This work proposes the Trainable Sampling Kaczmarz-Motzkin Network (T-SKM-Net) framework and, for the first time, systematically integrates SKM-type methods into neural network constraint satisfaction. The framework transforms mixed constraint problems into pure inequality problems through null space transformation, employs SKM for iterative solving, and maps solutions back to the original constraint space, efficiently handling both equality and inequality constraints. We provide theoretical proof of post-processing effectiveness in expectation and end-to-end trainability guarantees based on unbiased gradient estimators, demonstrating that despite non-differentiable operations, the framework supports standard backpropagation. On the DCOPF case118 benchmark, our method achieves 4.27ms/item GPU serial forward inference with 0.0025% max optimality gap with post-processing mode and 5.25ms/item with 0.0008% max optimality gap with joint training mode, delivering over 25× speedup compared to the pandapower solver while maintaining zero constraint violations under given tolerance.

AAAI Conference 2026 Conference Paper

Task-Specific Distance Correlation Matching for Few-Shot Action Recognition

  • Fei Long
  • Yao Zhang
  • Jiaming Lv
  • Jiangtao Xie
  • Peihua Li

Few-shot action recognition (FSAR) has recently made notable progress through set matching and efficient adaptation of large-scale pre-trained models. However, two key limitations persist. First, existing set matching metrics typically rely on cosine similarity to measure inter-frame linear dependencies and then perform matching with only instance-level information, thus failing to capture more complex patterns such as nonlinear relationships and overlooking task-specific cues. Second, for efficient adaptation of CLIP to FSAR, recent work performing fine-tuning via skip-fusion layers (which we refer to as side layers) has significantly reduced memory cost. However, the newly introduced side layers are often difficult to optimize under limited data conditions. To address these limitations, we propose TS-FSAR, a framework comprising three components: (1) a visual Ladder Side Network (LSN) for efficient CLIP fine-tuning; (2) a metric called Task-Specific Distance Correlation Matching (TS-DCM), which uses alpha-distance correlation to model both linear and nonlinear inter-frame dependencies and leverages a task prototype to enable task-specific matching; and (3) a Guiding LSN with Adapted CLIP (GLAC) module, which regularizes LSN using the adapted frozen CLIP to improve training for better α-distance correlation estimation under limited supervision. Extensive experiments on five widely-used benchmarks demonstrate that our TS-FSAR yields superior performance compared to prior state-of-the-arts.

AAAI Conference 2026 Conference Paper

VIL2C: Value-of-Information Aware Low-Latency Communication for Multi-Agent Reinforcement Learning

  • Qian Zhang
  • Zhuo Sun
  • Yao Zhang
  • Zhiwen Yu
  • Bin Guo
  • Jun Zhang

Inter-agent communication serves as an effective mechanism for enhancing performance in collaborative multi-agent reinforcement learning (MARL) systems. However, the inherent communication latency in practical systems induces both action decision delays and outdated information sharing, impeding MARL performance gains, particularly in time-critical applications like autonomous driving. In this work, we propose a Value-of-Information aware Low-latency Communication (VIL2C) scheme that proactively adjusts the latency distribution to mitigate its effects in MARL systems. Specifically, we define a Value of Information (VoI) metric to quantify the importance of delayed messages on the recipient agent's decision. We then design a VoI aware resource allocation method that dynamically prioritizes message transmission based on each delayed message's importance. Moreover, we propose a progressive message reception mechanism to adaptively adjust the reception duration based on received messages. We derive the optimized VoI aware resource allocation and theoretically prove the performance advantage of the proposed VIL2C scheme. Extensive experiments demonstrate that VIL2C outperforms existing approaches under various communication conditions. These gains are attributed to the low-latency transmission of high-VoI messages via resource allocation and the elimination of unnecessary waiting periods via adaptive reception duration.

AAAI Conference 2025 Conference Paper

DoGA: Enhancing Grounded Object Detection via Grouped Pre-Training with Attributes

  • Yang Liu
  • Feng Hou
  • Yunjie Peng
  • Gangjian Zhang
  • Yao Zhang
  • Dong Xie
  • Peng Wang
  • Yang Zhang

Recent advances in vision-language pre-training have significantly enhanced the model capabilities on grounded object detection. However, these studies often pre-train with coarse-grained text prompts, such as plain category names and brief grounded phrases. This limitation curtails the model's capacity for fine-grained linguistic comprehension and leads to a significant decline in performance when faced with detailed descriptions or contextual information. To tackle these problems, we develop DoGA: Detect objects with Grouped Attributes, which employs commonly apparent attributes to bridge different granular semantics and uses specific attributes to identify the object discrepancy. Our DoGA incorporates three principle components: 1) Generation of attribute-based prompts, consisting of linguistic definitions enriched with common-sense visible attributes and hard negative notations deriving from the image-specific attribute features; 2) Paralleled entity fusion and optimization, designed to manage long attribute-based descriptions and negative concepts efficiently; and 3) Prompt-wise grouped training to accommodate model to perform many-to-many assignments, facilitating simultaneous training and inferring with multiple attribute-based synonyms. Extensive experiments demonstrate that training with synonymous attribute-based prompts allows DoGA to generalize multi-granular prompts and surpass previous state-of-the-art approaches, yielding 50.2 on the COCO and 38.0 on the LVIS benchmarks under the zero-short setting. We will make our code publicly available upon acceptance.

EAAI Journal 2025 Journal Article

Enhanced multi-modal emotion recognition using the feature level fusion

  • Aziguli Wulamu
  • Yuheng Wu
  • Xin Liu
  • Yao Zhang
  • Jinghan Xu
  • Yang Zhang

Multi-modal human emotion recognition is a complex process of synthesizing information from various modalities to calculate emotion states. This field faces several challenges: (1) Acoustic is an essential component of emotion expression, but it often underperforms compared to visual and text in emotion recognition. (2) Capturing the feature interaction among different modalities is usually complex. (3) Processing high-definition videos can significantly reduce the efficiency of visual analysis. In this study, we presented a learning architecture designed to recognize human emotions effectively. For the first challenge, we implemented a multi-level acoustic encoder (MLAE) that enhances the extraction of acoustic information to improve the acoustic contribution in multi-modal emotion recognition. Facing the second challenge, we introduced the cross-attention block module, which adeptly captures the inter-modal interactions. To address the third challenge, we adopted the re-parameterized visual geometry group network (RepVGG) as the visual feature encoder, employing its multi-branch learning and single-branch reasoning structure to maintain high reasoning efficiency. Our model has demonstrated the state-of-the-art performance of the interactive emotional dyadic motion capture (IEMOCAP) dataset and the multi-modal opinion sentiment and emotion intensity of the Carnegie Mellon University (CMU-MOSEI) dataset.

IJCAI Conference 2025 Conference Paper

Incentives for Early Arrival in Cooperative Games (Extended Abstract)

  • Yaoxin Ge
  • Yao Zhang
  • Dengji Zhao
  • Zhihao Gavin Tang
  • Hu Fu
  • Pinyan Lu

We study cooperative games where players join sequentially, and the value generated by those who have joined at any point must be irrevocably divided among these players. We introduce two desiderata for the value division mechanism: that the players should have incentives to join as early as possible, and that the division should be considered fair. For the latter, we require that each player's expected share in the mechanism should equal her Shapley value if the players' arrival order is uniformly at random. When the value generation function is submodular, allocating the marginal value to the player satisfies these properties. This is no longer true for more general functions. Our main technical contribution is a complete characterization of 0-1 value games for which desired mechanisms exist. We show that a natural mechanism, Rewarding First Critical Player (RFC), is complete, in that a 0-1 value function admits a mechanism with the properties above if and only if RFC satisfies them; we analytically characterize all such value functions. Moreover, we give an algorithm that decomposes, in an online fashion, any value function into 0-1 value functions, on each of which RFC can be run. In this way, we design an extension of RFC for general monotone games, and the properties are proved to be maintained.

AAMAS Conference 2025 Conference Paper

Incentives for Early Arrival in Cost Sharing

  • Junyu Zhang
  • Yao Zhang
  • Yaoxin Ge
  • Dengji Zhao
  • Hu Fu
  • Zhihao Gavin Tang
  • Pinyan Lu

In cooperative games, we study how values created or costs incurred by a coalition are shared among the members within it, and the players may join the coalition in a online manner such as investors invest a startup. Recently, Ge et al. [10] proposed a new property called incentives for early arrival (I4EA) in such games, which says that the online allocation of values or costs should incentivize agents to join early in order to prevent mutual strategic waiting. Ideally, the allocation should also be fair, so that agents arriving in an order uniformly at random should expect to get/pay their Shapley values. Ge et al. [10] showed that not all monotone value functions admit such mechanisms in online value sharing games. In this work, we show a sharp contrast in online cost sharing games. We construct a mechanism with all the properties mentioned above, for every monotone cost function. To achieve this, we first solve 0-1 valued cost sharing games with a novel mechanism called Shapley-fair shuffle cost sharing mechanism (SFS-CS), and then extend SFS-CS to a family called generalized Shapley-fair shuffle cost sharing mechanisms (GSFS-CS). The critical technique we invented here is a mapping from one arrival order to another order so that we can directly apply marginal cost allocation on the shuffled orders to satisfy the properties. Finally, we solve general valued cost functions, by decomposing them into 0-1 valued functions in an online fashion.

AAAI Conference 2025 Conference Paper

Portcullis: A Scalable and Verifiable Privacy Gateway for Third-Party LLM Inference

  • Jiangou Zhan
  • Wenhui Zhang
  • Zheng Zhang
  • Huanran Xue
  • Yao Zhang
  • Ye Wu

Businesses using third-party LLMs face privacy risks from exposed prompts. This paper presents Portcullis, a privacy-preserving gateway that safeguards sensitive data while supporting efficient and accurate LLM responses. Portcullis functions as a mediator, anonymizing sensitive data in prompts through parallel substitution, securely interacting with LLMs, and accurately reconstructing responses. It ensures all data processing occurs within secure encrypted memory. The gateway is attested to ensure trustworthiness and protect user privacy. Portcullis is the first of its kind, offering a verifiable and scalable privacy gateway for third-party LLM inferences. We assess Portcullis's efficiency as a confidential container platform, demonstrating that its startup time scales linearly, ensuring scalability. Additionally, we evaluate its runtime performance using the PII and Enron Email Dataset. For masking and unmasking workloads, Portcullis outperforms Hide-and-Seek by 96x speed up, while maintaining equal or better false positive and false negative rates compared to existing solutions. On the Enron dataset, Portcullis achieves notably higher accuracy, surpassing Hide-and-Seek by over 0.1 for GPT-4o mini.

AAAI Conference 2025 Conference Paper

WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration

  • Yao Zhang
  • Zijian Ma
  • Yunpu Ma
  • Zhen Han
  • Yu Wu
  • Volker Tresp

LLM-based autonomous agents often fail to execute complex web tasks that require dynamic interaction, largely due to the inherent uncertainty and complexity of these environments. Existing LLM-based web agents typically rely on rigid, expert-designed policies specific to certain states and actions, lacking the flexibility and generalizability needed to adapt to unseen tasks. In contrast, humans excel by exploring unknowns, continuously adapting strategies based on new observations, and resolving ambiguities through exploration. To emulate human-like adaptability, web agents need strategic exploration and complex decision-making. Monte Carlo Tree Search (MCTS) is well-suited for this, but classical MCTS struggles with vast action spaces, unpredictable state transitions, and incomplete information in web tasks. In light of this, we develop WebPilot, a multi-agent system with a dual optimization strategy that improves MCTS to better handle complex web environments. Specifically, the Global Optimization phase involves generating a high-level plan by breaking down tasks into manageable subtasks, continuously refining this plan through reflective analysis of new observations and previous subtask attempts, thereby focusing the search process and mitigating challenges posed by vast action spaces in classical MCTS. Subsequently, the Local Optimization phase executes each subtask using a tailored MCTS designed for complex environments, effectively addressing uncertainties and managing incomplete information by iteratively refining decisions based on new observations. Experimental results on WebArena and MiniWoB++ demonstrate the effectiveness of WebPilot. Notably, on WebArena, WebPilot achieves SOTA performance with GPT-4, achieving a 93% relative increase in success rate over the concurrent tree search-based method. WebPilot advances autonomous agents, enabling more reliable decision-making in practical environments.

YNIMG Journal 2024 Journal Article

Development and validation of a perivascular space segmentation method in multi-center datasets

  • Peiyu Huang
  • Lingyun Liu
  • Yao Zhang
  • Siyan Zhong
  • Peng Liu
  • Hui Hong
  • Shuyue Wang
  • Linyun Xie

BACKGROUND: Perivascular spaces (PVS) visible on magnetic resonance imaging (MRI) are significant markers associated with various neurological diseases. Although quantitative analysis of PVS may enhance sensitivity and improve consistency across studies, the field lacks a universally validated method for analyzing images from multi-center studies. METHODS: We annotated PVS on multi-center 3D T1-weighted (T1w) images acquired using scanners from three major vendors (Siemens, General Electric, and Philips). A neural network, mcPVS-Net (multi-center PVS segmentation network), was trained using data from 40 subjects and then tested in a separate cohort of 15 subjects. We assessed segmentation accuracy against ground truth masks tailored for each scanner vendor. Additionally, we evaluated the agreement between segmented PVS volumes and visual scores for each scanner. We also explored correlations between PVS volumes and various clinical factors such as age, hypertension, and white matter hyperintensities (WMH) in a larger sample of 1020 subjects. Furthermore, mcPVS-Net was applied to a new dataset comprising both T1w and T2-weighted (T2w) images from a United Imaging scanner to investigate if PVS volumes could discriminate between subjects with differing visual scores. We also compared the mcPVS-Net with a previously published method that segments PVS from T1 images. RESULTS: In the test dataset, mcPVS-Net achieved a mean DICE coefficient of 0.80, with an average Precision of 0.81 and Recall of 0.79, indicating good specificity and sensitivity. The segmented PVS volumes were significantly associated with visual scores in both the basal ganglia (r = 0.541, p < 0.001) and white matter regions (r = 0.706, p < 0.001), and PVS volumes were significantly different among subjects with varying visual scores. Segmentation performance was consistent across different scanner vendors. PVS volumes exhibited significant associations with age, hypertension, and WMH. In the United Imaging scanner dataset, PVS volumes showed good associations with PVS visual scores evaluated on either T1w or T2w images. Compared to a previously published method, mcPVS-Net showed a higher accuracy and improved PVS segmentation in the basal ganglia region. CONCLUSION: The mcPVS-Net demonstrated good accuracy for segmenting PVS from 3D T1w images. It may serve as a useful tool for future PVS research.

AAAI Conference 2024 Conference Paper

FedDAT: An Approach for Foundation Model Finetuning in Multi-Modal Heterogeneous Federated Learning

  • Haokun Chen
  • Yao Zhang
  • Denis Krompass
  • Jindong Gu
  • Volker Tresp

Recently, foundation models have exhibited remarkable advancements in multi-modal learning. These models, equipped with millions (or billions) of parameters, typically require a substantial amount of data for finetuning. However, collecting and centralizing training data from diverse sectors becomes challenging due to distinct privacy regulations. Federated Learning (FL) emerges as a promising solution, enabling multiple clients to collaboratively train neural networks without centralizing their local data. To alleviate client computation burdens and communication overheads, previous works have adapted Parameter-efficient Finetuning (PEFT) methods for FL. Hereby, only a small fraction of the model parameters are optimized and communicated during federated communications. Nevertheless, most previous works have focused on a single modality and neglected one common phenomenon, i.e., the presence of data heterogeneity across the clients. Therefore, in this work, we propose a finetuning framework tailored to heterogeneous multi-modal FL, called Federated Dual-Aadapter Teacher (FedDAT). Specifically, our approach leverages a Dual-Adapter Teacher (DAT) to address data heterogeneity by regularizing the client local updates and applying Mutual Knowledge Distillation (MKD) for an efficient knowledge transfer. FedDAT is the first approach that enables an efficient distributed finetuning of foundation models for a variety of heterogeneous Vision-Language tasks. To demonstrate its effectiveness, we conduct extensive experiments on four multi-modality FL benchmarks with different types of data heterogeneity, where FedDAT substantially outperforms the existing centralized PEFT methods adapted for FL.

ICML Conference 2024 Conference Paper

GroupCover: A Secure, Efficient and Scalable Inference Framework for On-device Model Protection based on TEEs

  • Zheng Zhang
  • Na Wang 0003
  • Ziqi Zhang
  • Yao Zhang
  • Tianyi Zhang
  • Jianwei Liu 0001
  • Ye Wu

Due to the high cost of training DNN models, how to protect the intellectual property of DNN models, especially when the models are deployed to users’ devices, is becoming an important topic. One practical solution is to use Trusted Execution Environments (TEEs) and researchers have proposed various model obfuscation solutions to make full use of the high-security guarantee of TEEs and the high performance of collocated GPUs. In this paper, we first identify a common vulnerability, namely the fragility of randomness, that is shared by existing TEE-based model obfuscation solutions. This vulnerability benefits model-stealing attacks and allows the adversary to recover about 97% of the secret model. To improve the security of TEE-shielded DNN models, we further propose a new model obfuscation approach GroupCover, which uses sufficient randomization and mutual covering obfuscation to protect model weights. Experimental results demonstrate that GroupCover can achieve a comparable security level as the upper-bound (black-box protection), which is remarkably over 3x compared with existing solutions. Besides, GroupCover introduces 19% overhead and negligible accuracy loss compared to model unprotected scheme.

YNIMG Journal 2024 Journal Article

Higher intracranial arterial pulsatility is associated with presumed imaging markers of the glymphatic system: An explorative study

  • Linyun Xie
  • Yao Zhang
  • Hui Hong
  • Shan Xu
  • Lei Cui
  • Shuyue Wang
  • Jixuan Li
  • Lingyun Liu

BACKGROUND: Arterial pulsation has been suggested as a key driver of paravascular cerebrospinal fluid flow, which is the foundation of glymphatic clearance. However, whether intracranial arterial pulsatility is associated with glymphatic markers in humans has not yet been studied. METHODS: , and the presumed glymphatic markers, controlling for related covariates. RESULTS: in the ICA C2 (β, -0.239, p, 0.041) and C7 segments (β, -0.238, p, 0.037). CONCLUSIONS: Intracranial arterial pulsatility was associated with presumed neuroimaging markers of the glymphatic system, but the results were not consistent across different markers. Further studies are warranted to confirm these findings.

AAMAS Conference 2024 Conference Paper

Incentives for Early Arrival in Cooperative Games

  • Yaoxin Ge
  • Yao Zhang
  • Dengji Zhao
  • Zhihao Gavin Tang
  • Hu Fu
  • Pinyan Lu

We study cooperative games where players join sequentially, and the value generated by those who have joined at any point must be irrevocably divided among these players. We introduce two desiderata for the value division mechanism: that the players should have incentives to join as early as possible, and that the division should be considered fair. For the latter, we require that each player’s expected share in the mechanism should equal her Shapley value if the players’ arrival order is uniformly at random. When the value generation function is submodular, allocating the marginal value to the player satisfies these properties. This is no longer true for more general functions. Our main technical contribution is a complete characterization of 0-1 value games for which desired mechanisms exist. We show that a natural mechanism, Rewarding First Critical Player (RFC), is complete, in that a 0-1 value function admits a mechanism with the properties above if and only if RFC satisfies them; we analytically characterize all such value functions. Moreover, we give an algorithm that decomposes, in an online fashion, any value function into 0-1 value functions, on each of which RFC can be run. In this way, we design an extension of RFC for general monotone games, and the properties are proved to be maintained.

YNIMG Journal 2024 Journal Article

Neural correlates of working memory training: An fMRI meta-analysis

  • Yao Zhang
  • Junjun Fu
  • Xin Zhao

Working memory (WM) can be improved by cognitive training. Numerous studies examined neural mechanisms underlying WM training, although with differing conclusions. Therefore, we conducted a meta-analysis to examine the neural substrates underlying WM training in healthy adults. Findings from global analyses showed substantial neural changes in the frontoparietal and subcortical regions. Results from training dosage analyses of WM training showed that shorter WM training could produce neural changes in the frontoparietal regions, whereas longer WM training could produce changes in the subcortical regions (striatum, anterior cingulate cortex, and insula). WM training-induced neural changes were also moderated by the type of training task, with updating tasks inducing neural changes in more regions than maintenance tasks. Overall, these results indicate that the neural changes associated with WM training occur in the frontoparietal network and dopamine-related brain areas, extending previous meta-analyses on WM training and advancing our understanding of the neural underpinnings of WM training effects.

AAMAS Conference 2024 Conference Paper

Optimal Diffusion Auctions

  • Yao Zhang
  • Shanshan Zheng
  • Dengji Zhao

Diffusion auction design is a new trend in mechanism design for which the main goal is to incentivize existing buyers to invite their neighbors on a social network, to join an auction. With more buyers, a diffusion auction will be able to receive higher revenue. Existing studies have proposed many diffusion auctions to attract more buyers, but the seller’s revenue is not optimized. In this study, we investigate what optimal revenue the seller can achieve by attracting more buyers. Different from the traditional setting, the revenue can be achieved highly relies on the structure of the network. We propose a class of mechanisms, where for any given structure, an optimal diffusion mechanism can be found. Moreover, we show that an optimal mechanism that handles all structures does not exist. Therefore, we also propose mechanisms that have bounded approximations of the optimal revenue in all structures.

AAMAS Conference 2023 Conference Paper

A Redistribution Framework for Diffusion Auctions

  • Sizhe Gu
  • Yao Zhang
  • Yida Zhao
  • Dengji Zhao

Redistribution mechanism design aims to redistribute the revenue collected by a truthful auction back to its participants without affecting the truthfulness. We study redistribution mechanisms for diffusion auctions, which is a new trend in mechanism design [19]. The key property of a diffusion auction is that the existing participants are incentivized to invite new participants to join the auctions. Hence, when we design redistributions, we also need to maintain this incentive. Existing redistribution mechanisms in the traditional setting are targeted at modifying the payment design of a truthful mechanism, such as the Vickrey auction. In this paper, we do not focus on one specific mechanism. Instead, we propose a general framework to redistribute the revenue back for all truthful diffusion auctions for selling a single item. The framework treats the original truthful diffusion auction as a black box, and it does not affect its truthfulness. The framework can also distribute back almost all the revenue.

JBHI Journal 2023 Journal Article

Compressibility Analysis of Functional Near-Infrared Spectroscopy Signals in Children With Attention-Deficit/Hyperactivity Disorder

  • Yue Gu
  • Shuo Miao
  • Yao Zhang
  • Jian Yang
  • Xiaoli Li

Functional near-infrared spectroscopy (fNIRS) as an emerging optical neuroimaging technique has attracted the interest and attention of many investigators. With the growth of fNIRS data volume, effective data compression methods are urgent. Compressive sensing (CS) has been demonstrated a promising tool to deal with biomedical data. However, whether the compressibility of fNIRS data can discriminate different brain states is unclear. In this study, the fNIRS signals from fifteen attention-deficit/hyperactivity disorder (ADHD) children and fifteen typically developing (TD) children were recorded during an N-back task and a Go/NoGo task respectively. A block sparse Bayesian learning-based CS method was used to reconstruct the compressed fNIRS data. To assess the performance of the CS method, we adopted two metrics, structural similarity index (SSIM) and mean squared error (MSE), both of them effective in evaluating the compressibility of fNIRS data. Then, the two metrics were analyzed to discriminate the brain states of ADHD children and TD children during the two tasks using the multivariate pattern analysis (MVPA) method. As indicated by the results, the CS method could reconstruct the compressed fNIRS data with a high reconstruction quality at different compression ratio ( $\text{SSIM} > \text{0. 988}$ and $\text{MSE} < \text{1. 2} \times \text{10}^{-4}$ ). Furthermore, the MVPA method could distinguish different brain states with high accuracy, and identify that the prefrontal cortex is a key brain region for distinguishing ADHD vs. TD or N-back vs. Go/NoGo. These findings indicated that CS is very promising for the storage and transmission of massive fNIRS data, and the compressibility of fNIRS data is a potential biomarker for the diagnosis of ADHD.

AAMAS Conference 2023 Conference Paper

Distributed Mechanism Design in Social Networks

  • Haoxin Liu
  • Yao Zhang
  • Dengji Zhao

Designing auctions to incentivize buyers to invite new buyers via their social connections is a new trend in mechanism design [18]. The challenge is that buyers are competitors and we need to design proper incentives for them to invite each other. For selling a single item, many interesting mechanisms have been proposed. However, all the mechanisms require the seller or a third party to be trustworthy to execute the mechanisms. In addition, the owner of the mechanism will know all the connections of the network after the execution, which poses a potential privacy issue. Hence, distributed mechanisms to avoid the privacy issue are more appealing in practice. Therefore, in this paper, we propose the first distributed mechanism in social networks without revealing buyers’ private connections to anyone, and it achieves complete decentralization that does not rely on any trustworthy third party. Moreover, the centralized reduction of our mechanism also offers a novel way to compute players’ contributions compared to the existing solutions.

IJCAI Conference 2023 Conference Paper

Incentive-Compatible Selection for One or Two Influentials

  • Yuxin Zhao
  • Yao Zhang
  • Dengji Zhao

Selecting influentials in networks against strategic manipulations has attracted many researchers' attention and it also has many practical applications. Here, we aim to select one or two influentials in terms of progeny (the influential power) and prevent agents from manipulating their edges (incentive compatibility). The existing studies mostly focused on selecting a single influential for this setting. Zhang et al. [2021] studied the problem of selecting one agent and proved an upper bound of 1/(1+ln2) to approximate the optimal selection. In this paper, we first design a mechanism to actually reach the bound. Then, we move this forward to choosing two agents and propose a mechanism to achieve an approximation ratio of (3+ln2)/(4(1+ln2)) (approx. 0. 54).

IJCAI Conference 2023 Conference Paper

Task Allocation on Networks with Execution Uncertainty (Extended Abstract)∗

  • Yao Zhang
  • Xiuzhen Zhang
  • Dengji Zhao

We study a single task allocation problem where each worker connects to some other workers to form a network and the task requester only connects to some of the workers. The goal is to design an allocation mechanism such that each worker is incentivized to invite her neighbours to join the allocation, although they are competing for the task. Moreover, the performance of each worker is uncertain, which is modelled as the quality level of her task execution. The literature has proposed solutions to tackle the uncertainty problem by paying them after verifying their execution. Here, we extend the problem to the network setting. We propose a new mechanism that guarantees that inviting more workers and reporting/performing according to her true ability is a dominant strategy for each worker. We believe that the new solution can be widely applied in the digital economy powered by social connections such as crowdsourcing.

TIST Journal 2022 Journal Article

CAFE and SOUP: Toward Adaptive VDI Workload Prediction

  • Yao Zhang
  • Wenping Fan
  • Qichen Hao
  • Xinya Wu
  • Min-Ling Zhang

For Virtual Desktop Infrastructure (VDI) system, effective resource management is rather important where turning off spare virtual machines would help save running cost while maintaining sufficient virtual machines is essential to secure satisfactory user experience. Current VDI resource management strategy works in a passive manner by either reactively driving available capacity based on user demands or following manually configured schedules, which may lead to unnecessary running costs or unsatisfactory user experience. In this article, we propose a first attempt toward proactive VDI resource management, where two adaptive learning approaches for VDI workload prediction are proposed by learning from multi-grained historical features. For non-persistent desktop pool, based on the aggregation session count of pool-sharing users, the CAFE approach induces a pool-level workload predictive model by utilizing coarse-to-fine historical features extracted from aggregation workload data. For persistent desktop pool, based on the session connection status of individual users within the same pool, the SOUP approach induces user-level workload predictive model by incorporating encoded multi-grained features extracted from the logon behavior of individual users into an aggregation pool-level model. Extensive experiments on datasets of real VDI customers and electricity load evidently verify the effectiveness of the proposed adaptive approaches for VDI workload prediction as well as other workload prediction tasks.

IJCAI Conference 2022 Conference Paper

Diffusion Incentives in Cooperative Games

  • Yao Zhang

We study a cooperative game setting where we want to gather more players through their social connections. Social connections can be modeled as a graph, and initially, only a subset of the players are in the game. We want to introduce diffusion incentives in such a cooperative game, i. e. , incentivize the players to use their connections to invite more players to join the game. Our goal cannot be achieved by existing classical solutions, such as the Shapley value. Hence, to combat this problem, we have already proposed a solution called weighted permission Shapley value. Under this solution, for each player, inviting all of her neighbors is a dominant strategy in all monotone games. As one special application of the diffusion cooperative game, we also considered the diffusion incentives in query networks and the weighted permission Shapley value successfully characterizes the solution to the query network. Furthermore, we also characterize a Sybil-proof solution to the query network called the double geometric mechanism.

AAMAS Conference 2022 Conference Paper

Incentives to Invite Others to Form Larger Coalitions

  • Yao Zhang
  • Dengji Zhao

We study a cooperative game setting where players form a network and each player only knows the existence of the players to whom she connects. Initially, only a subset of the players are in the game. Our goal is to design a reward distribution mechanism to incentivize the players to use their connections to invite more players to join the game. We show that the existing solutions such as the Shapley value cannot achieve this. Hence, to combat this problem, we propose a solution called weighted permission Shapley value (inspired by permission structure and the weighted Shapley value). Under this solution, for each player, inviting all her neighbors is a dominant strategy in all monotone games. We further prove that the solution is unique for tree networks. Our solution offers the very first attempt to incentivize the players to invite others to form a larger coalition in cooperative games.

YNICL Journal 2022 Journal Article

Reduced coupling between the global blood-oxygen-level-dependent signal and cerebrospinal fluid inflow is associated with the severity of small vessel disease

  • Yao Zhang
  • Ruiting Zhang
  • Shuyue Wang
  • Hui Hong
  • Yeerfan Jiaerken
  • Kaicheng Li
  • Qingze Zeng
  • Xiao Luo

BACKGROUND: Small vessel disease (SVD) is highly prevalent in the elderly and associated with an increased risk of dementia and stroke. SVD may have disturbed cerebrospinal fluid (CSF) flow, which can compromise waste clearance and accelerate disease progression. METHODS: We retrospectively included 146 SVD patients from a prospectively collected dataset, with one- or two-year follow-up data in 61 patients. The coupling strength between the global blood-oxygen-level-dependent (gBOLD) signal and CSF inflow was used to reflect CSF dynamics. We performed regression analyses to investigate the association between the gBOLD-CSF coupling index and the severity of SVD and vascular risk factors. Longitudinal analysis was carried out to investigate causal relationships. RESULTS: Patients with severe SVD had significantly decreased gBOLD-CSF coupling (β = -0.180, p = 0.032). Dilation of perivascular spaces in the basal ganglia area (β = -0.172, p = 0.033) and diabetes (β = -0.204, p = 0.014) were associated with reduced gBOLD-CSF coupling. In longitudinal analyses, diabetes was associated with faster decline in gBOLD-CSF coupling (β = 0.20, p = 0.039), while perivascular space (PVS) dilation in the centrum semiovale showed a opposite relationship (β = -0.20, p = 0.041). The gBOLD-CSF coupling could not predict SVD progression. CONCLUSION: Altered CSF flow is associated with the severity of SVD.

JBHI Journal 2022 Journal Article

Voice Biomarkers of Recovery From Acute Respiratory Illness

  • Brian Tracey
  • Shyamal Patel
  • Yao Zhang
  • Kara Chappie
  • Dmitri Volfson
  • Federico Parisi
  • Catherine Adans-Dester
  • Francesco Bertacchi

Voice analysis is an emerging technology which has the potential to provide low-cost, at-home monitoring of symptoms associated with a variety of health conditions. While voice has received significant attention for monitoring neurological disease, few studies have focused on voice changes related to flu-like symptoms. Herein, we investigate the relationship between changes in acoustic features of voice and self-reported symptoms during recovery from a flu-like illness in a cohort of 29 subjects. Acoustic features were automatically extracted from “sick” and “well” visit data collected in the laboratory setting, and feature down-selection was used to identify those that change significantly between visits. The selected acoustic features were extracted from at-home data and used to construct a combined distance metric that correlated with self-reported symptoms (0. 63 rank correlation). Changes in self-reported symptoms corresponding to 10% of the ordinal scale used in the study were detected with an area under the curve of 0. 72. The results show that acoustic features derived from voice recordings may provide an objective measure for diagnosing and monitoring symptoms of respiratory illnesses.

AAAI Conference 2021 Conference Paper

Argument Mining Driven Analysis of Peer-Reviews

  • Michael Fromm
  • Evgeniy Faerman
  • Max Berrendorf
  • Siddharth Bhargava
  • Ruoxia Qi
  • Yao Zhang
  • Lukas Dennert
  • Sophia Selle

Peer reviewing is a central process in modern research and essential for ensuring high quality and reliability of published work. At the same time, it is a time-consuming process and increasing interest in emerging fields often results in a high review workload, especially for senior researchers in this area. How to cope with this problem is an open question and it is vividly discussed across all major conferences. In this work, we propose an Argument Mining based approach for the assistance of editors, meta-reviewers, and reviewers. We demonstrate that the decision process in the field of scientific publications is driven by arguments and automatic argument identification is helpful in various use-cases. One of our findings is that arguments used in the peer-review process differ from arguments in other domains making the transfer of pretrained models difficult. Therefore, we provide the community with a new peer-review dataset from different computer science conferences with annotated arguments. In our extensive empirical evaluation, we show that Argument Mining can be used to efficiently extract the most relevant parts from reviews, which are paramount for the publication decision. The process remains interpretable since the extracted arguments can be highlighted in a review without detaching them from their context.

IJCAI Conference 2021 Conference Paper

BAMBOO: A Multi-instance Multi-label Approach Towards VDI User Logon Behavior Modeling

  • Wenping Fan
  • Yao Zhang
  • Qichen Hao
  • Xinya Wu
  • Min-Ling Zhang

Different to traditional on-premise VDI, the virtual desktops in DaaS (Desktop as a Service) are hosted in public cloud where virtual machines are charged based on usage. Accordingly, an adaptive power management system which can turn off spare virtual machines without sacrificing end user experience is of significant customer value as it can greatly help reduce the running cost. Generally, logon behavior modeling for VDI users serves as the key enabling-technique to fulfill intelligent power management. Prior attempts work by modeling logon behavior in a user-dependent manner with tailored single-instance feature representation, where the strong relationships among pool-sharing VDI users are ignored in the modeling framework. In this paper, a novel formulation towards VDI user logon behavior modeling is proposed by employing the multi-instance multi-label (MIML) techniques. Specifically, each user is grouped with supporting users whose behaviors are jointly modeled in the feature space with multi-instance representation as well as in the output space with multi-label prediction. The resulting MIML formulation is optimized by adapting the popular MIML boosting procedure via balanced error-rate minimization. Experimental studies on real VDI customers' data clearly validate the effectiveness of the proposed MIML-based approach against state-of-the-art VDI user logon behavior modeling techniques.

AAAI Conference 2021 Conference Paper

Generalized Relation Learning with Semantic Correlation Awareness for Link Prediction

  • Yao Zhang
  • Xu Zhang
  • Jun Wang
  • Hongru Liang
  • Wenqiang Lei
  • Zhe Sun
  • Adam Jatowt
  • Zhenglu Yang

Developing link prediction models to automatically complete knowledge graphs has recently been the focus of significant research interest. The current methods for the link prediction task have two natural problems: 1) the relation distributions in KGs are usually unbalanced, and 2) there are many unseen relations that occur in practical situations. These two problems limit the training effectiveness and practical applications of the existing link prediction models. We advocate a holistic understanding of KGs and we propose in this work a unified Generalized Relation Learning framework GRL to address the above two problems, which can be plugged into existing link prediction models. GRL conducts a generalized relation learning, which is aware of semantic correlations between relations that serve as a bridge to connect semantically similar relations. After training with GRL, the closeness of semantically similar relations in vector space and the discrimination of dissimilar relations are improved. We perform comprehensive experiments on six benchmarks to demonstrate the superior capability of GRL in the link prediction task. In particular, GRL is found to enhance the existing link prediction models making them insensitive to unbalanced relation distributions and capable of learning unseen relations.

NeurIPS Conference 2021 Conference Paper

MIRACLE: Causally-Aware Imputation via Learning Missing Data Mechanisms

  • Trent Kyono
  • Yao Zhang
  • Alexis Bellot
  • Mihaela van der Schaar

Missing data is an important problem in machine learning practice. Starting from the premise that imputation methods should preserve the causal structure of the data, we develop a regularization scheme that encourages any baseline imputation method to be causally consistent with the underlying data generating mechanism. Our proposal is a causally-aware imputation algorithm (MIRACLE). MIRACLE iteratively refines the imputation of a baseline by simultaneously modeling the missingness generating mechanism, encouraging imputation to be consistent with the causal structure of the data. We conduct extensive experiments on synthetic and a variety of publicly available datasets to show that MIRACLE is able to consistently improve imputation over a variety of benchmark methods across all three missingness scenarios: at random, completely at random, and not at random.

NeurIPS Conference 2021 Conference Paper

Reinforcement Learning Enhanced Explainer for Graph Neural Networks

  • Caihua Shan
  • Yifei Shen
  • Yao Zhang
  • Xiang Li
  • Dongsheng Li

Graph neural networks (GNNs) have recently emerged as revolutionary technologies for machine learning tasks on graphs. In GNNs, the graph structure is generally incorporated with node representation via the message passing scheme, making the explanation much more challenging. Given a trained GNN model, a GNN explainer aims to identify a most influential subgraph to interpret the prediction of an instance (e. g. , a node or a graph), which is essentially a combinatorial optimization problem over graph. The existing works solve this problem by continuous relaxation or search-based heuristics. But they suffer from key issues such as violation of message passing and hand-crafted heuristics, leading to inferior interpretability. To address these issues, we propose a RL-enhanced GNN explainer, RG-Explainer, which consists of three main components: starting point selection, iterative graph generation and stopping criteria learning. RG-Explainer could construct a connected explanatory subgraph by sequentially adding nodes from the boundary of the current generated graph, which is consistent with the message passing scheme. Further, we design an effective seed locator to select the starting point, and learn stopping criteria to generate superior explanations. Extensive experiments on both synthetic and real datasets show that RG-Explainer outperforms state-of-the-art GNN explainers. Moreover, RG-Explainer can be applied in the inductive setting, demonstrating its better generalization ability.

NeurIPS Conference 2021 Conference Paper

SyncTwin: Treatment Effect Estimation with Longitudinal Outcomes

  • Zhaozhi Qian
  • Yao Zhang
  • Ioana Bica
  • Angela Wood
  • Mihaela van der Schaar

Most of the medical observational studies estimate the causal treatment effects using electronic health records (EHR), where a patient's covariates and outcomes are both observed longitudinally. However, previous methods focus only on adjusting for the covariates while neglecting the temporal structure in the outcomes. To bridge the gap, this paper develops a new method, SyncTwin, that learns a patient-specific time-constant representation from the pre-treatment observations. SyncTwin issues counterfactual prediction of a target patient by constructing a synthetic twin that closely matches the target in representation. The reliability of the estimated treatment effect can be assessed by comparing the observed and synthetic pre-treatment outcomes. The medical experts can interpret the estimate by examining the most important contributing individuals to the synthetic twin. In the real-data experiment, SyncTwin successfully reproduced the findings of a randomized controlled clinical trial using observational data, which demonstrates its usability in the complex real-world EHR.

NeurIPS Conference 2020 Conference Paper

CASTLE: Regularization via Auxiliary Causal Graph Discovery

  • Trent Kyono
  • Yao Zhang
  • Mihaela van der Schaar

Regularization improves generalization of supervised models to out-of-sample data. Prior works have shown that prediction in the causal direction (effect from cause) results in lower testing error than the anti-causal direction. However, existing regularization methods are agnostic of causality. We introduce Causal Structure Learning (CASTLE) regularization and propose to regularize a neural network by jointly learning the causal relationships between variables. CASTLE learns the causal directed acyclical graph (DAG) as an adjacency matrix embedded in the neural network's input layers, thereby facilitating the discovery of optimal predictors. Furthermore, CASTLE efficiently reconstructs only the features in the causal DAG that have a causal neighbor, whereas reconstruction-based regularizers suboptimally reconstruct all input features. We provide a theoretical generalization bound for our approach and conduct experiments on a plethora of synthetic and real publicly available datasets demonstrating that CASTLE consistently leads to better out-of-sample predictions as compared to other popular benchmark regularizers.

AAAI Conference 2020 Conference Paper

Fine-Grained Fashion Similarity Learning by Attribute-Specific Embedding Network

  • Zhe Ma
  • Jianfeng Dong
  • Zhongzi Long
  • Yao Zhang
  • Yuan He
  • Hui Xue
  • Shouling Ji

This paper strives to learn fine-grained fashion similarity. In this similarity paradigm, one should pay more attention to the similarity in terms of a specific design/attribute among fashion items, which has potential values in many fashion related applications such as fashion copyright protection. To this end, we propose an Attribute-Specific Embedding Network (ASEN) to jointly learn multiple attributespecific embeddings in an end-to-end manner, thus measure the fine-grained similarity in the corresponding space. With two attention modules, i. e. , Attribute-aware Spatial Attention and Attribute-aware Channel Attention, ASEN is able to locate the related regions and capture the essential patterns under the guidance of the specified attribute, thus make the learned attribute-specific embeddings better reflect the fine-grained similarity. Extensive experiments on four fashion-related datasets show the effectiveness of ASEN for fine-grained fashion similarity learning and its potential for fashion reranking. Code and data are available at https: //github. com/Maryeon/asen.

NeurIPS Conference 2020 Conference Paper

Gradient Regularized V-Learning for Dynamic Treatment Regimes

  • Yao Zhang
  • Mihaela van der Schaar

Deciding how to optimally treat a patient, including how to select treatments over time among the multiple available treatments, represents one of the most important issues that need to be addressed in medicine today. A dynamic treatment regime (DTR) is a sequence of treatment rules indicating how to individualize treatments for a patient based on the previously assigned treatments and the evolving covariate history. However, DTR evaluation and learning based on offline data remain challenging problems due to the bias introduced by time-varying confounders that affect treatment assignment over time; this may lead to suboptimal treatment rules being used in practice. In this paper, we introduce Gradient Regularized V-learning (GRV), a novel method for estimating the value function of a DTR. GRV regularizes the underlying outcome and propensity score models with respect to the optimality condition in semiparametric estimation theory. On the basis of this design, we construct estimators that are efficient and stable in finite samples regime. Using multiple simulation studies and one real-world medical dataset, we demonstrate that our method is superior in DTR evaluation and learning, thereby providing improved treatment options over time for patients.

NeurIPS Conference 2020 Conference Paper

Learning outside the Black-Box: The pursuit of interpretable models

  • Jonathan Crabbe
  • Yao Zhang
  • William Zame
  • Mihaela van der Schaar

Machine learning has proved its ability to produce accurate models -- but the deployment of these models outside the machine learning community has been hindered by the difficulties of interpreting these models. This paper proposes an algorithm that produces a continuous global interpretation of any given continuous black-box function. Our algorithm employs a variation of projection pursuit in which the ridge functions are chosen to be Meijer G-functions, rather than the usual polynomial splines. Because Meijer G-functions are differentiable in their parameters, we can "tune" the parameters of the representation by gradient descent; as a consequence, our algorithm is efficient. Using five familiar data sets from the UCI repository and two familiar machine learning algorithms, we demonstrate that our algorithm produces global interpretations that are both faithful (highly accurate) and parsimonious (involve a small number of terms). Our interpretations permit easy understanding of the relative importance of features and feature interactions. Our interpretation algorithm represents a leap forward from the previous state of the art.

NeurIPS Conference 2020 Conference Paper

Robust Recursive Partitioning for Heterogeneous Treatment Effects with Uncertainty Quantification

  • Hyun-Suk Lee
  • Yao Zhang
  • William Zame
  • Cong Shen
  • Jang-Won Lee
  • Mihaela van der Schaar

Subgroup analysis of treatment effects plays an important role in applications from medicine to public policy to recommender systems. It allows physicians (for example) to identify groups of patients for whom a given drug or treatment is likely to be effective and groups of patients for which it is not. Most of the current methods of subgroup analysis begin with a particular algorithm for estimating individualized treatment effects (ITE) and identify subgroups by maximizing the difference across subgroups of the average treatment effect in each subgroup. These approaches have several weaknesses: they rely on a particular algorithm for estimating ITE, they ignore (in)homogeneity within identified subgroups, and they do not produce good confidence estimates. This paper develops a new method for subgroup analysis, R2P, that addresses all these weaknesses. R2P uses an arbitrary, exogenously prescribed algorithm for estimating ITE and quantifies the uncertainty of the ITE estimation, using a construction that is more robust than other methods. Experiments using synthetic and semi-synthetic datasets (based on real data) demonstrate that R2P constructs partitions that are simultaneously more homogeneous within groups and more heterogeneous across groups than the partitions produced by other methods. Moreover, because R2P can employ any ITE estimator, it also produces much narrower confidence intervals with a prescribed coverage guarantee than other methods.

IJCAI Conference 2020 Conference Paper

Sybil-proof Answer Querying Mechanism

  • Yao Zhang
  • Xiuzhen Zhang
  • Dengji Zhao

We study a question answering problem on a social network, where a requester is seeking an answer from the agents on the network. The goal is to design reward mechanisms to incentivize the agents to propagate the requester's query to their neighbours if they don't have the answer. Existing mechanisms are vulnerable to Sybil-attacks, i. e. , an agent may get more reward by creating fake identities. Hence, we combat this problem by first proving some impossibility results to resolve Sybil-attacks and then characterizing a class of mechanisms which satisfy Sybil-proofness (prevents Sybil-attacks) as well as other desirable properties. Except for Sybil-proofness, we also consider cost minimization for the requester and agents' collusions.

NeurIPS Conference 2020 Conference Paper

VIME: Extending the Success of Self- and Semi-supervised Learning to Tabular Domain

  • Jinsung Yoon
  • Yao Zhang
  • James Jordon
  • Mihaela van der Schaar

Self- and semi-supervised learning frameworks have made significant progress in training machine learning models with limited labeled data in image and language domains. These methods heavily rely on the unique structure in the domain datasets (such as spatial relationships in images or semantic relationships in language). They are not adaptable to general tabular data which does not have the same explicit structure as image and language data. In this paper, we fill this gap by proposing novel self- and semi-supervised learning frameworks for tabular data, which we refer to collectively as VIME (Value Imputation and Mask Estimation). We create a novel pretext task of estimating mask vectors from corrupted tabular data in addition to the reconstruction pretext task for self-supervised learning. We also introduce a novel tabular data augmentation method for self- and semi-supervised learning frameworks. In experiments, we evaluate the proposed framework in multiple tabular datasets from various application domains, such as genomics and clinical data. VIME exceeds state-of-the-art performance in comparison to the existing baseline methods.

AAAI Conference 2019 Conference Paper

CAFE: Adaptive VDI Workload Prediction with Multi-Grained Features

  • Yao Zhang
  • Wen-Ping Fan
  • Xuan Wu
  • Hua Chen
  • Bin-Yang Li
  • Min-Ling Zhang

Virtual desktop infrastructure (VDI) is a virtualization technology that hosts desktop operating system on centralized server in a data center of private or public cloud. Effective resource management is of crucial importance for VDI customers, where maintaining sufficient virtual machines helps guarantee satisfactory user experience while turning off spare virtual machines helps save running cost. Generally, existing techniques work in passive manner by either driving available capacity reactively or configuring management schedules manually. In this paper, a novel proactive resource management approach is proposed which aims to predict VDI pool workload adaptively by utilizing CoArse to Fine historical dEscriptive (CAFE) features. Specifically, aggregate session count from pool end users serves as the basis for workload measurement and predictive model induction. Extensive experiments on real VDI customers data sets clearly validate the effectiveness of multi-grained features for VDI workload prediction. Furthermore, practical insights identified in our VDI data analytics are also discussed.

v2026.09.13