Arrow Research search

Author name cluster

Suyun Zhao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

5 papers
2 author rows

Possible papers

5

AAAI Conference 2025 Conference Paper

Personalized Clustering via Targeted Representation Learning

  • Xiwen Geng
  • Suyun Zhao
  • Yixin Yu
  • Borui Peng
  • Pan Du
  • Hong Chen
  • Cuiping Li
  • Mengdie Wang

Clustering traditionally aims to reveal a natural grouping structure within unlabeled data. However, this structure may not always align with users' preferences. In this paper, we propose a personalized clustering method that explicitly performs targeted representation learning by interacting with users via modicum task information (e.g., must-link or cannot-link pairs) to guide the clustering direction. We query users with the most informative pairs, i.e., those pairs most hard to cluster and those most easy to miscluster, to facilitate the representation learning in terms of the clustering preference. Moreover, by exploiting attention mechanism, the targeted representation is learned and augmented. By leveraging the targeted representation and constrained contrastive loss as well, personalized clustering is obtained. Theoretically, we verify that the risk of personalized clustering is tightly bounded, guaranteeing that active queries to users do mitigate the clustering risk. Experimentally, extensive results show that our method performs well across different clustering tasks and datasets, even when only a limited number of queries are available.

ICML Conference 2025 Conference Paper

Unsupervised Learning for Class Distribution Mismatch

  • Pan Du 0002
  • Wangbo Zhao
  • Xinai Lu
  • Nian Liu
  • Zhikai Li
  • Chaoyu Gong
  • Suyun Zhao
  • Hong Chen 0001

Class distribution mismatch (CDM) refers to the discrepancy between class distributions in training data and target tasks. Previous methods address this by designing classifiers to categorize classes known during training, while grouping unknown or new classes into an "other" category. However, they focus on semi-supervised scenarios and heavily rely on labeled data, limiting their applicability and performance. To address this, we propose Unsupervised Learning for Class Distribution Mismatch (UCDM), which constructs positive-negative pairs from unlabeled data for classifier training. Our approach randomly samples images and uses a diffusion model to add or erase semantic classes, synthesizing diverse training pairs. Additionally, we introduce a confidence-based labeling mechanism that iteratively assigns pseudo-labels to valuable real-world data and incorporates them into the training process. Extensive experiments on three datasets demonstrate UCDM’s superiority over previous semi-supervised methods. Specifically, with a 60% mismatch proportion on Tiny-ImageNet dataset, our approach, without relying on labeled data, surpasses OpenMatch (with 40 labels per class) by 35. 1%, 63. 7%, and 72. 5% in classifying known, unknown, and new classes.

AAAI Conference 2023 Conference Paper

Echo of Neighbors: Privacy Amplification for Personalized Private Federated Learning with Shuffle Model

  • Yixuan Liu
  • Suyun Zhao
  • Li Xiong
  • Yuhan Liu
  • Hong Chen

Federated Learning, as a popular paradigm for collaborative training, is vulnerable against privacy attacks. Different privacy levels regarding users' attitudes need to be satisfied locally, while a strict privacy guarantee for the global model is also required centrally. Personalized Local Differential Privacy (PLDP) is suitable for preserving users' varying local privacy, yet only provides a central privacy guarantee equivalent to the worst-case local privacy level. Thus, achieving strong central privacy as well as personalized local privacy with a utility-promising model is a challenging problem. In this work, a general framework (APES) is built up to strengthen model privacy under personalized local privacy by leveraging the privacy amplification effect of the shuffle model. To tighten the privacy bound, we quantify the heterogeneous contributions to the central privacy user by user. The contributions are characterized by the ability of generating “echos” from the perturbation of each user, which is carefully measured by proposed methods Neighbor Divergence and Clip-Laplace Mechanism. Furthermore, we propose a refined framework (S-APES) with the post-sparsification technique to reduce privacy loss in high-dimension scenarios. To the best of our knowledge, the impact of shuffling on personalized local privacy is considered for the first time. We provide a strong privacy amplification effect, and the bound is tighter than the baseline result based on existing methods for uniform local privacy. Experiments demonstrate that our frameworks ensure comparable or higher accuracy for the global model.

IJCAI Conference 2022 Conference Paper

Exploring Binary Classification Hidden within Partial Label Learning

  • Hengheng Luo
  • Yabin ZHANG
  • Suyun Zhao
  • Hong Chen
  • Cuiping Li

Partial label learning (PLL) is to learn a discriminative model under incomplete supervision, where each instance is annotated with a candidate label set. The basic principle of PLL is that the unknown correct label y of an instance x resides in its candidate label set s, i. e. , P(y ∈ s | x) = 1. On which basis, current researches either directly model P(x | y) under different data generation assumptions or propose various surrogate multiclass losses, which all aim to encourage the model-based Pθ(y ∈ s | x)→1 implicitly. In this work, instead, we explicitly construct a binary classification task toward P(y ∈ s | x) based on the discriminative model, that is to predict whether the model-output label of x is one of its candidate labels. We formulate a novel risk estimator with estimation error bound for the proposed PLL binary classification risk. By applying logit adjustment based on disambiguation strategy, the practical approach directly maximizes Pθ(y ∈ s | x) while implicitly disambiguating the correct one from candidate labels simultaneously. Thorough experiments validate that the proposed approach achieves competitive performance against the state-of-the-art PLL methods.

ECAI Conference 2020 Conference Paper

Partial Label Learning via Generative Adversarial Nets

  • Yabin Zhang 0005
  • Guang Yang 0040
  • Suyun Zhao
  • Peng Ni
  • Hairong Lian
  • Hong Chen 0001
  • Cuiping Li 0001

Partial label learning (PLL) is a weakly supervised learning framework, in which each sample is provided with multiple candidate labels while only one of them is correct. Most of the existing methods are designed based on some conventional machine learning techniques, from kNN to SVM and logistic regression. Till now, it is still unclear whether we can use adversarial networks to solve partial label problems. This paper gives a positive answer to this question for the first time. We are the first to solve partial label learning with the network structure of CGAN combine SSGAN. In partial label learning with adversarial networks, it is interesting to find that some fake samples close to real sample distribution are generated, and then all these samples gradually promote discriminator to disambiguate the candidate labels of real samples. We give theoretical justifications of PL-GAN on challenging partial label data classification. Numerical experiments on artificial and real-world partial label datasets show that our approach significantly outperforms state-of-the-art counter-parts.

v2026.09.13