Arrow Research search

Author name cluster

Yue Han

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
2 author rows

Possible papers

8

ICML Conference 2025 Conference Paper

Flow Matching for Denoised Social Recommendation

  • Yinxuan Huang
  • Ke Liang
  • Zhuofan Dong
  • Xiaodong Qu
  • Tianxiang Wang
  • Yue Han
  • Jingao Xu
  • Bin Zhou 0004

Graph-based social recommendation (SR) models suffer from various noises of the social graphs, hindering their recommendation performances. Either graph-level redundancy or graph-level missing will indeed influence the social graph structures, further influencing the message propagation procedure of graph neural networks (GNNs). Generative models, especially diffusion-based models, are usually used to reconstruct and recover the data in better quality from original data with noises. Motivated by it, a few works take attempts on it for social recommendation. However, they can only handle isotropic Gaussian noises but fail to leverage the anisotropic ones. Meanwhile the anisotropic relational structures in social graphs are commonly seen, so that existing models cannot sufficiently utilize the graph structures, which constraints the capacity of noise removal and recommendation performances. Compared to the diffusion strategy, the flow matching strategy shows better ability to handle the data with anisotropic noises since they can better preserve the data structures during the learning procedure. Inspired by it, we propose RecFlow which is the first flow-matching based SR model. Concretely, RecFlow performs flow matching on the structure representations of social graphs. Then, a conditional learning procedure is designed for optimization. Extensive performances prove the promising performances of our RecFlow from six aspects, including superiority, effectiveness, robustnesses, sensitivity, convergence and visualization.

IROS Conference 2025 Conference Paper

IHGSL: Interpretable Heuristic Graph Structure Learning for Multi-Robot Autonomous Collaborative Systems

  • Yue Han
  • Hanqi Li
  • Cuiwei Liu
  • Chen Liang
  • Zhixiao Sun

In multi-robot systems, capturing the complex and dynamic interaction relationships is essential for enhancing autonomous collaboration. However, existing learning-based approaches usually overlook the understanding of these relationships, leading to reliability issues and hindering their application to real-world scenarios. This paper proposes a novel approach called Interpretable Heuristic Graph Structure Learning (IHGSL) to better comprehend the complex collaborative relationships in multi-robot systems. We first construct a predicate space to define diverse predicates that express fundamental relationships. Then we employ the variational information bottleneck technique to acquire a latent representation of the current observation by aligning it with the historical trajectory. On this basis, the predicates that the robot should currently focus on the most are learned, and some interaction relationships are established accordingly. Thereby an interpretable relationship graph is generated heuristically to guide the achievement of multi-robot autonomous collaborative decision-making. Through experimental evaluation, we demonstrate the process of relationship inference, thus validating the interpretability of IHGSL. Compared with existing methods, IHGSL also achieves superior collaboration performance, which highlights the effectiveness of the learned heuristic graph structure.

TMLR Journal 2025 Journal Article

LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects

  • Guangyi Liu
  • Pengxiang Zhao
  • Yaozhen Liang
  • Liang Liu
  • Yaxuan Guo
  • Han Xiao
  • Weifeng Lin
  • Yuxiang Chai

With the rapid rise of large language models (LLMs), phone automation has undergone transformative changes. This paper systematically reviews LLM-driven phone GUI agents, highlighting their evolution from script-based automation to intelligent, adaptive systems. We first contextualize key challenges, (i) limited generality, (ii) high maintenance overhead, and (iii) weak intent comprehension, and show how LLMs address these issues through advanced language understanding, multimodal perception, and robust decision-making. We then propose a taxonomy covering fundamental agent frameworks (single-agent, multi-agent, plan-then-act), modeling approaches (prompt engineering, training-based), and essential datasets and benchmarks. Furthermore, we detail task-specific architectures, supervised fine-tuning, and reinforcement learning strategies that bridge user intent and GUI operations. Finally, we discuss open challenges such as dataset diversity, on-device deployment efficiency, user-centric adaptation, and security concerns, offering forward-looking insights into this rapidly evolving field. By providing a structured overview and identifying pressing research gaps, this paper serves as a definitive reference for researchers and practitioners seeking to harness LLMs in designing scalable, user-friendly phone GUI agents. The collection of papers reviewed in this survey will be hosted and regularly updated on the GitHub repository: \url{https://github.com/PhoneLLM/Awesome-LLM-Powered-Phone-GUI-Agents}

AAAI Conference 2025 Conference Paper

Metric Distortion of Line-up Elections: The Right Person for the Right Job

  • Christopher Jerrett
  • Yue Han
  • Elliot Anshelevich

We provide mechanisms and new metric distortion bounds for line-up elections. In such elections, a set of n voters, k candidates, and ell positions are all located in a metric space. The goal is to choose a set of candidates and assign them to different positions, so as to minimize the total cost of the voters. The cost of each voter consists of the distances from itself to the chosen candidates (measuring how much the voter likes the chosen candidates, or how similar it is to them), as well as the distances from the candidates to the positions they are assigned to (measuring the fitness of the candidates for their positions). Our mechanisms, however, do not know the exact distances, and instead produce good outcomes while only using a smaller amount of information, resulting in small distortion. We consider several different types of information: ordinal voter preferences, ordinal position preferences, and knowing the exact locations of candidates and positions, but not those of voters. In each of these cases, we provide constant distortion bounds, thus showing that only a small amount of information is enough to form outcomes close to optimum in line-up elections.

NeurIPS Conference 2025 Conference Paper

MS-Bench: Evaluating LMMs in Ancient Manuscript Study through a Dunhuang Case Study

  • Yuqing Zhang
  • Yue Han
  • Shuanghe Zhu
  • Haoxiang Wu
  • Hangqi Li
  • Shengyu Zhang
  • Junchi Yan
  • Zemin Liu

Analyzing ancient manuscripts has traditionally been a labor-intensive and time-consuming task for philologists. While recent advancements in LMMs have demonstrated their potential across diverse domains, their effectiveness in manuscript study remains underexplored. In this paper, we introduce MS-Bench, the first comprehensive benchmark co-developed with archaeologists, comprising 5, 076 high-resolution images from 4th to 14th century and 9, 982 expert-curated questions across nine sub-tasks aligned with archaeological workflows. Through four prompting strategies, we systematically evaluate 32 LMMs on their effectiveness, robustness, and cultural contextualization. Our analysis reveals scale-driven performance and reliability improvements, prompting strategies' impact on performance (CoT has two-sides effect, while visual retrieval-augmented prompts provide consistent boost), and task-specific preferences depending on LMM’s visual capabilities. Although current LMMs are not yet capable of replacing domain expertise, they demonstrate promising potential to accelerate manuscript research through future human–AI collaboration.

YNIMG Journal 2024 Journal Article

Clinical characteristics of post-stroke basal ganglia aphasia and the study of language-related white matter tracts based on diffusion spectrum imaging

  • Yue Han
  • Yuanyuan Jing
  • Xuewei Li
  • Hongwei Zhou
  • Fang Deng

BACKGROUND: Stroke often damages the basal ganglia, leading to atypical and transient aphasia, indicating that post-stroke basal ganglia aphasia (PSBGA) may be related to different anatomical structural damage and functional remodeling rehabilitation mechanisms. The basal ganglia contain dense white matter tracts (WMTs). Hence, damage to the functional tract may be an essential anatomical structural basis for the development of PSBGA. METHODS: We first analyzed the clinical characteristics of PSBGA in 28 patients and 15 healthy controls (HCs) using the Western Aphasia Battery and neuropsychological test batteries. Moreover, we investigated white matter injury during the acute stage using diffusion magnetic resonance imaging scans for differential tractography. Finally, we used multiple regression models in correlation tractography to analyze the relationship between various language functions and quantitative anisotropy (QA) of WMTs. RESULTS: Compared with HCs, patients with PSBGA showed lower scores for fluency, comprehension (auditory word recognition and sequential commands), naming (object naming and word fluency), reading comprehension of sentences, Mini-Mental State Examination, and Montreal Cognitive Assessment, along with increased scores in Hamilton Anxiety Scale-17 and Hamilton Depression Scale-17 within 7 days after stroke onset (P < 0.05). Differential tractography revealed that patients with PSBGA had damaged fibers, including in the body fibers of the corpus callosum, left cingulum bundles, left parietal aslant tracts, bilateral superior longitudinal fasciculus II, bilateral thalamic radiation tracts, left fornix, corpus callosum tapetum, and forceps major, compared with HCs (FDR < 0.02). Correlation tractography highlighted that better comprehension was correlated with a higher QA of the left inferior fronto-occipital fasciculus (IFOF), corpus callosum forceps minor, and left extreme capsule (FDR < 0.0083). Naming was positively associated with the QA of the left IFOF, forceps minor, left arcuate fasciculus, and uncinate fasciculus (UF) (FDR < 0.0083). Word fluency of naming was also positively associated with the QA of the forceps minor, left IFOF, and thalamic radiation tracts (FDR < 0.0083). Furthermore, reading was positively correlated with the QA of the forceps minor, left IFOF, and UF (FDR < 0.0083). CONCLUSION: PSBGA is primarily characterized by significantly impaired word fluency of naming and preserved repetition abilities, as well as emotional and cognitive dysfunction. Damaged limbic pathways, dorsally located tracts in the left hemisphere, and left basal ganglia pathways are involved in PSBGA pathogenesis. The results of connectometry analysis further refine the current functional localization model of higher-order neural networks associated with language functions.

AAAI Conference 2023 Conference Paper

Optimizing Multiple Simultaneous Objectives for Voting and Facility Location

  • Yue Han
  • Christopher Jerrett
  • Elliot Anshelevich

We study the classic facility location setting, where we are given n clients and m possible facility locations in some arbitrary metric space, and want to choose a location to build a facility. The exact same setting also arises in spatial social choice, where voters are the clients and the goal is to choose a candidate or outcome, with the distance from a voter to an outcome representing the cost of this outcome for the voter (e.g., based on their ideological differences). Unlike most previous work, we do not focus on a single objective to optimize (e.g., the total distance from clients to the facility, or the maximum distance, etc.), but instead attempt to optimize several different objectives simultaneously. More specifically, we consider the l-centrum family of objectives, which includes the total distance, max distance, and many others. We present tight bounds on how well any pair of such objectives (e.g., max and sum) can be simultaneously approximated compared to their optimum outcomes. In particular, we show that for any such pair of objectives, it is always possible to choose an outcome which simultaneously approximates both objectives within a factor of 1 plus square root of 2, and give a precise characterization of how this factor improves as the two objectives being optimized become more similar. For q>2 different centrum objectives, we show that it is always possible to approximate all q of these objectives within a small constant, and that this constant approaches 3 as q increases. Our results show that when optimizing only a few simultaneous objectives, it is always possible to form an outcome which is a significantly better than 3 approximation for all of these objectives.

AAAI Conference 2022 Conference Paper

SCSNet: An Efficient Paradigm for Learning Simultaneously Image Colorization and Super-resolution

  • Jiangning Zhang
  • Chao Xu
  • Jian Li
  • Yue Han
  • Yabiao Wang
  • Ying Tai
  • Yong Liu

In the practical application of restoring low-resolution grayscale images, we generally need to run three separate processes of image colorization, super-resolution, and dowssampling operation for the target device. However, this pipeline is redundant and inefficient for the independent processes, and some inner features could have been shared. Therefore, we present an efficient paradigm to perform Simultaneously Image Colorization and Super-resolution (SCS) and propose an end-to-end SCSNet to achieve this goal. The proposed method consists of two parts: colorization branch for learning color information that employs the proposed plug-and-play Pyramid Valve Cross Attention (PV- CAttn) module to aggregate feature maps between source and reference images; and super-resolution branch for integrating color and texture information to predict target images, which uses the designed Continuous Pixel Mapping (CPM) module to predict high-resolution images at continuous magnification. Furthermore, our SCSNet supports both automatic and referential modes that is more flexible for practical application. Abundant experiments demonstrate the superiority of our method for generating authentic images over state-of-theart methods, e. g. , averagely decreasing FID by 1. 8↓ and 5. 1 ↓ compared with current best scores for automatic and referential modes, respectively, while owning fewer parameters (more than ×2↓) and faster running speed (more than ×3↑).

v2026.09.13