Arrow Research search

Author name cluster

Jiayi Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

16 papers
2 author rows

Possible papers

16

AAAI Conference 2026 Conference Paper

A Tale of Two Identities: An Ethical Audit of AI-Crafted Synthetic Personas

  • Pranav Narayanan Venkit
  • Jiayi Li
  • Yingfan Zhou
  • Sarah Rajtmajer
  • Shomir Wilson

As LLMs (large language models) are increasingly used to generate synthetic personas, particularly in data-limited domains such as health, privacy, and HCI, it becomes necessary to understand how these narratives represent identity, especially that of minority communities. In this paper, we audit synthetic personas generated by 3 LLMs (GPT4o, Gemini 1.5 Pro, Deepseek v2.5) through the lens of representational harm, focusing specifically on racial identity. Using a mixed-methods approach combining close reading, lexical analysis, and a parameterized creativity framework, we compare 1,512 LLM-generated persona to human-authored responses. Our findings reveal that LLMs disproportionately foreground racial markers, overproduce culturally coded language, and construct personas that are syntactically elaborate yet narratively reductive. These patterns result in a range of sociotechnical harms, including stereotyping, exoticism, erasure, and benevolent bias, that are often obfuscated by superficially positive narrations. We formalize this phenomenon as algorithmic othering, where minoritized identities are rendered hypervisible but less authentic.

AAAI Conference 2026 Conference Paper

DS-ProGen: A Dual-Structure Deep Language Model for Functional Protein Design

  • Yanting Li
  • Zikang Wang
  • Jiyue Jiang
  • Ziqian Lin
  • Dongchen He
  • Yuheng Shan
  • Yanruisheng Shao
  • Jiayi Li

Inverse Protein Folding (IPF) is a critical subtask in the field of protein design, aiming to engineer amino acid sequences capable of folding correctly into a specified three-dimensional (3D) conformation. Although substantial progress has been achieved in recent years, existing methods generally rely on either backbone coordinates or molecular surface features alone, which restricts their ability to fully capture the complex chemical and geometric constraints necessary for precise sequence prediction. To address this limitation, we present DS-ProGen, a dual-structure deep language model for functional protein design, which integrates both backbone geometry and surface-level representations. By incorporating backbone coordinates as well as surface chemical and geometric descriptors into a next-amino-acid prediction paradigm, DS-ProGen is able to generate functionally relevant and structurally stable sequences while satisfying both global and local conformational constraints. On the PRIDE dataset, DS-ProGen attains the current state-of-the-art recovery rate of 61.47%, demonstrating the synergistic advantage of multi-modal structural encoding in protein design. Furthermore, DS-ProGen excels in predicting interactions with a variety of biological partners, including ligands, ions, and RNA, confirming its robust functional retention capabilities.

AAAI Conference 2026 Conference Paper

From Single to Societal: Analyzing Persona-Induced Bias in Multi-Agent Interactions

  • Jiayi Li
  • Xiao Liu
  • Yansong Feng

Large Language Model (LLM)-based multi-agent systems are increasingly used to simulate human interactions and solve collaborative tasks. A common practice is to assign agents with personas to encourage behavioral diversity. However, this raises a critical yet underexplored question: do personas introduce biases into multi-agent interactions? This paper presents a systematic investigation into persona-induced biases in multi-agent interactions, with a focus on social traits like trustworthiness (how an agent's opinion is received by others) and insistence (how strongly an agent advocates for its opinion). Through a series of controlled experiments in collaborative problem-solving and persuasion tasks, we reveal that (1) LLM-based agents exhibit biases in both trustworthiness and insistence, with personas from historically advantaged groups (e.g., men and White individuals) perceived as less trustworthy and demonstrating less insistence; and (2) agents exhibit significant in-group favoritism, showing a higher tendency to conform to others who share the same persona. These biases persist across various LLMs, group sizes, and numbers of interaction rounds, highlighting an urgent need for awareness and mitigation to ensure the fairness and reliability of multi-agent systems.

AAAI Conference 2026 Conference Paper

RMSAGen: Integrating Multiple Sequence Alignment for Function RNA Design

  • Jiyue Jiang
  • Yanyu Chen
  • Qingchuan Zhang
  • Jiayi Li
  • Xiangyu Shi
  • Chang Zhou
  • Ziqian Lin
  • Jiuming Wang

Biological sequences, including RNAs and proteins, share similarities with natural languages, enabling the application of advanced language models to various biological tasks. However, due to its flexibility and lack of experimental data, RNA is a particularly challenging biological ``language'' compared to other biological sequences like proteins. RNA multiple sequence alignments (MSAs), which align evolutionarily related RNA sequences, can greatly enhance RNA biology modeling, as evidenced by their significant roles in structure prediction and function annotation. This raises the question of whether RNA MSAs can also benefit RNA design, which remains unexplored. This paper introduces RMSAGen, a model comprising RMSA-Encoder and RMSA-Decoder, that leverages MSAs to design functional RNA sequences. RMSA-Encoder effectively extracts MSA features, enhancing performance in functional prediction and solvent accessibility prediction tasks and supporting RMSA-Decoder in accurate RNA generation. RMSAGen can design RNA sequences that effectively bind to target RNA-binding proteins, and the design performance improves with an increasing number of sequences. In addition, the ribozymes designed with structural features by RMSAGen show strong computational metrics and exhibit biological activity during gel electrophoresis. These results highlight the effectiveness of RMSAGen, establishing it as a powerful tool and a new direction for RNA design.

IJCAI Conference 2025 Conference Paper

DASS: A Dual-Branch Attention-based Framework for Trajectory Similarity Learning with Spatial and Semantic Fusion

  • Jiayi Li
  • Junhua Fang
  • Pingfu Chao
  • Jiajie Xu
  • Pengpeng Zhao

Trajectory similarity aims to identify pairs of similar trajectories, serving as a crucial operation in spatial-temporal data mining. Although several approaches have been proposed, they encounter the following two issues: 1) An overemphasis on spatial similarity in road networks while the rich semantic information embedded in trajectories is not fully exploited; 2) Dependence on Recurrent Neural Network (RNN) architectures would struggle to capture long-term dependencies. To address these limitations, we propose a Dual-branch Attention-based framework with Spatial and Semantic information (DASS) based on self-supervised learning. Specifically, DASS comprises two core components: 1) A trajectory representation module that models spatial-temporal adjacent relationships in the form of graph and converts semantics into numerical embeddings. 2) A backbone encoder with a co-attention module to independently process two features before they are integrated. Extensive experiments on real-world datasets demonstrate that DASS outperforms state-of-the-art methods, establishing itself as a novel paradigm.

YNIMG Journal 2025 Journal Article

Motor-cognitive aging: The role of motor cortex and its pathways

  • Jiaqi Wen
  • Zifei Liang
  • Chenyang Li
  • Huize Pang
  • Li Jiang
  • Jiayi Li
  • Xiaojun Guan
  • Jiangyang Zhang

BACKGROUND: Motor and cognitive decline are hallmark features of aging. In the primary motor cortex (M1), pyramidal neurons project to the corticospinal tract (CST), a well-established motor pathway, and send collaterals to the ipsilateral striatum, forming the corticostriatal tract (CStrT). While the CST has been extensively studied, the role of the CStrT in motor and cognitive aging remains poorly understood. METHODS: We analyzed T1- and T2-weighted MRI, multi-delay arterial spin labeling, and multi-shell diffusion MRI data from 339 right-handed healthy adults (aged 36-90 years) in the Human Connectome Project-Aging dataset. Age-related trajectories of M1 structure and hemodynamics, as well as CST and CStrT microstructure, were assessed. Segment-wise along-tract analyses were conducted to identify localized tract degeneration. Mediation analyses were performed to examine whether tract integrity linked M1 atrophy to motor and cognitive performance. RESULTS: With age, M1 exhibited reduced volume and hemodynamics, altered T1/T2 ratio, and increased cortical curvature, reflecting structural and hemodynamic alterations. Along-tract analyses revealed localized microstructural degeneration in the CST adjacent to M1, whereas the CStrT showed more extensive degeneration along its trajectory. These tract changes were associated with structural and hemodynamic alterations in M1. Furthermore, integrity of the dominant (left) CST and CStrT mediated the relationship between ipsilateral M1 atrophy and motor decline. Notably, CStrT integrity also mediated the association between M1 atrophy and motor cognition decline. CONCLUSION: These findings establish age-related structural and functional degeneration of M1 and its pathways, highlighting the CStrT as a critical mediator between motor cortical atrophy and both motor and cognitive decline. These normative imaging markers of healthy aging may help inform the early detection of neurodegenerative diseases.

JBHI Journal 2025 Journal Article

Prediction of Drug-Target Interactions Based on Hypergraph Neural Networks With Multimodal Feature Fusion

  • Yufang Zhang
  • Jiayi Li
  • Shangqing Zhao
  • Hong Tan
  • Heqi Sun
  • Yi Xiong
  • Dong-Qing Wei

Accurate drug-target interaction prediction is vital for drug discovery and optimization. Traditional experimental methods, while effective, are time-intensive and costly. HyperGCN-DTI, a novel framework that explicitly advances beyond existing models such as CHL-DTI and HHDTI by leveraging hypergraph neural networks with a multimodal feature fusion strategy. While exsiting methods primarily focuses on low-order graph representations and fixed heterogeneous network structures, HyperGCN-DTI incorporates richer multimodal fused features including embeddings from pretrained language models and diverse biological networks and build robust hypergraphs that capture high-order multi-entity relationships within drug-target pairs. This dual-channel architecture effectively captures both local topological connections and higher-order structural dependencies. HyperGCN-DTI outperforms state-of-the-art DTI prediction models across multiple datasets and remains robust under imbalanced and large-scale real-world datasets, demonstrating its superior predictive power. The model demonstrates significant improvements when using multimodal features and hypergraph-based message passing, with sensitivity analysis confirming stability across hyperparameter variations. Top-ranked predictions are validated through biomedical literature and molecular docking, underscoring the reliability and practical relevance of our approach. HyperGCN-DTI is the first DTI prediction model to jointly integrate such a wide range of heterogeneous information sources with hypergraph representation, significantly enhancing accuracy and robustness, particularly in sparse or noisy settings. The proposed model offers a powerful and generalizable tool for accelerating drug development and target identification.

IJCAI Conference 2024 Conference Paper

Accelerating Diffusion Models for Inverse Problems through Shortcut Sampling

  • Gongye Liu
  • Haoze Sun
  • Jiayi Li
  • Fei Yin
  • Yujiu Yang

Diffusion models have recently demonstrated an impressive ability to address inverse problems in an unsupervised manner. While existing methods primarily focus on modifying the posterior sampling process, the potential of the forward process remains largely unexplored. In this work, we propose Shortcut Sampling for Diffusion(SSD), a novel approach for solving inverse problems in a zero-shot manner. Instead of initiating from random noise, the core concept of SSD is to find a specific transitional state that bridges the measurement image y and the restored image x. By utilizing the shortcut path of "input - transitional state - output", SSD can achieve precise restoration with fewer steps. To derive the transitional state during the forward process, we introduce Distortion Adaptive Inversion. Moreover, we apply back projection as additional consistency constraints during the generation process. Experimentally, we demonstrate SSD's effectiveness on multiple representative IR tasks. Our method achieves competitive results with only 30 NFEs compared to state-of-the-art zero-shot methods(100 NFEs) and outperforms them with 100 NFEs in certain tasks. Code is available at https: //github. com/GongyeLiu/SSD.

EAAI Journal 2024 Journal Article

Image colour application rules of Shanghai style Chinese paintings based on machine learning algorithm

  • Rongrong Fu
  • Jiayi Li
  • Chaoxiang Yang
  • Junxuan Li
  • Xiaowen Yu

Colour is an important factor in the expression of recognizability and cultural identity of regional cultural and creative design. At present, the colour recognition of regional characteristic and the colour association of regional culture mainly rely on the designer's subjective perception. To obtain the target colour resources with reference for regional cultural and design need, this study proposes a scientific method of colour extraction and strong colour association matching of Shanghai style Chinese paintings by machine learning, and applies the related results to the colour design of cultural and creative products. Firstly, using the SLIC superpixel algorithm and Mean shift algorithm to realize the overall dimensionality reduction of the image features and the colour aggregation gradually, so as to extract the characteristic colours of Shanghai style Chinese paintings; Secondly, we introduce a data mining algorithm (Apriori) to mine out the association rules from multiple characteristics colours and filter out strongly associated colour combinations; Finally, we apply the colour combinations and colour tones to the colour design of the creative products. In order to verify the scientificity of the colour extraction and colour matching method proposed in this paper, we selected another painter's paintings in the same school as the algorithm experimental validation sample, similar results were obtained. In addition, we measured user satisfaction using the degree of awakening to Shanghai style culture and the propensity to make consumer decisions as the evaluation dimensions, which proves that the method of this study is effective.

YNIMG Journal 2024 Journal Article

Noninvasive focused ultrasound-mediated delivery of rAAV9-EGFP vectors for neuronal targeting in rats

  • Rui Wang
  • Jiayi Li

OBJECTIVE: To evaluate the synergistic potential of Focused Ultrasound (FUS) in conjunction with microbubbles (MB) and recombinant adeno-associated virus serotype 9 (rAAV9) vectors for targeted gene delivery to neuronal cells in rats, optimizing gene expression conditions and assessing any adverse effects. METHODS: The parameters for permeability enhancement of the rat's blood-brain barrier (BBB) were established using FUS+MB, with MRI scans and Evans Blue (EB) dye assisting in the evaluation. Rats underwent FUS-mediated transfection using rAAV9-Syn-EGFP vectors produced via a triple-transfection in HEK293T cells. Following this, the uptake and expression of GFP in targeted brain regions were evaluated using confocal fluorescence microscopy at various time intervals. Inflammatory responses post-FUS treatment were tracked by observing levels of GFAP, a marker for astrocytic activation, and TNF-α, a pro-inflammatory cytokine. Motor behavior effects post-intervention were gauged using the Rotarod test across multiple groups over a span of four weeks. RESULTS: FUS+MB affected BBB permeability, with optimal results at 4 W for 200 s showing 85 % permeability and evident Gd-DTPA leakage. Settings beyond these resulted in tissue damage. Control groups exhibited a basal GFP expression of 2 % ± 0.5 %, whereas FUS+MB with rAAV-EGFP injections substantially increased GFP expression to about 67 % ± 6 % in targeted neurons. This GFP expression peaked at three weeks post-treatment and remained evident six months later. Following FUS treatment, both GFAP and TNF-α levels underwent fluctuations before eventually nearing their baseline values. The Rotarod test revealed no significant behavioral differences post-treatments among the groups. CONCLUSIONS: Combining FUS+MB with rAAV offers an innovative approach to enhance therapeutic delivery to the central nervous system (CNS) by transiently adjusting BBB permeability.

TMLR Journal 2024 Journal Article

Pull-back Geometry of Persistent Homology Encodings

  • Shuang Liang
  • Renata Turkes
  • Jiayi Li
  • Nina Otter
  • Guido Montufar

Persistent homology (PH) is a method for generating topology-inspired representations of data. Empirical studies that investigate the properties of PH, such as its sensitivity to perturbations or ability to detect a feature of interest, commonly rely on training and testing an additional model on the basis of the PH representation. To gain more intrinsic insights about PH, independently of the choice of such a model, we propose a novel methodology based on the pull-back geometry that a PH encoding induces on the data manifold. The spectrum and eigenvectors of the induced metric help to identify the most and least significant information captured by PH. Furthermore, the pull-back norm of tangent vectors provides insights about the sensitivity of PH to a given perturbation, or its potential to detect a given feature of interest, and in turn its ability to solve a given classification or regression problem. Experimentally, the insights gained through our methodology align well with the existing knowledge about PH. Moreover, we show that the pull-back norm correlates with the performance on downstream tasks, and can therefore guide the choice of a suitable PH encoding.

EAAI Journal 2023 Journal Article

An improved spherical evolution with enhanced exploration capabilities to address wind farm layout optimization problem

  • Haichuan Yang
  • Shangce Gao
  • Zhenyu Lei
  • Jiayi Li
  • Yang Yu
  • Yirui Wang

The utilization of metaheuristics for optimizing wind farm layouts (WFLOP) has emerged as a popular research area in recent years. However, effectively screening and improving metaheuristics to obtain optimal layouts remain a challenging task. Traditional metaheuristic screening methods require testing numerous algorithms, resulting in high computational resource consumption and trial-and-error costs due to the lack of theoretical guidance. To overcome this challenge, this study proposes a complex network-based metaheuristic screening method. Population interaction networks are utilized to classify metaheuristics into two categories: biased exploitation and biased exploration. The results of several metaheuristics on WFLOP suggest that exploration-biased algorithms generally outperform exploitation-biased ones. This discovery holds great significance as it has the potential to predict the performance of various algorithms on WFLOP to a certain degree. Additionally, it provides valuable suggestions for algorithm selection and improvement. Building upon this new methodology, we screen and improve the spherical evolution algorithm to enhance its exploration capabilities. Experimental results demonstrate that the improved spherical evolution algorithm significantly outperforms its competitors on WFLOP.

ICLR Conference 2022 Conference Paper

Meta-Imitation Learning by Watching Video Demonstrations

  • Jiayi Li
  • Tao Lu 0006
  • Xiaoge Cao
  • Yinghao Cai
  • Shuo Wang 0001

Meta-Imitation Learning is a promising technique for the robot to learn a new task from observing one or a few human demonstrations. However, it usually requires a significant number of demonstrations both from humans and robots during the meta-training phase, which is a laborious and hard work for data collection, especially in recording the actions and specifying the correspondence between human and robot. In this work, we present an approach of meta-imitation learning by watching video demonstrations from humans. In comparison to prior works, our approach is able to translate human videos into practical robot demonstrations and train the meta-policy with adaptive loss based on the quality of the translated data. Our approach relies only on human videos and does not require robot demonstration, which facilitates data collection and is more in line with human imitation behavior. Experiments reveal that our method achieves the comparable performance to the baseline on fast learning a set of vision-based tasks through watching a single video demonstration.

ICRA Conference 2021 Conference Paper

DIMSAN: Fast Exploration with the Synergy between Density-based Intrinsic Motivation and Self-adaptive Action Noise

  • Jiayi Li
  • Boyao Li
  • Tao Lu 0006
  • Ning Lu
  • Yinghao Cai
  • Shuo Wang 0001

Exploration in environments with sparse rewards remains a challenging problem in Deep Reinforcement Learning (DRL). For the off-policy method, it usually needs a large number of training samples. With the growing dimensions of state and action space, this method becomes more and more sample-inefficient. In this paper, we propose a novel fast exploration method for off-policy reinforcement learning, called Density-based Intrinsic Motivation and Self-adaptive Action Noise (DIMSAN). Our main contribution is twofold: (1) We propose a Density-based Intrinsic Motivation (DIM) method. It introduces a new intrinsic-reward generation mechanism based on samples’ density estimation during experience replay and encourages the agent to seek novel and unfamiliar states. (2) We propose a Self-adaptive Action Noise (SAN) to deal with the exploration-exploitation tradeoffs, which could automatically change the exploration step through adding adaptive action space noise. The synergy between DIM and SAN could guide the agent to search the state and action space with high efficiency. We evaluate our method on the benchmark manipulation tasks and the designed challenging ones. Empirical results show that our method outperforms the existing methods in terms of convergence speed and sample efficiency, especially in challenging tasks.

ICRA Conference 2021 Conference Paper

Hierarchical Learning from Demonstrations for Long-Horizon Tasks

  • Boyao Li
  • Jiayi Li
  • Tao Lu 0006
  • Yinghao Cai
  • Shuo Wang 0001

Although reinforcement learning (RL) has achieved great success in robotic manipulation skills learning, it is still challenging for long-horizon tasks. Combining RL with demonstrations is an effective solution. In this paper, we propose a novel hierarchical learning from demonstrations method for long-horizon tasks, which leverages (i) object-centered segmentation of demonstrations to automatically segment the teaching trajectories into episodes. (ii) a bi-level hierarchical imitation learning method with a parallel training mechanism to train the two-level policies simultaneously. Experimental results on three challenging long-horizon tasks with sparse rewards show that our proposed method significantly outperforms state-of-art approaches in terms of both sample-efficiency and success rate. Moreover, our method is the only one which achieves satisfactory performance in tasks of multi-object stack and multi-object push&stack.

ICRA Conference 2020 Conference Paper

ACDER: Augmented Curiosity-Driven Experience Replay

  • Boyao Li
  • Tao Lu 0006
  • Jiayi Li
  • Ning Lu
  • Yinghao Cai
  • Shuo Wang 0001

Exploration in environments with sparse feed-back remains a challenging research problem in reinforcement learning (RL). When the RL agent explores the environment randomly, it results in low exploration efficiency, especially in robotic manipulation tasks with high dimensional continuous state and action space. In this paper, we propose a novel method, called Augmented Curiosity-Driven Experience Replay (ACDER), which leverages (i) a new goal-oriented curiosity-driven exploration to encourage the agent to pursue novel and task-relevant states more purposefully and (ii) the dynamic initial states selection as an automatic exploratory curriculum to further improve the sample-efficiency. Our approach complements Hindsight Experience Replay (HER) by introducing a new way to pursue valuable states. Experiments conducted on four challenging robotic manipulation tasks with binary rewards, including Reach, Push, Pick&Place and Multi-step Push. The empirical results show that our proposed method significantly outperforms existing methods in the first three basic tasks and also achieves satisfactory performance in multi-step robotic task learning.

v2026.09.13