Arrow Research search

Author name cluster

Chunyang Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
2 author rows

Possible papers

6

ICLR Conference 2025 Conference Paper

Deep Signature: Characterization of Large-Scale Molecular Dynamics

  • Tiexin Qin
  • Mengxu Zhu
  • Chunyang Li
  • Terry Lyons
  • Hong Yan 0001
  • Haoliang Li

Understanding protein dynamics are essential for deciphering protein functional mechanisms and developing molecular therapies. However, the complex high-dimensional dynamics and interatomic interactions of biological processes pose significant challenge for existing computational techniques. In this paper, we approach this problem for the first time by introducing Deep Signature, a novel computationally tractable framework that characterizes complex dynamics and interatomic interactions based on their evolving trajectories. Specifically, our approach incorporates soft spectral clustering that locally aggregates cooperative dynamics to reduce the size of the system, as well as signature transform that collects iterated integrals to provide a global characterization of the non-smooth interactive dynamics. Theoretical analysis demonstrates that Deep Signature exhibits several desirable properties, including invariance to translation, near invariance to rotation, equivariance to permutation of atomic coordinates, and invariance under time reparameterization. Furthermore, experimental results on three benchmarks of biological processes verify that our approach can achieve superior performance compared to baseline methods.

TMLR Journal 2025 Journal Article

The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning

  • Tianshi Zheng
  • Yixiang Chen
  • Chengxi Li
  • Chunyang Li
  • Qing Zong
  • Haochen Shi
  • Baixuan Xu
  • Yangqiu Song

Chain-of-Thought (CoT) prompting has been widely recognized for its ability to enhance reasoning capabilities in large language models (LLMs). However, our study reveals a surprising contradiction to this prevailing perspective within the fundamental domain of pattern-based in-context learning (ICL). Through extensive experiments involving 16 state-of-the-art LLMs and nine diverse pattern-based ICL datasets, we demonstrate that CoT and its reasoning variants consistently underperform direct answering across varying model scales and benchmark complexities. To systematically investigate this unexpected phenomenon, we designed extensive experiments to validate several hypothetical explanations. Our analysis uncovers a fundamental hybrid mechanism of explicit-implicit reasoning driving CoT’s performance in pattern-based ICL: while explicit reasoning falters due to LLMs’ struggles to infer underlying patterns from demonstrations, implicit reasoning—disrupted by the increased contextual distance of CoT rationales—often compensates, delivering correct answers despite flawed rationales. This hybrid mechanism explains CoT’s relative underperformance, as noise from weak explicit inference undermines the process, even as implicit mechanisms partially salvage outcomes. Notably, even long-CoT reasoning models, which excel in abstract and symbolic reasoning, fail to fully overcome these limitations despite higher computational costs. Our findings challenge existing assumptions regarding the universal efficacy of CoT, yielding novel insights into its limitations and guiding future research toward more nuanced and effective reasoning methodologies for LLMs.

ICLR Conference 2024 Conference Paper

KoLA: Carefully Benchmarking World Knowledge of Large Language Models

  • Jifan Yu
  • Xiaozhi Wang
  • Shangqing Tu
  • Shulin Cao
  • Daniel Zhang-Li
  • Xin Lv
  • Hao Peng 0015
  • Zijun Yao 0002

The unprecedented performance of large language models (LLMs) necessitates improvements in evaluations. Rather than merely exploring the breadth of LLM abilities, we believe meticulous and thoughtful designs are essential to thorough, unbiased, and applicable evaluations. Given the importance of world knowledge to LLMs, we construct a Knowledge-oriented LLM Assessment benchmark (KoLA), in which we carefully design three crucial factors: (1) For ability modeling, we mimic human cognition to form a four-level taxonomy of knowledge-related abilities, covering 19 tasks. (2) For data, to ensure fair comparisons, we use both Wikipedia, a corpus prevalently pre-trained by LLMs, along with continuously collected emerging corpora, aiming to evaluate the capacity to handle unseen data and evolving knowledge. (3) For evaluation criteria, we adopt a contrastive system, including overall standard scores for better numerical comparability across tasks and models, and a unique self-contrast metric for automatically evaluating knowledge-creating ability. We evaluate 21 open-source and commercial LLMs and obtain some intriguing findings. The KoLA dataset will be updated every three months to provide timely references for developing LLMs and knowledge-related systems.

JBHI Journal 2022 Journal Article

Bayesian Comorbidity Network and Cost Analysis for Asthma

  • Zhilin Yong
  • Li Luo
  • Yonghong Gu
  • Chunyang Li

The evolving disease spectrum poses significant challenges to the asthma management, thus worsening health quality and increased financial burden on patients. However, potential dependency pattern in comorbidity spectrum remains unclear. We built comorbidity networks based on Bayesian networks utilizing 19604 asthma-patient hospitalization data to investigate dependency patterns among asthma comorbidities. We analyze static properties and trajectory behaviors of gender- and age-stratified asthmatic comorbidity networks. Results suggest that chronic obstructive pulmonary disease, respiratory failure, hypertension, atherosclerosis, and gastritis and duodenitis are the hubs of the asthma comorbidity network. They have a strong dependency pattern, while most of the associations among other comorbidities are sparse and weak. The strength of association between comorbidities is higher in female asthmatics than in males. Although the comorbidity network in children with asthma is simple and stable, the onset of common comorbidities as they age will enhance the association between comorbidities and thus increase the risk of developing other comorbidities. Furthermore, the more attributes of comorbidities, the stronger association with each other, and the greater risk of causing high treatment costs. Our study will help to dissect the asthma co-morbidity network and provide a basis for improving asthma management and cost control.

JBHI Journal 2021 Journal Article

Design Comorbidity Portfolios to Improve Treatment Cost Prediction of Asthma Using Machine Learning

  • Li Luo
  • Xinzhu Yu
  • Zhilin Yong
  • Chunyang Li
  • Yonghong Gu

Comorbidity is an important factor to consider when trying to predict the cost of treating asthma patients. When an asthmatic patient suffered from comorbidity, the cost of treating such a patient becomes dependent on the nature of the comorbidity. Therefore, lack of recognition of comorbidity on asthmatic patient poses a challenge in predicting the cost of treatment. In this study, we proposed a comorbidity portfolio design that improves the prediction cost of treating asthmatic patients by regrouping frequently occurred comorbidities in different cost groups. In the experiment, predictive models, including logistic regression, random forest, support vector machine, classification regression tree, and backpropagation neural network were trained with real-world data of asthmatic patients from 2012 to 2014 in a large city of China. The 10-fold cross validation and random search algorithm were employed to optimize the hyper-parameters. We recorded significant improvements using our model, which are attributed to comorbidity portfolios in area under curve (AUC) and sensitivity increase of 46. 89% (standard deviation: 4. 45%) and 101. 07% (standard deviation: 44. 94%), respectively. In risk analysis of comorbidity on cost, respiratory diseases with a cumulative proportion in the adjusted odds ratio of 36. 38% (95%CI: 27. 61%, 47. 86%) and circulatory diseases with a cumulative proportion in the adjusted odds ratio of 23. 83% (95%CI: 15. 95%, 35. 22%) are the dominant risks of asthmatic patients that affects the treatment cost. It is found that the comorbidity portfolio is robust, and provides a better prediction of the high-cost of treating asthmatic patients. The preliminary characterization of the joint risk of multiple comorbidities posed on cost are also reported. This study will be of great help in improving cost prediction and comorbidity management.

AAAI Conference 2020 Conference Paper

Learning to Generate Maps from Trajectories

  • Sijie Ruan
  • Cheng Long
  • Jie Bao
  • Chunyang Li
  • Zisheng Yu
  • Ruiyuan Li
  • Yuxuan Liang
  • Tianfu He

Accurate and updated road network data is vital in many urban applications, such as car-sharing, and logistics. The traditional approach to identifying the road network, i. e. , field survey, requires a significant amount of time and effort. With the wide usage of GPS embedded devices, a huge amount of trajectory data has been generated by different types of mobile objects, which provides a new opportunity to extract the underlying road network. However, the existing trajectory-based map recovery approaches require many empirical parameters and do not utilize the prior knowledge in existing maps, which over-simplifies or overcomplicates the reconstructed road network. To this end, we propose a deep learning-based map generation framework, i. e. , DeepMG, which learns the structure of the existing road network to overcome the noisy GPS positions. More specifically, DeepMG extracts features from trajectories in both spatial view and transition view and uses a convolutional deep neural network T2RNet to infer road centerlines. After that, a trajectory-based post-processing algorithm is proposed to re- fine the topological connectivity of the recovered map. Extensive experiments on two real-world trajectory datasets con- firm that DeepMG significantly outperforms the state-of-theart methods.

v2026.09.13