Arrow Research search

Author name cluster

Xiucai Ye

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
1 author row

Possible papers

6

JBHI Journal 2026 Journal Article

A Dual-Language-Model Framework for Reproducibility in Small Molecule-RNA Binding Site Prediction

  • Shixuan Guan
  • Xiucai Ye
  • Tetsuya Sakurai

Single-seed evaluation—the dominant reporting practice in small-dataset molecular learning—can substantially inflate performance estimates yet remains largely unexamined. We present the first systematic reproducibility analysis for RNA–ligand binding site prediction by integrating two large pretrained RNA language models (RNA-FM and RiNALMo) across multiple fusion architectures and replicated training runs on the TR60/TE18 benchmark. Our analysis reveals a pronounced Peak–SOTA Paradox: a favorable initialization in the Reverse Cross-Attention model reached an MCC of 0. 353, surpassing the reported state-of-the-art (0. 327), whereas multi-seed replication yielded only 0. 266 $\pm$ 0. 020—a 32. 8% overestimation. Across architectures, mean accuracy remained tightly clustered, yet reproducibility varied substantially. Simple concat fusion strategies exhibited markedly higher stability than attention-based models, indicating that architectural entanglement rather than parameter count governs variance under data scarcity. Collectively, these findings establish reproducibility as a primary evaluation criterion for small-sample molecular prediction and motivate a dual-reporting standard in which mean $\pm$ SD serves as the principal metric and peak scores as supplementary evidence. This variance-aware perspective highlights that single-seed evaluations can misrepresent expected performance by 20–30% in limited-sample regimes.

EAAI Journal 2026 Journal Article

High-performance computing enhanced task recommendation strategy based on mobile prediction in mobile crowdsensing

  • Jing Zhang
  • Xiangxuan Zhong
  • Zhenhan Huang
  • Li Xu
  • Xiucai Ye

With the development of High-Performance Computing (HPC) and Artificial Intelligence (AI) technologies, Mobile CrowdSensing (MCS) plays an important role in large-scale data processing and analysis. By combining the parallel computing capabilities of HPC, MCS can quickly process complex spatiotemporal data and utilize AI to optimize task recommendations and resource allocation. However, existing task allocation and recommendation models are inefficient due to limited consideration of users’ movement, location preferences, and collaboration needs. To address these issues, a High-performance computing Enhanced Task Recommendation Strategy based on Mobile Prediction (HEtrs-MP) is proposed in this paper. It combines HPC acceleration with deep learning models. Firstly, the User Trajectory Prediction algorithm based on Convolutional Neural Network - Long Short-Term Memory (UTPCL) uses Convolutional Neural Network - Long Short-Term Memory and HPC to predict the users’ location and achieve intelligent task allocation. Secondly, the Time Fuzzy Clustering algorithm based on User Time Preference (TFC-UTP) clusters users’ time preferences, optimizes task time recommendations, and reduces disruption to users. Thirdly, the Task Recommendation algorithm based on User Collaboration (TRUC) uses HPC to analyze user similarity and form collaborative groups, thereby improving task execution efficiency. Finally, extensive experiments are conducted on the GeoLife and T-Driver datasets to validate the effectiveness of the HEtrs-MP strategy. Compared with other task recommendation strategies, the prediction accuracy of the HEtrs-MP strategy can reach up to 96%. The accuracy and hit rate increase by more than 5%. Additionally, the sensing users’ mobility costs are reduced.

JBHI Journal 2026 Journal Article

LLM-Enhanced Knowledge Distillation for Sequence-Based Protein-Ligand Interaction Prediction

  • Wenyu Xi
  • Ruheng Wang
  • Xiucai Ye
  • Tetsuya Sakurai
  • Leyi Wei

Accurate prediction of protein-ligand interactions is essential for drug discovery, supporting critical stages from lead optimization to therapeutic development. Many existing methods depend on high-resolution protein-ligand complex structures, which limits scalability and reduces robustness in structure-limited settings. To address these challenges, we introduce Multi-Combinatorial Knowledge Distillation (MCKD), a sequence-based framework that predicts protein-ligand interactions without requiring explicit three-dimensional structures at inference time. MCKD represents proteins and ligands as two-dimensional molecular graphs derived from their sequences and physicochemical properties, enabling effective learning from readily available inputs. To incorporate structural knowledge beyond sequence information, MCKD employs a hybrid distillation strategy that combines cross-modal distillation from a structure-based teacher with self-distillation to improve representation consistency across layers. To model protein-ligand interactions explicitly, MCKD integrates a bilinear attention network that captures residue-atom level associations and supports both binding affinity regression and binary interaction classification. Evaluations on multiple public benchmark datasets show that MCKD consistently outperforms existing sequence-based methods and achieves performance comparable to structure-based approaches. The model also generalizes well to unseen proteins and novel ligand scaffolds, while providing interpretable insights into key molecular interaction regions. These results suggest that MCKD offers a scalable and effective solution for protein-ligand interaction prediction, particularly for structure-free and data-limited drug discovery applications.

JBHI Journal 2025 Journal Article

PKAN: Leveraging Kolmogorov–Arnold Networks and Multi-Modal Learning for Peptide Prediction With Advanced Language Models

  • Li Wang
  • Xiangzheng Fu
  • Xiucai Ye
  • Tetsuya Sakurai
  • Xiangxiang Zeng
  • Yiping Liu

Peptides can offer highly specific biological activities, serving as essential mediators of intercellular signaling, which are critical for advancing precision medicine and drug development. Their primary structure can be depicted either as an amino acid sequence or as a chemical molecules consisting of atoms and chemical bonds. Large language models (LLMs) hold the potential to thoroughly elucidate the intricate intrinsic properties of peptides. Here we present the Peptide Kolmogorov-Arnold Network (PKAN), a framework leveraging multi-modal representations inspired by advanced language models for peptide activity and functionality prediction. Comparative experiments across tasks show that PKAN outperforms state-of-the-art models while maintaining a streamlined design with superior predictive capabilities. The multi-modal feature importance scoring, anchored in global structures and the significant marginal impacts of derived features on the model, coupled with intricate symbolic regression of specific activation functions, further demonstrates the robustness and precision of the PKAN framework in identifying and elucidating key determinants of peptide functionality. This work provides scientific evidence for investigating the complex mechanisms of peptide materials and supports the progression of peptide language paradigms in biology.

AAAI Conference 2019 Conference Paper

Complex Moment-Based Supervised Eigenmap for Dimensionality Reduction

  • Akira Imakura
  • Momo Matsuda
  • Xiucai Ye
  • Tetsuya Sakurai

Dimensionality reduction methods that project highdimensional data to a low-dimensional space by matrix trace optimization are widely used for clustering and classification. The matrix trace optimization problem leads to an eigenvalue problem for a low-dimensional subspace construction, preserving certain properties of the original data. However, most of the existing methods use only a few eigenvectors to construct the low-dimensional space, which may lead to a loss of useful information for achieving successful classification. Herein, to overcome the deficiency of the information loss, we propose a novel complex moment-based supervised eigenmap including multiple eigenvectors for dimensionality reduction. Furthermore, the proposed method provides a general formulation for matrix trace optimization methods to incorporate with ridge regression, which models the linear dependency between covariate variables and univariate labels. To reduce the computational complexity, we also propose an efficient and parallel implementation of the proposed method. Numerical experiments indicate that the proposed method is competitive compared with the existing dimensionality reduction methods for the recognition performance. Additionally, the proposed method exhibits high parallel efficiency.

IJCAI Conference 2019 Conference Paper

Distributed Collaborative Feature Selection Based on Intermediate Representation

  • Xiucai Ye
  • Hongmin Li
  • Akira Imakura
  • Tetsuya Sakurai

Feature selection is an efficient dimensionality reduction technique for artificial intelligence and machine learning. Many feature selection methods learn the data structure to select the most discriminative features for distinguishing different classes. However, the data is sometimes distributed in multiple parties and sharing the original data is difficult due to the privacy requirement. As a result, the data in one party may be lack of useful information to learn the most discriminative features. In this paper, we propose a novel distributed method which allows collaborative feature selection for multiple parties without revealing their original data. In the proposed method, each party finds the intermediate representations from the original data, and shares the intermediate representations for collaborative feature selection. Based on the shared intermediate representations, the original data from multiple parties are transformed to the same low dimensional space. The feature ranking of the original data is learned by imposing row sparsity on the transformation matrix simultaneously. Experimental results on real-world datasets demonstrate the effectiveness of the proposed method.

v2026.09.13