Arrow Research search

Author name cluster

Xinyu Wu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
2 author rows

Possible papers

9

AAAI Conference 2026 Conference Paper

BrainLMM: A Label-Free Framework for Mapping Multi-Semantic Representation in the Human Visual Cortex

  • Tan Gao
  • Mufan Xue
  • Haofang Zheng
  • Shuo Lv
  • Jia Xu
  • Dabin Sheng
  • Ziming Mao
  • Xinyu Wu

Previous studies leveraging artificial neural networks have been used to investigate the semantic coding within human visual cortex. However, building an interpretable label-free framework that can effectively map brain responses to multiple coexisting semantic concepts remains largely unexplored. Here, we propose BrainLMM, a label-free framework for multi-semantic mapping of voxel responses by combining diverse vision encoders with the Describe-and-Dissect strategy, enabling a hypothesis-free analysis of the human high-level visual cortex. First, we construct voxel-wise encoding models leveraging diverse vision encoders to predict visual cortical responses to natural scene images. Then, we use BrainLMM to map individual brain voxels to multiple semantics without requiring any predefined labels. To evaluate the effectiveness of our method, we compute Pearson correlation coefficients to compare the multi-semantic mappings produced by BrainLMM and CLIP-MSM with ground-truth voxel responses within selective cortical areas. Our findings indicate that BrainLMM achieves more accurate predictions of visual responses compared to CLIP-MSM. Finally, to demonstrate the multi-semantic mapping capability of our method, we project multiple representative semantic concepts onto the cortical surface for visualization. Our method enables the discovery of voxels that exhibit strong activation in response to previously undefined semantic concepts across two independent datasets: the Natural Scenes Dataset (NSD) and the Natural Object Dataset (NOD).

EAAI Journal 2026 Journal Article

Precise weed identification and differentiated laser weeding strategies for Salvia miltiorrhiza fields based on an enhanced object detection network

  • Xianlin Cao
  • Jinkai Zhang
  • Kaidong Liu
  • Xinyu Wu
  • Yatuan Ma
  • Jifeng Ning
  • Shuqin Yang

Effective weed control is crucial for Salvia miltiorrhiza cultivation, yet traditional methods are often inefficient, costly, or polluting. To address this, this study developed a laser weeding robot based on an improved object detection model capable of identifying weeds and implementing targeted strategies. First, a self-propelled laser weeding robot was constructed for Salvia miltiorrhiza fields to meet operational requirements. Second, a real-world field dataset was established for Salvia miltiorrhiza and five weed families. The detection model, optimized from the You Only Look Once (YOLO) architecture, integrates attention-based feature interaction, dynamic spatial attention, and small object feature enhancement modules. These improvements enhanced the features of small objects, improved occluded target localization, and strengthened similar object discrimination. Third, drawing on weed biological characteristics, a multi-level, differentiated laser weeding strategy was developed to precisely target growth points while ensuring crop safety. Finally, the model and strategy were deployed on the robot to perform real-time detection and intelligent laser weeding. Test results demonstrate the superior performance of the proposed model: the precision of object detection reached 78. 09% (2. 54% over baseline) and that of keypoint detection stood at 80. 69% (8. 22% over baseline). The mean average precision (mAP50) metrics improved to 78. 14% and 80. 56%, representing increases of 2. 34% and 2. 88% respectively. Field tests achieved a 90. 2% weed control rate alongside a low 1. 9% damage rate to Salvia miltiorrhiza. These results validate the system's effectiveness and practicality, providing crucial technical support for intelligent weed management in Salvia miltiorrhiza and other high-value medicinal crops.

JBHI Journal 2025 Journal Article

A Fusion Network With Stacked Denoise Autoencoder and Meta Learning for Lateral Walking Gait Phase Recognition and Multi-Step-Ahead Prediction

  • Wujing Cao
  • Changyu Li
  • Lijun Yang
  • Meng Yin
  • Chunjie Chen
  • Worawarit Kobsiriphat
  • Thanak Utakapan
  • Yizhuang Yang

Lateral walking gait phase recognition and prediction are the premise of hip exoskeleton application in lateral resistance walk exercise. We presented a fusion network with stacked denoise autoencoder and meta learning (SDA-NN-ML) to recognize gait phase and predict gait percentage from IMU signals. Experiments were conducted to detect the four lateral walking gait phases and predict their percentage across different speeds. The performance of SDA-NN-ML and Support Vector Machine (SVM), Adaptive Boosting (AdaBoost) and Long Short Term Memory (LSTM) were evaluated. The cross-subject recognition accuracy of SDA-NN-ML (89. 94%) decreased by 4. 62% compared to the training accuracy, which outperformed SVM (8. 60%), AdaBoost (5. 61%), and LSTM (7. 12%). For real-time and cross-subject prediction of gait phase percentage, the RMSE of SDA-NN-ML (0. 2043) outperformed that of a single regression network (0. 2426). With a signal noise ratio of 100: 30, the cross-subject recognition accuracy decreased by a mere 5. 70%, while the prediction result (RMSE) of SDA-NN-ML increased by 0. 0167 when compared to the noise-free results. SDA-NN-ML demonstrates a stable multi-step-ahead prediction ability with an accuracy higher than 82. 50% and an RMSE of less than 0. 23 when the ahead time is less than 200 ms. The results demonstrated that the proposed method has high accuracy and robust performance in lateral walking gait recognition and prediction.

JBHI Journal 2025 Journal Article

Dual Transformer Network for Predicting Joint Angles and Torques From Multi-Channel EMG Signals in the Lower Limbs

  • Zhuo Wang
  • Chunjie Chen
  • Hui Chen
  • Yizhe Zhou
  • Xiangyang Wang
  • Xinyu Wu

Accurate estimation of lower limb joint kinematics and kinetics using wearable sensors enables biomechanical analysis beyond laboratory settings and facilitates real-time adaptation of exoskeleton assistance profiles. This study introduces a Dual Transformer Network (DTN) designed to concurrently estimate multiple joint angles and moments from multi-channel surface electromyography (sEMG) signals in the lower limbs. The performance evaluation of the predicted joint angles for the hip, knee, and ankle showed average root mean square error ( RMSE ) values of 1. 1827 $^{\circ }$, 1. 4312 $^{\circ }$, and 0. 8113 $^{\circ }$, Pearson correlation coefficients ( $\boldsymbol{\rho }$ ) of 0. 9992, 0. 9993, and 0. 9991, and coefficients of determination ( $\mathbf {{\mathit{R}}}^{2}$ ) of 0. 9847, 0. 9858, and 0. 9838, respectively. For the predicted joint moments, the corresponding values were RMSE of 0. 0458, 0. 0341, and 0. 0522 Nm/kg, $\boldsymbol{\rho }$ of 0. 9978, 0. 9972, and 0. 9990, and $\mathbf {{\mathit{R}}}^{2}$ of 0. 9825, 0. 9801, and 0. 9902. Angular velocities, derived by differentiating the estimated joint angles, achieved an RMSE below 0. 6530 rd/s, $\boldsymbol{\rho }$ exceeding 0. 9534, and $\mathbf {{\mathit{R}}}^{2}$ above 0. 9552. Additionally, joint power, computed as the dot product of predicted joint moments and angular velocities, resulted in RMSE below 0. 3823W/kg, $\boldsymbol{\rho }$ above 0. 9771, and $\mathbf {{\mathit{R}}}^{2}$ above 0. 8925. These results demonstrate the effectiveness of the proposed network in continuously estimating lower limb kinematics and kinetics, contributing to advancements in assist-as-needed exoskeleton control strategies.

TMLR Journal 2025 Journal Article

Multi-Modal Foundation Models for Computational Pathology: A Survey

  • Dong Li
  • Guihong Wan
  • Xintao Wu
  • Xinyu Wu
  • Xiaohui Chen
  • Yi He
  • Zhong Chen
  • Peter K Sorger

Foundation models have emerged as a powerful paradigm in computational pathology (CPath), enabling scalable and generalizable analysis of histopathological images. While early developments centered on uni-modal models trained solely on visual data, recent advances have highlighted the promise of multi-modal foundation models that integrate heterogeneous data sources such as textual reports, structured domain knowledge, and molecular profiles. In this survey, we provide a comprehensive and up-to-date review of multi-modal foundation models in CPath, with a particular focus on models built upon hematoxylin and eosin (H&E) stained whole slide images (WSIs) and tile-level representations. We categorize 34 state-of-the-art multi-modal foundation models into three major paradigms: vision-language, vision-knowledge graph, and vision-gene expression. We further divide vision-language models into non-LLM-based and LLM-based approaches. Additionally, we analyze 30 available multi-modal datasets tailored for pathology, grouped into image-text pairs, instruction datasets, and image-other modality pairs. Our survey also presents a taxonomy of downstream tasks, highlights training and evaluation strategies, and identifies key challenges and future directions. We aim for this survey to serve as a valuable resource for researchers and practitioners working at the intersection of pathology and AI.

AAAI Conference 2024 Conference Paper

A Convolutional Neural Network Interpretable Framework for Human Ventral Visual Pathway Representation

  • Mufan Xue
  • Xinyu Wu
  • Jinlong Li
  • Xuesong Li
  • Guoyuan Yang

Recently, convolutional neural networks (CNNs) have become the best quantitative encoding models for capturing neural activity and hierarchical structure in the ventral visual pathway. However, the weak interpretability of these black-box models hinders their ability to reveal visual representational encoding mechanisms. Here, we propose a convolutional neural network interpretable framework (CNN-IF) aimed at providing a transparent interpretable encoding model for the ventral visual pathway. First, we adapt the feature-weighted receptive field framework to train two high-performing ventral visual pathway encoding models using large-scale functional Magnetic Resonance Imaging (fMRI) in both goal-driven and data-driven approaches. We find that network layer-wise predictions align with the functional hierarchy of the ventral visual pathway. Then, we correspond feature units to voxel units in the brain and successfully quantify the alignment between voxel responses and visual concepts. Finally, we conduct Network Dissection along the ventral visual pathway including the fusiform face area (FFA), and discover variations related to the visual concept of `person'. Our results demonstrate the CNN-IF provides a new perspective for understanding encoding mechanisms in the human ventral visual pathway, and the combination of ante-hoc interpretable structure and post-hoc interpretable approaches can achieve fine-grained voxel-wise correspondence between model and brain. The source code is available at: https://github.com/BIT-YangLab/CNN-IF.

FOCS Conference 2020 Conference Paper

Explicit near-fully X-Ramanujan graphs

  • Ryan O'Donnell
  • Xinyu Wu

Let p(Y1, .. ., Yd, Z1, .. ., Ze) be a self-adjoint noncommutative polynomial, with coefficients from C r×r, in the indeterminates Y1, .. ., Yd (considered to be self-adjoint), the indeterminates Z1, .. ., Ze, and their adjoints Z1*, .. ., Ze*. Suppose Y1, .. ., Yd are replaced by independent random n x n matching matrices, and Z1, .. ., Ze are replaced by independent random n x n permutation matrices. Assuming for simplicity that p's coefficients are 0-1 matrices, the result can be thought of as a kind of random rn-vertex graph G. As n goes to infinity, there will be a natural limiting infinite graph X that covers any finite outcome for G. A recent landmark result of Bordenave and Collins shows that for any, with high probability the spectrum of a random G will be eps-close in Hausdorff distance to the spectrum of X (once the suitably defined “trivial” eigenvalues are excluded). We say that G is “eps-near fully X-Ramanujan”. Our work has two contributions: First we study and clarify the class of infinite graphs X that can arise in this way. Second, we derandomize the Bordenave-Collins result: for any X, we provide explicit, arbitrarily large graphs G that are covered by X and that have (nontrivial) spectrum at Hausdorff distance at most eps from that of X. This significantly generalizes the recent work of Mohanty et al. , which provided explicit near-Ramanujan graphs for every degree d (meaning d-regular graphs with all nontrivial eigenvalues bounded in magnitude by 2sqrt(d-1) + eps). As an application of our main technical theorem, we are also able to determine the “eigenvalue relaxation value” for a wide class of average-case degree-2 constraint satisfaction problems.

SAT Conference 2020 Conference Paper

Mycielski Graphs and PR Proofs

  • Emre Yolcu
  • Xinyu Wu
  • Marijn J. H. Heule

Abstract Mycielski graphs are a family of triangle-free graphs \(M_k\) with arbitrarily high chromatic number. \(M_k\) has chromatic number k and there is a short informal proof of this fact, yet finding proofs of it via automated reasoning techniques has proved to be a challenging task. In this paper, we study the complexity of clausal proofs of the uncolorability of \(M_k\) with \(k-1\) colors. In particular, we consider variants of the \(\mathrm {PR}\) (propagation redundancy) proof system that are without new variables, and with or without deletion. These proof systems are of interest due to their potential uses for proof search. As our main result, we present a sublinear-length and constant-width \(\mathrm {PR}\) proof without new variables or deletion. We also implement a proof generator and verify the correctness of our proof. Furthermore, we consider formulas extended with clauses from the proof until a short resolution proof exists, and investigate the performance of CDCL in finding the short proof. This turns out to be difficult for CDCL with the standard heuristics. Finally, we describe an approach inspired by SAT sweeping to find proofs of these extended formulas.

NeurIPS Conference 2004 Conference Paper

Mistake Bounds for Maximum Entropy Discrimination

  • Philip Long
  • Xinyu Wu

We establish a mistake bound for an ensemble method for classification based on maximizing the entropy of voting weights subject to margin constraints. The bound is the same as a general bound proved for the Weighted Majority Algorithm, and similar to bounds for other variants of Winnow. We prove a more refined bound that leads to a nearly opti- mal algorithm for learning disjunctions, again, based on the maximum entropy principle. We describe a simplification of the on-line maximum entropy method in which, after each iteration, the margin constraints are replaced with a single linear inequality. The simplified algorithm, which takes a similar form to Winnow, achieves the same mistake bounds.

v2026.09.13