Arrow Research search

Author name cluster

Lisha Yu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

4 papers
1 author row

Possible papers

4

AAAI Conference 2026 Conference Paper

UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model

  • Changxin Huang
  • Lv Tang
  • Zhaohuan Zhan
  • Lisha Yu
  • Runhao Zeng
  • Zun Liu
  • Zhengjie Wang
  • Jianqiang Li

Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions—remains highly challenging. Recent research on enhancing language-guided navigation reasoning using pre-trained large language models (LLMs) has shown promising prospects. However, the reasoning of such methods is limited to the linguistic modality, lacking visual reasoning capabilities. Moreover, existing reasoning modules are optimized separately from navigation policies, leading to incompatibility and potential conflicts in optimization objectives. To tackle these challenges, we introduce UNeMo, a novel framework designed for the collaborative optimization of visual state reasoning and navigational decision-making. It introduces a Multimodal World Model (MWM) that takes visual features, language instructions, and navigational actions as inputs to jointly predict subsequent visual states, enabling cross-modal reasoning. Via a Hierarchical Prediction-Feedback (HPN) mechanism, MWM collaborates with navigation policies: the first layer generates actions using current vision-and-language features; MWM then infers post-action visual states to guide the second layer’s fine-grained decisions. This forms a dynamic bidirectional promotion mechanism where MWM reasoning optimizes navigation policies, while policy decisions feedback to improve MWM’s reasoning accuracy. Experiments on R2R and REVERIE datasets show UNeMo outperforms state-of-the-art methods by 2.1% and 0.7% in navigation accuracy for unseen scenes, validating its effectiveness.

JBHI Journal 2025 Journal Article

A Novel Dynamic Latent Variables-Based Framework for Enhancing Freezing of Gait Detection in Parkinson's Disease Patients

  • Xuan Wang
  • Lisha Yu
  • S. Joe Qin
  • Yang Zhao

Freezing of Gait (FOG) is one of the most severe symptoms of Parkinson's disease (PD), which often lead to life-threatening falls. Wearable sensor-based technologies coupled with data driven methods have advanced the detection of FOG in a timely fashion. However, most existing monitoring methods overlook the dynamics of processes when extracting effective information from high-dimensional sensor data. To tackle these problems, we develop a novel framework for FOG detection by integrating Dynamic Latent Variable (DLV)-based dimensionality reduction strategies and personalized monitoring. First, a multi-channel sliding window mechanism is adopted to extract the multiple potentially effective feature sequences. Second, an interpretable DLV-based method incorporating time-lagged terms is designed for the subspace representation of complex high-dimensional sequences. Third, the extracted DLVs are integrated with threshold-based methods or the Statistical Process Control (SPC) method for anomaly detection. We identified distinct variations in gait patterns among individuals, underscoring the importance of personalized approaches. The proposed framework demonstrates its effectiveness in FOG detection via validating on real world dataset, achieving a sensitivity of $\mathbf {0. 845} \pm \mathbf {0. 254}$ and a specificity of $\mathbf {0. 842} \pm \mathbf {0. 211}$.

JBHI Journal 2025 Journal Article

MSTG-Transformer: Multivariate Spatial-Temporal Gated Transformer Model for 3D Skeleton Data-based Fall Risk Prediction

  • Junjie Cao
  • Xuan Wang
  • Keyi Huang
  • Lisha Yu
  • Xiaomao Fan
  • Yang Zhao

As the aging population continues to grow, falls among older adults have become a significant public health concern worldwide. Data-driven approaches for effective fall risk prediction, which integrate standard functional tests with 3D skeleton data from depth sensors, are gaining increasing attention. However, the complex physiological and functional interactions among skeletal keypoints during ambulation pose challenges for multidimensional feature extraction in most predictive models. In this study, we developed a novel approach based on preprocessed 3D skeleton data, named Multivariate SpatialTemporal Gated Transformer (MSTG-Transformer). This approach consists of three main stages. First, gait cycle sequences are constructed to sophisticatedly depict the movement patterns of subjects, amplifying the distinctions between groups. Then, spatial and topological features are extracted via convolutional modules, and a dual-stream encoder block is employed to encode the features of 3D skeleton data across both time steps and time channels. Finally, a voting scheme is used to determine fall risk by integrating the classification results of individual gait cycle segments. Validation experiments on a real-world dataset demonstrate that our proposed approach outperforms classical methods, achieving a superior prediction accuracy of 0. 9510 ± 0. 0240. Additionally, our study highlights the crucial role of potential interactions between skeletal keypoints in accurately predicting fall risk

JBHI Journal 2024 Journal Article

Sensor-Based Multifaceted Feature Extraction and Ensemble Elastic Net Approach for Assessing Fall Risk in Community-Dwelling Older Adults

  • Xuan Wang
  • Lisha Yu
  • Hailiang Wang
  • Kwok Leung Tsui
  • Yang Zhao

Accurate identification of community-dwelling older adults at high fall risk can facilitate timely intervention and significantly reduce fall incidents. Analyzing gait and balance capabilities via feature extraction and modeling through sensor-based motion data has emerged as a viable approach for fall risk assessment. However, the existing approaches for extracting key features related to fall risk lack inclusiveness, with limited consideration of the non-linear characteristics of sensor signals, such as signal complexity, self-similarity, and local stability. In this study, we developed a multifaceted feature extraction scheme employing diverse feature types, including demographic, descriptive statistical, non-linear, spatiotemporal and spectral features, derived from three-axis accelerometers and gyroscope data. This study is the first attempt to investigate non-linear features related to fall risk in multi-task scenarios from a dynamic system perspective. Based on the extracted multifaceted features, we propose an ensemble elastic net (E-E-N) approach for handling imbalanced data and offering high model interpretability. The E-E-N utilizes bootstrap sampling to construct base classifiers and employs a weighting mechanism to aggregate the base classifiers. We conducted a set of validation experiments using real-world data for comprehensive comparative analysis. The results demonstrate that the E-E-N approach exhibits superior predictive performance on fall risk classification. Our proposed approach offers a cost-effective tool for accurately assessing fall risk and alleviating the burden of continuous health monitoring in the long term.

v2026.09.13