Arrow Research search

Author name cluster

Yan Ma

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
2 author rows

Possible papers

8

EAAI Journal 2025 Journal Article

A personalized human–machine cooperative approach with transformer-based recognition for longitudinal and lateral control of intelligent vehicles

  • Yan Ma
  • Jingjing Xie
  • Liang He
  • Kailong Zhang
  • Xiongmei Zeng
  • Quan Ouyang
  • Danwei Wang

Human–machine interaction brings challenges for vehicle control design due to individual differences, a personalized cooperative approach with driving style recognition is proposed to achieve lateral and longitudinal control of intelligent vehicles in this paper. An improved Transformer-based method with an unsupervised pre-training and window-based multi-head self-attention is proposed to enhance the recognition accuracy and speed of driving styles, and thereby to capture the controller parameters under various driving styles. To achieve the lateral and longitudinal control of human-machine cooperative system, an integrated driver–vehicle model is established by considering driving styles and vehicle planar dynamics. Then, a Takagi–Sugeno fuzzy controller is developed to handle time-varying parameters and eliminate human-machine conflicts. Especially, stability conditions are exploited by Lyapunov arguments to achieve the control objective. Finally, simulation results show that the designed Transformer-based method has better classification accuracy and computational efficiency than other baselines on the same dataset. Based on recognition results, the designed controller can effectively improve the driving performance under various driving styles and time-varying parameters compared with other methods.

NeurIPS Conference 2024 Conference Paper

MMLONGBENCH-DOC: Benchmarking Long-context Document Understanding with Visualizations

  • Yubo Ma
  • Yuhang Zang
  • Liangyu Chen
  • Meiqi Chen
  • Yizhu Jiao
  • Xinze Li
  • Xinyuan Lu
  • Ziyu Liu

Understanding documents with rich layouts and multi-modal components is a long-standing and practical task. Recent Large Vision-Language Models (LVLMs) have made remarkable strides in various tasks, particularly in single-page document understanding (DU). However, their abilities on long-context DU remain an open problem. This work presents MMLONGBENCH-DOC, a long-context, multi- modal benchmark comprising 1, 082 expert-annotated questions. Distinct from previous datasets, it is constructed upon 135 lengthy PDF-formatted documents with an average of 47. 5 pages and 21, 214 textual tokens. Towards comprehensive evaluation, answers to these questions rely on pieces of evidence from (1) different sources (text, image, chart, table, and layout structure) and (2) various locations (i. e. , page number). Moreover, 33. 7\% of the questions are cross-page questions requiring evidence across multiple pages. 20. 6\% of the questions are designed to be unanswerable for detecting potential hallucinations. Experiments on 14 LVLMs demonstrate that long-context DU greatly challenges current models. Notably, the best-performing model, GPT-4o, achieves an F1 score of only 44. 9\%, while the second-best, GPT-4V, scores 30. 5\%. Furthermore, 12 LVLMs (all except GPT-4o and GPT-4V) even present worse performance than their LLM counterparts which are fed with lossy-parsed OCR documents. These results validate the necessity of future research toward more capable long-context LVLMs.

NeurIPS Conference 2024 Conference Paper

OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

  • Zhen Huang
  • Zengzhi Wang
  • Shijie Xia
  • Xuefeng Li
  • Haoyang Zou
  • Ruijie Xu
  • Run-Ze Fan
  • Lyumanshan Ye

The evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showcasing potential cognitive reasoning abilities in problem-solving and scientific discovery (i. e. , AI4Science) once exclusive to human intellect. To comprehensively evaluate current models' performance in cognitive reasoning abilities, we introduce OlympicArena, which includes 11, 163 bilingual problems across both text-only and interleaved text-image modalities. These challenges encompass a wide range of disciplines spanning seven fields and 62 international Olympic competitions, rigorously examined for data leakage. We argue that the challenges in Olympic competition problems are ideal for evaluating AI's cognitive reasoning due to their complexity and interdisciplinary nature, which are essential for tackling complex scientific challenges and facilitating discoveries. Beyond evaluating performance across various disciplines using answer-only criteria, we conduct detailed experiments and analyses from multiple perspectives. We delve into the models' cognitive reasoning abilities, their performance across different modalities, and their outcomes in process-level evaluations, which are vital for tasks requiring complex reasoning with lengthy solutions. Our extensive evaluations reveal that even advanced models like GPT-4o only achieve a 39. 97\% overall accuracy (28. 67\% for mathematics and 29. 71\% for physics), illustrating current AI limitations in complex reasoning and multimodal integration. Through the OlympicArena, we aim to advance AI towards superintelligence, equipping it to address more complex challenges in science and beyond. We also provide a comprehensive set of resources to support AI research, including a benchmark dataset, an open-source annotation platform, a detailed evaluation tool, and a leaderboard with automatic submission features.

AAAI Conference 2023 Conference Paper

Open-Ended Diverse Solution Discovery with Regulated Behavior Patterns for Cross-Domain Adaptation

  • Kang Xu
  • Yan Ma
  • Bingsheng Wei
  • Wei Li

While Reinforcement Learning can achieve impressive results for complex tasks, the learned policies are generally prone to fail in downstream tasks with even minor model mismatch or unexpected perturbations. Recent works have demonstrated that a policy population with diverse behavior characteristics can generalize to downstream environments with various discrepancies. However, such policies might result in catastrophic damage during the deployment in practical scenarios like real-world systems due to the unrestricted behaviors of trained policies. Furthermore, training diverse policies without regulation of the behavior can result in inadequate feasible policies for extrapolating to a wide range of test conditions with dynamics shifts. In this work, we aim to train diverse policies under the regularization of the behavior patterns. We motivate our paradigm by observing the inverse dynamics in the environment with partial state information and propose Diversity in Regulation (DiR) training diverse policies with regulated behaviors to discover desired patterns that benefit the generalization. Considerable empirical results on various variations of different environments indicate that our method attains improvements over other diversity-driven counterparts.

JBHI Journal 2022 Journal Article

Non-Invasive Glucose Metabolism Quantification Method Based on Unilateral ICA Image Derived Input Function by Hybrid PET/MR in Ischemic Cerebrovascular Disease

  • Min Wang
  • Bixiao Cui
  • Yi Shan
  • Hongwei Yang
  • Zhuangzhi Yan
  • Lalith Kumar Shiyam Sundar
  • Ian Alberts
  • Axel Rominger

The non-invasive quantification of the cerebral metabolic rate for glucose (CMRGlc) and the characterization of cerebral metabolism in the cerebrovascular territories are helpful in understanding ischemic cerebrovascular disease (ICVD). Firstly, we investigated a non-invasive quantification approach based on an image-derived input function (IDIF) in ICVD. Second, we studied the metabolic changes in CMRGlc after surgical intervention. We evaluated the hypothesis that the IDIF method based on the unilateral internal carotid artery could address challenges in ICVD quantification. The CMRGlc and standardized uptake value ratio (SUVR) were used to measure glucose metabolism activity. Healthy controls showed no significant differences in CMRGlc values between bilateral and unilateral IDIF measurements (intraclass correlation coefficient [ICC]: 0. 91–0. 98). Patients with ICVD showed significantly increased CMRGlc values after surgical intervention for all territories (percentage changes: 7. 4%–22. 5%). In contrast, SUVR showed minor differences between postoperative and preoperative patients, indicating that it was a poor biomarker for the diagnosis of ICVD. A significant association between CMRGlc and the National Institutes of Health Stroke Scale (NIHSS) scores was observed ( r =-0. 54). Our findings suggested that IDIF could be a valuable tool for CMRGlc quantification in patients with ICVD and may advance personalized precision interventions.

YNICL Journal 2013 Journal Article

Comparison of randomized multifocal mapping and temporal phase mapping of visual cortex for clinical use

  • Yan Ma
  • B. Douglas Ward
  • Kristina M. Ropella
  • Edgar A. DeYoe

fMRI is becoming an important clinical tool for planning and guidance of surgery to treat brain tumors, arteriovenous malformations, and epileptic foci. For visual cortex mapping, the most popular paradigm by far is temporal phase mapping, although random multifocal stimulation paradigms have drawn increased attention due to their ability to identify complex response fields and their random properties. In this study we directly compared temporal phase and multifocal vision mapping paradigms with respect to clinically relevant factors including: time efficiency, mapping completeness, and the effects of noise. Randomized, multifocal mapping accurately decomposed the response of single voxels to multiple stimulus locations and made correct retinotopic assignments as noise levels increased despite decreasing sensitivity. Also, multifocal mapping became less efficient as the number of stimulus segments (locations) increased from 13 to 25 to 49 and when duty cycle was increased from 25% to 50%. Phase mapping, on the other hand, activated more extrastriate visual areas, was more time efficient in achieving statistically significant responses, and had better sensitivity as noise increased, though with an increase in systematic retinotopic mis-assignments. Overall, temporal phase mapping is likely to be a better choice for routine clinical applications though random multifocal mapping may offer some unique advantages for selected applications.

YNIMG Journal 2012 Journal Article

Adaptive Kalman filtering for real-time mapping of the visual field

  • B. Douglas Ward
  • John Janik
  • Yousef Mazaheri
  • Yan Ma
  • Edgar A. DeYoe

This paper demonstrates the feasibility of real-time mapping of the visual field for clinical applications. Specifically, three aspects of this problem were considered: (1) experimental design, (2) statistical analysis, and (3) display of results. Proper experimental design is essential to achieving a successful outcome, particularly for real-time applications. A random-block experimental design was shown to have less sensitivity to measurement noise, as well as greater robustness to error in modeling of the hemodynamic impulse response function (IRF) and greater flexibility than common alternatives. In addition, random encoding of the visual field allows for the detection of voxels that are responsive to multiple, not necessarily contiguous, regions of the visual field. Due to its recursive nature, the Kalman filter is ideally suited for real-time statistical analysis of visual field mapping data. An important feature of the Kalman filter is that it can be used for nonstationary time series analysis. The capability of the Kalman filter to adapt, in real time, to abrupt changes in the baseline arising from subject motion inside the scanner and other external system disturbances is important for the success of clinical applications. The clinician needs real-time information to evaluate the success or failure of the imaging run and to decide whether to extend, modify, or terminate the run. Accordingly, the analytical software provides real-time displays of (1) brain activation maps for each stimulus segment, (2) voxel-wise spatial tuning profiles, (3) time plots of the variability of response parameters, and (4) time plots of activated volume.

IROS Conference 2011 Conference Paper

Coordinated landing of a quadrotor on a skid-steered ground vehicle in the presence of time delays

  • John Michael Daly
  • Yan Ma
  • Steven L. Waslander

This work presents a control technique to autonomously coordinate a landing between a quadrotor UAV and a skid-steered UGV. Local controllers to feedback linearize the models are presented, and a joint decentralized controller is developed to coordinate a rendezvous for the two vehicles. The effects of time delays on closed loop stability are examined using a Retarded Functional Differential Equation (RFDE) formulation of the problem, and delay margins are determined for particular closed loop setups. Simulation results are presented, which demonstrate the feasibility of this approach for autonomous outdoor coordinated landing.

v2026.09.13