Arrow Research search

Author name cluster

Rui Dai

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
2 author rows

Possible papers

9

AAAI Conference 2026 Conference Paper

Do Retrieval Augmented Language Models Know When They Don’t Know?

  • Youchao Zhou
  • Heyan Huang
  • Yicheng Liu
  • Rui Dai
  • Xinglin Wang
  • Xingchen Zhang
  • Shumin Shi
  • Yang Deng

Existing large language models (LLMs) occasionally generate plausible yet factually incorrect responses, known as hallucinations. Two main approaches have been proposed to mitigate hallucinations: retrieval-augmented language models (RALMs) and refusal post-training. However, current research predominantly focuses on their individual effectiveness while overlooking the evaluation of the refusal capability of RALMs. Ideally, if RALMs know when they do not know, they should refuse to answer. In this study, we ask the fundamental question: Do RALMs know when they don’t know? Specifically, we investigate three questions. First, are RALMs well calibrated with respect to different internal and external knowledge states? We examine the influence of various factors. Contrary to expectations, when all retrieved documents are irrelevant, RALMs still tend to refuse questions they could have answered correctly. Next, given the model's pronounced over-refusal behavior, we raise a second question: How does a RALM's refusal ability align with its calibration quality? Our results show that the over-refusal problem can be mitigated through in-context fine-tuning. However, we observe that improved refusal behavior does not necessarily imply better calibration or higher overall accuracy. Finally, we ask: Can we combine refusal-aware RALMs with uncertainty-based answer abstention to mitigate over-refusal? We develop a simple yet effective refusal mechanism for refusal-post-trained RALMs that improves their overall answer quality by balancing refusal and correct answers. Our study provides a more comprehensive understanding of the factors influencing RALM behavior. Meanwhile, we emphasize that uncertainty estimation for RALMs remains an open problem deserving deeper investigation.

ICLR Conference 2025 Conference Paper

Leveraging Submodule Linearity Enhances Task Arithmetic Performance in LLMs

  • Rui Dai
  • Sile Hu
  • Xu Shen 0001
  • Yonggang Zhang 0003
  • Xinmei Tian 0001
  • Jieping Ye

Task arithmetic is a straightforward yet highly effective strategy for model merging, enabling the resultant model to exhibit multi-task capabilities. Recent research indicates that models demonstrating linearity enhance the performance of task arithmetic. In contrast to existing methods that rely on the global linearization of the model, we argue that this linearity already exists within the model's submodules. In particular, we present a statistical analysis and show that submodules (e.g., layers, self-attentions, and MLPs) exhibit significantly higher linearity than the overall model. Based on these findings, we propose an innovative model merging strategy that independently merges these submodules. Especially, we derive a closed-form solution for optimal merging weights grounded in the linear properties of these submodules. Experimental results demonstrate that our method consistently outperforms the standard task arithmetic approach and other established baselines across different model scales and various tasks. This result highlights the benefits of leveraging the linearity of submodules and provides a new perspective for exploring solutions for effective and practical multi-task model merging.

IROS Conference 2025 Conference Paper

RoboNurse-VLA: Robotic Scrub Nurse System based on Vision-Language-Action Model

  • Shunlei Li
  • Jin Wang
  • Rui Dai
  • Wanyu Ma
  • Wing Yin Ng
  • Yingbai Hu
  • Zheng Li 0012

In modern healthcare, the demand for autonomous robotic assistants has grown significantly, particularly in the operating room, where surgical tasks require precision and reliability. Robotic scrub nurses have emerged as a promising solution to improve efficiency and reduce human error during surgery. However, challenges remain in terms of accurately grasping and handing over surgical instruments, especially when dealing with complex objects in dynamic environments. In this work, we introduce RoboNurse-VLA, a novel robotic scrub nurse system based on a Vision-Language-Action (VLA) model. RoboNurse-VLA integrates Segment Anything Model 2 (SAM 2) and Llama 2, leveraging an LLM head to enhance reasoning capabilities. By combining SAM 2’s mask generation with Llama 2’s advanced reasoning, RoboNurse-VLA can accurately interpret task requirements, identify optimal grasping points, and determine appropriate handover poses. Designed for real-time operation, RoboNurse-VLA enables precise grasping and seamless handover of surgical instruments based on voice commands from the surgeon. Utilizing state-of-the-art vision and language models, it effectively addresses challenges related to object detection, pose optimization, and handling difficult-to-grasp instruments. Extensive evaluations demonstrate that RoboNurse-VLA outperforms existing models, achieving high success rates in surgical instrument handovers, even for previously unseen tools and complex objects. This work represents a significant advancement in autonomous surgical assistance, highlighting the potential of VLA models for real-world medical applications. More details can be found at https:// robonurse-vla.github.io.

YNIMG Journal 2023 Journal Article

Classical and non-classical psychedelic drugs induce common network changes in human cortex

  • Rui Dai
  • Tony E. Larkin
  • Zirui Huang
  • Vijay Tarnal
  • Paul Picton
  • Phillip E. Vlisides
  • Ellen Janke
  • Amy McKinney

The neurobiology of the psychedelic experience is not fully understood. Identifying common brain network changes induced by both classical (i. e. , acting at the 5-HT2 receptor) and non-classical psychedelics would provide mechanistic insight into state-specific characteristics. We analyzed whole-brain functional connectivity based on resting-state fMRI data in humans, acquired before and during the administration of nitrous oxide, ketamine, and lysergic acid diethylamide. We report that, despite distinct molecular mechanisms and modes of delivery, all three psychedelics reduced within-network functional connectivity and enhanced between-network functional connectivity. More specifically, all three drugs increased connectivity between right temporoparietal junction and bilateral intraparietal sulcus as well as between precuneus and left intraparietal sulcus. These regions fall within the posterior cortical “hot zone, ” posited to mediate the qualitative aspects of experience. Thus, both classical and non-classical psychedelics modulate networks within an area of known relevance for consciousness, identifying a biologically plausible candidate for their subjective effects.

ICML Conference 2023 Conference Paper

Moderately Distributional Exploration for Domain Generalization

  • Rui Dai
  • Yonggang Zhang 0003
  • Zhen Fang 0001
  • Bo Han 0003
  • Xinmei Tian 0001

Domain generalization (DG) aims to tackle the distribution shift between training domains and unknown target domains. Generating new domains is one of the most effective approaches, yet its performance gain depends on the distribution discrepancy between the generated and target domains. Distributionally robust optimization is promising to tackle distribution discrepancy by exploring domains in an uncertainty set. However, the uncertainty set may be overwhelmingly large, leading to low-confidence prediction in DG. It is because a large uncertainty set could introduce domains containing semantically different factors from training domains. To address this issue, we propose to perform a $\textit{mo}$derately $\textit{d}$istributional $\textit{e}$xploration (MODE) for domain generalization. Specifically, MODE performs distribution exploration in an uncertainty $\textit{subset}$ that shares the same semantic factors with the training domains. We show that MODE can endow models with provable generalization performance on unknown target domains. The experimental results show that MODE achieves competitive performance compared to state-of-the-art baselines.

YNIMG Journal 2022 Journal Article

Early visual exposure primes future cross-modal specialization of the fusiform face area in tactile face processing in the blind

  • Rui Dai
  • Zirui Huang
  • Xuchu Weng
  • Sheng He

The fusiform face area (FFA) is a core cortical region for face information processing. Evidence suggests that its sensitivity to faces is largely innate and tuned by visual experience. However, how experience in different time windows shape the plasticity of the FFA remains unclear. In this study, we investigated the role of visual experience at different time points of an individual's early development in the cross-modal face specialization of the FFA. Participants (n = 74) were classified into five groups: congenital blind, early blind, late blind, low vision, and sighted control. Functional magnetic resonance imaging data were acquired when the participants haptically processed carved faces and other objects. Our results showed a robust and highly consistent face-selective activation in the FFA region in the early blind participants, invariant to size and level of abstraction of the face stimuli. The cross-modal face activation in the FFA was much less consistent in other groups. These results suggest that early visual experience primes cross-modal specialization of the FFA, and even after the absence of visual experience for more than 14 years in early blind participants, their FFA can engage in cross-modal processing of face information.

AIIM Journal 2019 Journal Article

Classifying medical relations in clinical text via convolutional neural networks

  • Bin He
  • Yi Guan
  • Rui Dai

Deep learning research on relation classification has achieved solid performance in the general domain. This study proposes a convolutional neural network (CNN) architecture with a multi-pooling operation for medical relation classification on clinical records and explores a loss function with a category-level constraint matrix. Experiments using the 2010 i2b2/VA relation corpus demonstrate these models, which do not depend on any external features, outperform previous single-model methods and our best model is competitive with the existing ensemble-based method.

YNIMG Journal 2016 Journal Article

Decoupled temporal variability and signal synchronization of spontaneous brain activity in loss of consciousness: An fMRI study in anesthesia

  • Zirui Huang
  • Jun Zhang
  • Jinsong Wu
  • Pengmin Qin
  • Xuehai Wu
  • Zhiyao Wang
  • Rui Dai
  • Yuan Li

Two aspects of the low frequency fluctuations of spontaneous brain activity have been proposed which reflect the complex and dynamic features of resting-state activity, namely temporal variability and signal synchronization. The relationship between them, especially its role in consciousness, nevertheless remains unclear. Our study examined the temporal variability and signal synchronization of spontaneous brain activity, as well as their relationship during loss of consciousness. We applied an intra-subject design of resting-state functional magnetic resonance imaging (rs-fMRI) in two conditions: during wakefulness, and under anesthesia with clinical unconsciousness. In addition, an independent group of patients with disorders of consciousness (DOC) was included in order to test the reliability of our findings. We observed a global reduction in the temporal variability, local and distant brain signal synchronization for subjects during anesthesia. Importantly, we found a link between temporal variability and both local and distant signal synchronizations during wakefulness: the higher the degree of temporal variability, the higher its intra-regional homogeneity and inter-regional functional connectivity. In contrast, this link was broken down under anesthesia, implying a decoupling between temporal variability and signal synchronization; this decoupling was reproduced in patients with DOC. Our results suggest that there exist some as yet unclear physiological mechanisms of consciousness which “couple” the two mathematically independent measures, temporal variability and signal synchronization of spontaneous brain activity. Our findings not only extend our current knowledge of the neural correlates of anesthetic-induced unconsciousness, but have implications for both computational neural modeling and clinical practice, such as in the diagnosis of loss of consciousness in patients with DOC.

YNIMG Journal 2014 Journal Article

Using fMRI to decode true thoughts independent of intention to conceal

  • Zhi Yang
  • Zirui Huang
  • Javier Gonzalez-Castillo
  • Rui Dai
  • Georg Northoff
  • Peter Bandettini

Multi-variate pattern analysis (MVPA) applied to BOLD-fMRI has proven successful at decoding complicated fMRI signal patterns associated with a variety of cognitive processes. One cognitive process, not yet investigated, is the mental representation of “Yes/No” thoughts that precede the actual overt response to a binary “Yes/No” question. In this study, we focus on examining: (1) whether spatial patterns of the hemodynamic response carry sufficient information to allow reliable decoding of “Yes/No” thoughts; and (2) whether decoding of “Yes/No” thoughts is independent of the intention to respond honestly or dishonestly. To achieve this goal, we conducted two separate experiments. Experiment 1, collected on a 3T scanner, examined the whole brain to identify regions that carry sufficient information to permit significantly above-chance prediction of “Yes/No” thoughts at the group level. In Experiment 2, collected on a 7T scanner, we focused on the regions identified in Experiment 1 to examine the capability of achieving high decoding accuracy at the single subject level. A set of regions – namely right superior temporal gyrus, left supra-marginal gyrus, and left middle frontal gyrus – exhibited high decoding power. Decoding accuracy for these regions increased with trial averaging. When 18 trials were averaged, the median accuracies were 82. 5%, 77. 5%, and 79. 5%, respectively. When trials were separated according to deceptive intentions (set via experimental cues), and classifiers were trained on honest trials, but tested on trials where subjects were asked to deceive, the median accuracies of these regions still reached 66%, 75%, and 78. 5%. These results provide evidence that concealed “Yes/No” thoughts are encoded in the BOLD signal, retaining some level of independence from the subject’s intentions to answer honestly or dishonestly. These findings also suggest the theoretical possibility for more efficient brain-computer interfaces where subjects only need to think their answers to communicate.

v2026.09.13