Arrow Research search

Author name cluster

Qin Ni

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

5 papers
1 author row

Possible papers

5

AAAI Conference 2026 Conference Paper

MovieGraph-ToM: Evaluating Long-Range Theory of Mind in Large Language Models via Implicit Social-Causal Graphs

  • Tingjiang Wei
  • Qin Ni
  • Rong Gao
  • Yingying Wang
  • Liang He

The capacity for social reasoning, particularly Theory of Mind (ToM), is a foundational prerequisite for aligning Large Language Models (LLMs) with human values. However, current evaluations are predominantly confined to simplistic, short-text scenarios, obscuring their true capabilities and potential failure modes in complex, long-range social dynamics. To address this deficit, we introduce MovieGraph-ToM, a large-scale benchmark for evaluating long-range ToM and social cognition within extended, multimodal narratives. We employ a "scaffold-and-probe" methodology: we construct a ground-truth Social-Causal Graph offline, which maps the narrative's latent mental states and causal chains. During evaluation, the model is denied access to this graph and must reason directly from raw multimodal inputs. This decoupling forces genuine inference over superficial pattern matching. Reasoning is probed via a hierarchical questioning framework designed to differentiate spontaneous understanding from logical robustness. Our empirical results reveal systematic vulnerabilities in even state-of-the-art models. We identify a critical "multiple-choice pitfall," where accuracy plummets against well-crafted distractors, and a stark "generative-discriminative divide," where models fail to construct coherent explanations for answers they correctly identify. These findings highlight a latent risk, as models that feign comprehension could lead to unpredictable and misaligned behaviors. MovieGraph-ToM thus offers a rigorous platform for assessing and advancing the robust social intelligence required for safely aligned AI systems.

AAAI Conference 2026 Conference Paper

Shaping Human–AI Collaboration in Education: Effects of AI-Assisted Decision-Making Paradigms and Human–AI Decision Consistency on Pre-Service Teachers’ Psychological States and Performance

  • Yingying Wang
  • Qin Ni
  • Haoxin Xu
  • Jiaqi Yin
  • Tingjiang Wei

Artificial intelligence is playing an increasingly important role in supporting decision-making, particularly in educational contexts, where it serves as a critical tool to assist teacher judgment and optimize instructional decisions. However, limited research has examined how different AI-assisted decision-making paradigms influence the Performance of human-AI collaboration, as well as the underlying psychological mechanisms and causal pathways. Therefore, this study investigated 59 pre-service teachers to examine how AI-assisted decision-making paradigms and human-AI consistency influenced their psychological states and task performance. Specifically, this study employed a two-factor mixed experimental design, with the AI-assisted decision-making paradigms as the between-subjects factor and human-AI consistency as the within-subjects factor. Data were analyzed using the Bayesian cumulative link mixed model and structural equation modeling. The results reveal that AI-assisted decision-making paradigms do not have a significant direct effect on task performance. However, when the moderating role of human-AI decision consistency is taken into account, the effect of AI-assisted decision-making paradigms on task performance can exert its influence indirectly through a sequential psychological pathway involving users’ confidence and their trust in the AI. Consistency between human and AI decisions not only significantly enhances users’ trust in AI, confidence, and task performance, but the proportion of consistent decisions also significantly moderates the impact of AI-assisted decision-making paradigms on users’ confidence levels. Notably, our findings indicate that users maintain a moderately level of trust in AI even when their decisions diverge from those of AI. In summary, this study highlights the mediating mechanism by which AI-assisted decision-making paradigms influence task performance through psychological states and identifies the moderating role of human-AI consistency in this pathway. These findings advance the theoretical understanding of human-AI interaction models in educational contexts and offer mechanistic insights to guide the optimization of instructional AI systems.

YNIMG Journal 2025 Journal Article

Screening tools for subjective cognitive decline and mild cognitive impairment based on task-state prefrontal functional connectivity: a functional near-infrared spectroscopy study

  • Zhengping Pu
  • Hongna Huang
  • Man Li
  • Hongyan Li
  • Xiaoyan Shen
  • Lizhao Du
  • Qingfeng Wu
  • Xiaomei Fang

BACKGROUND: Subjective cognitive decline (SCD) and mild cognitive impairment (MCI) carry the risk of progression to dementia, and accurate screening methods for these conditions are urgently needed. Studies have suggested the potential ability of functional near-infrared spectroscopy (fNIRS) to identify MCI and SCD. The present fNIRS study aimed to develop an early screening method for SCD and MCI based on activated prefrontal functional connectivity (FC) during the performance of cognitive scales and subject-wise cross-validation via machine learning. METHODS: Activated prefrontal FC data measured by fNIRS were collected from 55 normal controls, 80 SCD patients, and 111 MCI patients. Differences in FC were analyzed among the groups, and FC strength and cognitive scale performance were extracted as features to build classification and predictive models through machine learning. Model performance was assessed based on accuracy, specificity, sensitivity, and area under the curve (AUC) with 95 % confidence interval (CI) values. RESULTS: Statistical analysis revealed a trend toward more impaired prefrontal FC with declining cognitive function. Prediction models were built by combining features of prefrontal FC and cognitive scale performance and applying machine learning models, The models showed generally satisfactory abilities to differentiate among the three groups, especially those employing linear discriminant analysis, logistic regression, and support vector machine. Accuracies of 92.0 % for MCI vs. NC, 80.0 % for MCI vs. SCD, and 76.1 % for SCD vs. NC were achieved, and the highest AUC values were 97.0 % (95 % CI: 94.6 %-99.3 %) for MCI vs. NC, 87.0 % (95 % CI: 81.5 %-92.5 %) for MCI vs. SCD, and 79.2 % (95 % CI: 71.0 %-87.3 %) for SCD vs. NC. CONCLUSION: The developed screening method based on fNIRS and machine learning has the potential to predict early-stage cognitive impairment based on prefrontal FC data collected during cognitive scale-induced activation.

TIST Journal 2025 Journal Article

The Social Cognition Ability Evaluation of LLMs: A Dynamic Gamified Assessment and Hierarchical Social Learning Measurement Approach

  • Qin Ni
  • Yangze Yu
  • Yiming Ma
  • Xin Lin
  • Ciping Deng
  • Tingjiang Wei
  • Mo Xuan

Large Language Model (LLM) has shown amazing abilities in reasoning tasks, theory of mind (ToM) has been tested in many studies as part of reasoning tasks, and social learning, which is closely related to ToM, is still lack of investigation. However, the test methods and materials make the test results unconvincing. We propose a dynamic gamified assessment (DGA) and hierarchical social learning measurement to test ToM and social learning capacities in LLMs. The test for ToM consists of five parts. First, we extract ToM tasks from ToM experiments and then design game rules to satisfy the ToM task requirement. After that, we design ToM questions to match the game’s rules and use these to generate test materials. Finally, we go through the above steps to test the model. To assess the social learning ability, we introduce a novel set of social rules (three in total). Experiment results demonstrate that, except GPT-4, LLMs performed poorly on the ToM test but showed a certain level of social learning ability in social learning measurement.

AAAI Conference 2024 Conference Paper

BDIQA: A New Dataset for Video Question Answering to Explore Cognitive Reasoning through Theory of Mind

  • Yuanyuan Mao
  • Xin Lin
  • Qin Ni
  • Liang He

As a foundational component of cognitive intelligence, theory of mind (ToM) can make AI more closely resemble human thought processes, thereby enhancing their interaction and collaboration with human. In particular, it can significantly improve a model's comprehension of videos in complex scenes. However, current video question answer (VideoQA) datasets focus on studying causal reasoning within events, few of them genuinely incorporating human ToM. Consequently, there is a lack of development in ToM reasoning tasks within the area of VideoQA. This paper presents BDIQA, the first benchmark to explore the cognitive reasoning capabilities of VideoQA models in the context of ToM. BDIQA is inspired by the cognitive development of children's ToM and addresses the current deficiencies in machine ToM within datasets and tasks. Specifically, it offers tasks at two difficulty levels, assessing Belief, Desire and Intention (BDI) reasoning in both simple and complex scenarios. We conduct evaluations on several mainstream methods of VideoQA and diagnose their capabilities with zero-shot, few-shot and supervised learning. We find that the performance of pre-trained models on cognitive reasoning tasks remains unsatisfactory. To counter this challenge, we undertake thorough analysis and experimentation, ultimately presenting two guidelines to enhance cognitive reasoning derived from ablation analysis.

v2026.09.13