Arrow Research search

Author name cluster

Lei Ji

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

14 papers
1 author row

Possible papers

14

NeurIPS Conference 2025 Conference Paper

PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning

  • Yizhen Zhang
  • Yang Ding
  • Shuoshuo Zhang
  • Xinchen Zhang
  • Haoling Li
  • Zhong-Zhi Li
  • Peijie Wang
  • Jie Wu

Inspired by the impressive reasoning capabilities demonstrated by reinforcement learning approaches like DeepSeek-R1, recent emerging research has begun exploring the use of reinforcement learning (RL) to enhance vision-language models (VLMs) for multimodal reasoning tasks. However, most existing multimodal reinforcement learning approaches remain limited to spatial reasoning within single-image contexts, yet still struggle to generalize to more complex and real-world scenarios involving multi-image positional reasoning, where understanding the relationships across images is crucial. To address this challenge, we propose a general reinforcement learning approach PeRL tailored for interleaved multimodal tasks, and a multi-stage strategy designed to enhance the exploration-exploitation trade-off, thereby improving learning efficiency and task performance. Specifically, we introduce permutation of image sequences to simulate varied positional relationships to explore more spatial and positional diversity. Furthermore, we design a rollout filtering mechanism for resampling to focus on trajectories that contribute most to learning optimal behaviors to exploit learned policies effectively. We evaluate our model on 5 widely-used multi-image benchmarks and 3 single-image benchmarks. Our experiments confirm that PeRL trained model consistently surpasses R1-related and interleaved VLM baselines by a large margin, achieving state-of-the-art performance on multi-image benchmarks, while preserving comparable performance on single-image tasks.

AIIM Journal 2024 Journal Article

Abnormal recognition-assisted and onset-offset aware network for pathological wearable ECG delineation

  • Yue Zhang
  • Jiewei Lai
  • Chenyu Zhao
  • Jinliang Wang
  • Yong Yan
  • Mingyang Chen
  • Lei Ji
  • Jun Guo

Electrocardiogram (ECG) delineation is essential to the identification of abnormal cardiac status, especially when ECG signals are remotely monitored with wearable devices. The complexity and diversity of cardiac conditions generate numerous pathological ECG patterns, not only requiring the recognition of normal ECG but also addressing an extensive range of abnormal ECG patterns, posing a challenging task. Therefore, we propose an abnormal recognition-assisted network to integrate supplementary information on diverse ECG patterns. Simultaneously, we design an onset-offset aware loss to enhance precise waveform localization. Specifically, we establish a two-branch framework where ECG delineation serves as the target task, producing the final segmentation results. Additionally, the abnormal recognition-assisted network serves as an auxiliary task, extracting multi-label pathological information from ECGs. This joint learning approach establishes crucial correlations between ECG delineation and associated ECG abnormalities. The correlations enable the model to demonstrate sufficient generalization in the presence of diverse abnormal ECG patterns. Besides, onset-offset aware loss focuses intensively on wave onsets and offsets by applying biased weights to various waveform positions. This approach ensures a focus on precise localization, facilitating seamless integration into cross-entropy loss function. A large-scale wearable 12-lead dataset containing 4, 913 signals is collected, offering an extensive range of ECG data for model training. Results demonstrate that our method achieves outstanding performance on two test datasets, attaining sensitivity of 94. 97% and 94. 27% and an error tolerance lower than 20 ms. Furthermore, our method is effective for various aberrant ECG signals, including ST-segment changes, atrial premature beats, and right and left bundle branch blocks.

AAAI Conference 2024 Conference Paper

HORIZON: High-Resolution Semantically Controlled Panorama Synthesis

  • Kun Yan
  • Lei Ji
  • Chenfei Wu
  • Jian Liang
  • Ming Zhou
  • Nan Duan
  • Shuai Ma

Panorama synthesis endeavors to craft captivating 360-degree visual landscapes, immersing users in the heart of virtual worlds. Nevertheless, contemporary panoramic synthesis techniques grapple with the challenge of semantically guiding the content generation process. Although recent breakthroughs in visual synthesis have unlocked the potential for semantic control in 2D flat images, a direct application of these methods to panorama synthesis yields distorted content. In this study, we unveil an innovative framework for generating high-resolution panoramas, adeptly addressing the issues of spherical distortion and edge discontinuity through sophisticated spherical modeling. Our pioneering approach empowers users with semantic control, harnessing both image and text inputs, while concurrently streamlining the generation of high-resolution panoramas using parallel decoding. We rigorously evaluate our methodology on a diverse array of indoor and outdoor datasets, establishing its superiority over recent related work, in terms of both quantitative and qualitative performance metrics. Our research elevates the controllability, efficiency, and fidelity of panorama synthesis to new levels.

NeurIPS Conference 2024 Conference Paper

Voila-A: Aligning Vision-Language Models with User's Gaze Attention

  • Kun Yan
  • Zeyu Wang
  • Lei Ji
  • Yuntao Wang
  • Nan Duan
  • Shuai Ma

In recent years, the integration of vision and language understanding has led to significant advancements in artificial intelligence, particularly through Vision-Language Models (VLMs). However, existing VLMs face challenges in handling real-world applications with complex scenes and multiple objects, as well as aligning their focus with the diverse attention patterns of human users. In this paper, we introduce gaze information, feasibly collected by ubiquitous wearable devices such as MR glasses, as a proxy for human attention to guide VLMs. We propose a novel approach, Voila-A, for gaze alignment to enhance the effectiveness of these models in real-world applications. First, we collect hundreds of minutes of gaze data to demonstrate that we can mimic human gaze modalities using localized narratives. We then design an automatic data annotation pipeline utilizing GPT-4 to generate the VOILA-COCO dataset. Additionally, we introduce a new model VOILA-A that integrate gaze information into VLMs while maintain pretrained knowledge from webscale dataset. We evaluate Voila-A using a hold-out validation set and a newly collected VOILA-GAZE testset, which features real-life scenarios captured with a gaze-tracking device. Our experimental results demonstrate that Voila-A significantly outperforms several baseline models. By aligning model attention with human gaze patterns, Voila-A paves the way for more intuitive, user-centric VLMs and fosters engaging human-AI interaction across a wide range of applications.

NeurIPS Conference 2023 Conference Paper

EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images

  • Seongsu Bae
  • Daeun Kyung
  • Jaehee Ryu
  • Eunbyeol Cho
  • Gyubok Lee
  • Sunjun Kweon
  • Jungwoo Oh
  • Lei Ji

Electronic Health Records (EHRs), which contain patients' medical histories in various multi-modal formats, often overlook the potential for joint reasoning across imaging and table modalities underexplored in current EHR Question Answering (QA) systems. In this paper, we introduce EHRXQA, a novel multi-modal question answering dataset combining structured EHRs and chest X-ray images. To develop our dataset, we first construct two uni-modal resources: 1) The MIMIC- CXR-VQA dataset, our newly created medical visual question answering (VQA) benchmark, specifically designed to augment the imaging modality in EHR QA, and 2) EHRSQL (MIMIC-IV), a refashioned version of a previously established table-based EHR QA dataset. By integrating these two uni-modal resources, we successfully construct a multi-modal EHR QA dataset that necessitates both uni-modal and cross-modal reasoning. To address the unique challenges of multi-modal questions within EHRs, we propose a NeuralSQL-based strategy equipped with an external VQA API. This pioneering endeavor enhances engagement with multi-modal EHR sources and we believe that our dataset can catalyze advances in real-world medical scenarios such as clinical decision-making and research. EHRXQA is available at https: //github. com/baeseongsu/ehrxqa.

NeurIPS Conference 2021 Conference Paper

Learning from Inside: Self-driven Siamese Sampling and Reasoning for Video Question Answering

  • Weijiang Yu
  • Haoteng Zheng
  • Mengfei Li
  • Lei Ji
  • Lijun Wu
  • Nong Xiao
  • Nan Duan

Recent advances in the video question answering (i. e. , VideoQA) task have achieved strong success by following the paradigm of fine-tuning each clip-text pair independently on the pretrained transformer-based model via supervised learning. Intuitively, multiple samples (i. e. , clips) should be interdependent to capture similar visual and key semantic information in the same video. To consider the interdependent knowledge between contextual clips into the network inference, we propose a Siamese Sampling and Reasoning (SiaSamRea) approach, which consists of a siamese sampling mechanism to generate sparse and similar clips (i. e. , siamese clips) from the same video, and a novel reasoning strategy for integrating the interdependent knowledge between contextual clips into the network. The reasoning strategy contains two modules: (1) siamese knowledge generation to learn the inter-relationship among clips; (2) siamese knowledge reasoning to produce the refined soft label by propagating the weights of inter-relationship to the predicted candidates of all clips. Finally, our SiaSamRea can endow the current multimodal reasoning paradigm with the ability of learning from inside via the guidance of soft labels. Extensive experiments demonstrate our SiaSamRea achieves state-of-the-art performance on five VideoQA benchmarks, e. g. , a significant +2. 1% gain on MSRVTT-QA, +2. 9% on MSVD-QA, +1. 0% on ActivityNet-QA, +1. 8% on How2QA and +4. 3% (action) on TGIF-QA.

AAAI Conference 2020 Conference Paper

Functionality Discovery and Prediction of Physical Objects

  • Lei Ji
  • Botian Shi
  • Xianglin Guo
  • Xilin Chen

Functionality is a fundamental attribute of an object which indicates the capability to be used to perform specific actions. It is critical to empower robots the functionality knowledge in discovering appropriate objects for a task e. g. cut cake using knife. Existing research works have focused on understanding object functionality through human-object-interaction from extensively annotated image or video data and are hard to scale up. In this paper, we (1) mine object-functionality knowledge through pattern-based and model-based methods from text, (2) introduce a novel task on physical objectfunctionality prediction, which consumes an image and an action query to predict whether the object in the image can perform the action, and (3) propose a method to leverage the mined functionality knowledge for the new task. Our experimental results show the effectiveness of our methods.

AAAI Conference 2020 Conference Paper

Segment-Then-Rank: Non-Factoid Question Answering on Instructional Videos

  • Kyungjae Lee
  • Nan Duan
  • Lei Ji
  • Jason Li
  • Seung-won Hwang

We study the problem of non-factoid QA on instructional videos. Existing work focuses either on visual or textual modality of video content, to find matching answers to the question. However, neither is flexible enough for our problem setting of non-factoid answers with varying lengths. Motivated by this, we propose a two-stage model: (a) multimodal segmentation of video into span candidates and (b) lengthadaptive ranking of the candidates to the question. First, for segmentation, we propose Segmenter for generating span candidates of diverse length, considering both textual and visual modality. Second, for ranking, we propose Ranker to score the candidates, dynamically combining the two models with complementary strength for both short and long spans respectively. Experimental result demonstrates that our model achieves state-of-the-art performance.

IJCAI Conference 2019 Conference Paper

Knowledge Aware Semantic Concept Expansion for Image-Text Matching

  • Botian Shi
  • Lei Ji
  • Pan Lu
  • Zhendong Niu
  • Nan Duan

Image-text matching is a vital cross-modality task in artificial intelligence and has attracted increasing attention in recent years. Existing works have shown that learning semantic concepts is useful to enhance image representation and can significantly improve the performance of both image-to-text and text-to-image retrieval. However, existing models simply detect semantic concepts from a given image, which are less likely to deal with long-tail and occlusion concepts. Frequently co-occurred concepts in the same scene, e. g. bedroom and bed, can provide common-sense knowledge to discover other semantic-related concepts. In this paper, we develop a Scene Concept Graph (SCG) by aggregating image scene graphs and extracting frequently co-occurred concept pairs as scene common-sense knowledge. Moreover, we propose a novel model to incorporate this knowledge to improve image-text matching. Specifically, semantic concepts are detected from images and then expanded by the SCG. After learning to select relevant contextual concepts, we fuse their representations with the image embedding feature to feed into the matching module. Extensive experiments are conducted on Flickr30K and MSCOCO datasets, and prove that our model achieves state-of-the-art results due to the effectiveness of incorporating the external SCG.

JBHI Journal 2018 Journal Article

Aligning Event Logs to Task-Time Matrix Clinical Pathways in BPMN for Variance Analysis

  • Hui Yan
  • Pieter Van Gorp
  • Uzay Kaymak
  • Xudong Lu
  • Lei Ji
  • Choo Chiap Chiau
  • Hendrikus H. M. Korsten
  • Huilong Duan

Clinical pathways (CPs) are popular healthcare management tools to standardize care and ensure quality. Analyzing CP compliance levels and variances is known to be useful for training and CP redesign purposes. Flexible semantics of the business process model and notation (BPMN) language has been shown to be useful for the modeling and analysis of complex protocols. However, in practical cases one may want to exploit that CPs often have the form of task-time matrices. This paper presents a new method parsing complex BPMN models and aligning traces to the models heuristically. A case study on variance analysis is undertaken, where a CP from the practice and two large sets of patients data from an electronic medical record (EMR) database are used. The results demonstrate that automated variance analysis between BPMN task-time models and real-life EMR data are feasible, whereas that was not the case for the existing analysis techniques. We also provide meaningful insights for further improvement.

TIST Journal 2018 Journal Article

Concept and Attention-Based CNN for Question Retrieval in Multi-View Learning

  • Pengwei Wang
  • Lei Ji
  • Jun Yan
  • Dejing Dou
  • Nisansa De Silva
  • Yong Zhang
  • Lianwen Jin

Question retrieval, which aims to find similar versions of a given question, is playing a pivotal role in various question answering (QA) systems. This task is quite challenging, mainly in regard to five aspects: synonymy, polysemy, word order, question length, and data sparsity. In this article, we propose a unified framework to simultaneously handle these five problems. We use the word combined with corresponding concept information to handle the synonymy problem and the polysemous problem. Concept embedding and word embedding are learned at the same time from both the context-dependent and context-independent views. To handle the word-order problem, we propose a high-level feature-embedded convolutional semantic model to learn question embedding by inputting concept embedding and word embedding. Due to the fact that the lengths of some questions are long, we propose a value-based convolutional attentional method to enhance the proposed high-level feature-embedded convolutional semantic model in learning the key parts of the question and the answer. The proposed high-level feature-embedded convolutional semantic model nicely represents the hierarchical structures of word information and concept information in sentences with their layer-by-layer convolution and pooling. Finally, to resolve data sparsity, we propose using the multi-view learning method to train the attention-based convolutional semantic model on question–answer pairs. To the best of our knowledge, we are the first to propose simultaneously handling the above five problems in question retrieval using one framework. Experiments on three real question-answering datasets show that the proposed framework significantly outperforms the state-of-the-art solutions.

AAAI Conference 2015 Conference Paper

Acronym Disambiguation Using Word Embedding

  • Chao Li
  • Lei Ji
  • Jun Yan

According to the website AcronymFinder. com which is one of the world's largest and most comprehensive dictionaries of acronyms, an average of 37 new human-edited acronym definitions are added every day. There are 379, 918 acronyms with 4, 766, 899 definitions on that site up to now, and each acronym has 12. 5 definitions on average. It is a very important research topic to identify what exactly an acronym means in a given context for document comprehension as well as for document retrieval. In this paper, we propose two word embedding based models for acronym disambiguation. Word embedding is to represent words in a continuous and multidimensional vector space, so that it is easy to calculate the semantic similarity between words by calculating the vector distance. We evaluate the models on MSH Dataset and ScienceWISE Dataset, and both models outperform the state-of-art methods on accuracy. The experimental results show that word embedding helps to improve acronym disambiguation.

AIIM Journal 2015 Journal Article

On local anomaly detection and analysis for clinical pathways

  • Zhengxing Huang
  • Wei Dong
  • Lei Ji
  • Liangying Yin
  • Huilong Duan

Objective Anomaly detection, as an imperative task for clinical pathway (CP) analysis and improvement, can provide useful and actionable knowledge of interest to clinical experts to be potentially exploited. Existing studies mainly focused on the detection of global anomalous inpatient traces of CPs using the similarity measures in a structured manner, which brings order in the chaos of CPs, may decline the accuracy of similarity measure between inpatient traces, and may distort the efficiency of anomaly detection. In addition, local anomalies that exist in some subsegments of events or behaviors in inpatient traces are easily overlooked by existing approaches since they are designed for detecting global or large anomalies. Method In this study, we employ a probabilistic topic model to discover underlying treatment patterns, and assume any significant unexplainable deviations from the normal behaviors surmised by the derived patterns are strongly correlated with anomalous behaviours. In this way, we can figure out the detailed local abnormal behaviors and the associations between these anomalies such that diagnostic information on local anomalies can be provided. Results The proposed approach is evaluated via a clinical data-set, including 2954 unstable angina patient traces and 483, 349 clinical events, extracted from a Chinese hospital. Using the proposed method, local anomalies are detected from the log. In addition, the identified associations between the detected local anomalies are derived from the log, which lead to clinical concern on the reason resulting in these anomalies in CPs. The correctness of the proposed approach has been evaluated by three experience cardiologists of the hospital. For four types of local anomalies (i. e. , unexpected events, early events, delay events, and absent events), the proposed approach achieves 94%, 71% 77%, and 93. 2% in terms of recall. This is quite remarkable as we do not use a prior knowledge. Conclusion Substantial experimental results show that the proposed approach can effectively detect local anomalies in CPs, and also provide diagnostic information on the detected anomalies in an informative manner.

AAAI Conference 2011 Conference Paper

Collaborative Users’ Brand Preference Mining across Multiple Domains from Implicit Feedbacks

  • Jian Tang
  • Jun Yan
  • Lei Ji
  • Ming Zhang
  • Shaodan Guo
  • Ning Liu
  • Xianfang Wang
  • Zheng Chen

Advanced e-applications require comprehensive knowledge about their users’ preferences in order to provide accurate personalized services. In this paper, we propose to learn users’ preferences to product brands from their implicit feedbacks such as their searching and browsing behaviors in user Web browsing log data. The user brand preference learning problem is challenge since (1) the users’ implicit feedbacks are extremely sparse in various product domains; and (2) we can only observe positive feedbacks from users’ behaviors. In this paper, we propose a latent factor model to collaboratively mine users’ brand preferences across multiple domains simultaneously. By collective learning, the learning processes in all the domains are mutually enhanced and hence the problem of data scarcity in each single domain can be effectively addressed. On the other hand, we learn our model with an adaption of the Bayesian personalized ranking (BPR) optimization criterion which is a general learning framework for collaborative filtering from implicit feedbacks. Experiments with both synthetic and real world datasets show that our proposed model significantly outperforms the baselines.

v2026.09.13