Arrow Research search

Author name cluster

Jiahui Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

10 papers
2 author rows

Possible papers

10

AAAI Conference 2026 Conference Paper

Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation Comprehension

  • Lin Li
  • Wei Chen
  • Jiahui Li
  • Kwang-Ting Cheng
  • Long Chen

Recent advances in multi-modal large language models (MLLMs) have significantly improved object-level grounding and region captioning. However, they remain limited in visual relation understanding, struggling even with binary relation detection, let alone N-ary relations involving multiple semantic roles. The core reason is the lack of modeling for structural semantic dependencies among multi-entities, leading to over-reliance on language priors (e.g., defaulting to "person drinks a milk" if a person is merely holding it). To this end, we propose Relation-R1, the first unified relation comprehension framework that explicitly integrates cognitive chain-of-thought (CoT)-guided supervised fine-tuning (SFT) and group relative policy optimization (GRPO) within a reinforcement learning (RL) paradigm. Specifically, we first establish foundational reasoning capabilities via SFT, enforcing structured outputs with thinking processes. Then, GRPO is utilized to refine these outputs via multi-rewards optimization, prioritizing visual-semantic grounding over language-induced biases, thereby improving generalization capability. Furthermore, we investigate the impact of various CoT strategies within this framework, demonstrating that a specific-to-general progressive approach in CoT guidance further improves generalization, especially in capturing synonymous N-ary relations. Extensive experiments on widely-used PSG and SWiG datasets demonstrate that Relation-R1 achieves state-of-the-art performance in both binary and N-ary relation understanding.

EAAI Journal 2026 Journal Article

Secure and energy-efficient unmanned aerial vehicle-enabled visible light communication via a multi-objective optimization approach

  • Lingling Liu
  • Aimin Wang
  • Jing Wu
  • Jiao Lu
  • Jiahui Li
  • Geng Sun

This research investigates a unique approach for providing communication service to terrestrial receivers via using unmanned aerial vehicle-enabled visible light communication. Specifically, we consider a scenario involving multiplex transmitters, multiplex receivers, and a single eavesdropper, each of which is equipped with a single photodetector. Then, a unmanned aerial vehicle deployment multi-objective optimization problem is formulated to simultaneously enhance the uniformity of the optical power received by the receivers, minimize the amount of information collected by the eavesdropper, and minimize the energy consumption of unmanned aerial vehicles, by jointly optimizing the locations and transmission power of unmanned aerial vehicles under specific constraints. Due to the complexity and nonlinearity of the formulated unmanned aerial vehicle deployment multi-objective optimization problem, conventional methods are inadequate for solving it efficiently. Therefore, a multi-objective evolutionary algorithm based on decomposition with chaos initiation and crossover mutation is proposed. Simulation outcomes demonstrate that the proposed approach outperforms other approaches, providing significant improvements in both security and energy efficiency for visible light communication systems. In addition, extensive simulations are conducted to verify the convergence, robustness, adaptability, and scalability of the proposed method under various influencing factors, including terrestrial receiver distributions, unmanned aerial vehicle height variations, visible light communication channel uncertainty, unmanned aerial vehicle position dynamics, as well as scenario scale changes. The results show that the proposed method exhibits superior convergence, robustness, adaptability and scalability, and can flexibly adapt to the dynamic changes of channels and unmanned aerial vehicle positions. Moreover, it ensures coverage fairness, system security and energy efficiency of unmanned aerial vehicles, which further makes it well-suited for practical unmanned aerial vehicle-enabled visible light communication systems operating in dynamic and uncertain environments.

AILAW Journal 2025 Journal Article

A method of legal judgment prediction via prompt learning and charge keywords fusion

  • Jiahui Li
  • Jianquan Ouyang

Abstract Legal Judgment Prediction (LJP) is a research hotspot in legal intelligence, which aim to predict the judgment result based on the fact description. The existing research mainly focuses on multiclass classification methods and single-label learning by analyzing case facts, but neglect the semantic correlation between the fact descriptions and legal keyword labels, and the significance of legal keywords, which leads to unsatisfactory judicial prediction results. To address this limitation, we introduce a novel framework for LJP that enhances the utilization of legal concept keywords through prompt learning. We propose a method based on legal charge keywords and prompt engineering to enhance the performance for LJP. Our approach first integrates legal keywords with fact descriptions to improve the representation capacity of case fact vectors. We have incorporated the legal keywords into the language model to enhance the model’s ability to understand and process legal texts. And we design a prompt template to guide the reasoning process of the pre-trained language model through structured instructions. This strengthens the semantical relevance between the fact description and legal labels, so that the model can more accurately capture the logical connection between fact descriptions and charges. Experimental results on CAIL2018 datasets across different tasks show an improvement in F1 scores ranging from at least 1. 59% to a maximum of 9. 28%, illustrating the effectiveness of our method compared with the state of the art models such as LADAN, NeurJudge and CL4LJP.

AAAI Conference 2025 Conference Paper

Learning Causal Transition Matrix for Instance-dependent Label Noise

  • Jiahui Li
  • Tai-Wei Chang
  • Kun Kuang
  • Ximing Li
  • Long Chen
  • Jun Zhou

Noisy labels are both inevitable and problematic in machine learning methods, as they negatively impact models' generalization ability by causing overfitting. In the context of learning with noise, the transition matrix plays a crucial role in the design of statistically consistent algorithms. However, the transition matrix is often considered unidentifiable. One strand of methods typically addresses this problem by assuming that the transition matrix is instance-independent; that is, the probability of mislabeling a particular instance is not influenced by its characteristics or attributes. This assumption is clearly invalid in complex real-world scenarios. To better understand the transition relationship and relax this assumption, we propose to study the data generation process of noisy labels from a causal perspective. We discover that an unobservable latent variable can affect either the instance itself, the label annotation procedure, or both, which complicates the identification of the transition matrix. To address various scenarios, we have unified these observations within a new causal graph. In this graph, the input instance is divided into a noise-resistant component and a noise-sensitive component based on whether they are affected by the latent variable. These two components contribute to identifying the “causal transition matrix”, which approximates the true transition matrix with theoretical guarantee. In line with this, we have designed a novel training framework that explicitly models this causal relationship and, as a result, achieves a more accurate model for inferring the clean label.

EAAI Journal 2025 Journal Article

Multi-objective deployment optimization for integrated sensing and communication-enabled unmanned aerial vehicle swarm

  • Hongjuan Li
  • Haiyuan Chen
  • Miao Wang
  • Jiahui Li
  • Hui Kang
  • Yuzhuo Guan
  • Xu Lin

With the convergence of mobile communication, sensing, and computational networks in sixth-generation technology, the integration of sensing and communication with unmanned aerial vehicles (UAVs) is promising. This paper focuses on the contribution of artificial intelligence in optimizing the deployment of UAV swarms for multi-objective target detection applications in sixth-generation networks. Specifically, the artificial intelligence contribution lies in the development of an improved multi-objective particle swarm optimization (IMOPSO) algorithm for solving a complex multi-objective deployment problem. The problem aims to simultaneously optimize communication rate, sensing quality, and energy consumption in the deployment of UAV swarms. To address this, the proposed IMOPSO incorporates chaotic initialization, Lévy flight mutation, dynamic mutation rate, and an elimination mechanism based on opposition-based learning. These innovations are designed to enhance the algorithm’s ability to explore the solution space effectively, overcome premature convergence to local solutions, and improve solution quality. In terms of engineering applications, the IMOPSO is applied to the deployment of UAV swarms for target detection, demonstrating its ability to enhance communication and sensing performance while reducing energy consumption in practical scenarios. Through extensive simulations, we show that the IMOPSO outperforms traditional optimization methods and other baseline algorithms, achieving superior results across all optimization objectives. Specifically, the IMOPSO achieves approximately 5% higher transmission data rate, 9% better sensing quality, and 19% lower energy consumption compared to baseline algorithms across multiple test scenarios. Furthermore, the solutions obtained are not only closer to the optimal front but also more concentrated, indicating higher-quality results.

ICRA Conference 2025 Conference Paper

RMP-YOLO: A Robust Motion Predictor for Partially Observable Scenarios Even if You Only Look Once

  • Jiawei Sun 0006
  • Jiahui Li
  • Tingchen Liu
  • Chengran Yuan
  • Shuo Sun
  • Zefan Huang
  • Anthony Wong
  • Keng Peng Tee

We introduce RMP-YOLO, a unified framework designed to provide robust motion predictions even with incomplete input data. Our key insight stems from the observation that complete and reliable historical trajectory data plays a pivotal role in ensuring accurate motion prediction. Therefore, we propose a new paradigm that prioritizes the reconstruction of intact historical trajectories before feeding them into the prediction modules. Our approach introduces a novel scene tokenization module to enhance the extraction and fusion of spatial and temporal features. Following this, our proposed recovery module reconstructs agents' incomplete historical trajectories by leveraging local map topology and interactions with nearby agents. The reconstructed, clean historical data is then integrated into the downstream prediction modules. Our framework is able to effectively handle missing data of varying lengths and remains robust against observation noise while maintaining high prediction accuracy. Furthermore, our recovery module is compatible with existing prediction models, ensuring seamless integration. Extensive experiments validate the effectiveness of our approach, and deployment in real-world autonomous vehicles confirms its practical utility. In the 2024 Waymo Motion Prediction Competition, our method, RMP-YOLO, achieves state-of-the-art performance, securing third place. Our code is open-source at https://github.com/ggosjw/RMP-YOLO.

ECAI Conference 2025 Conference Paper

Temporal Knowledge Graph Question Answering via Sub-Question Decomposition and Cross-Checking

  • Jiahui Li
  • Xiangzhen Meng
  • Peilin Jiang
  • Xilang Tang

Temporal Knowledge Graph Question Answering (TKGQA) requires precise alignment between natural language questions and temporally structured knowledge. While Large Language Models (LLMs) have shown promise in complex reasoning tasks, they often struggle with ambiguous temporal cues, resulting in hallucinations and limited robustness. To address these issues, we propose DeCompKGQA, a two-stage framework that leverages sub-question decomposition and a dedicated cross-checking module for iterative temporal reasoning. In the first stage, the input question is decomposed into a temporally constrained sub-question, enabling finer control over temporal conditions and logical dependencies while reducing ambiguity. In the second stage, the sub-question guides the selection of executable query templates for candidate retrieval, after which a cross-checking module compares the retrieved candidates against the outputs of an embedding-based TKGQA model on the same sub-question to determine the intermediate answer. This hybrid approach aligns the implicit reasoning capabilities of LLMs with explicit temporal facts, improving robustness and mitigating hallucinations. Experimental results on multiple benchmark datasets demonstrate that DeCompKGQA consistently outperforms strong baselines, achieving state-of-the-art performance on complex multi-hop temporal queries. Code is available at: https: //github. com/lijh01/DeCompKGQA.

NeurIPS Conference 2023 Conference Paper

Two Heads are Better Than One: A Simple Exploration Framework for Efficient Multi-Agent Reinforcement Learning

  • Jiahui Li
  • Kun Kuang
  • Baoxiang Wang
  • Xingchen Li
  • Fei Wu
  • Jun Xiao
  • Long Chen

Exploration strategy plays an important role in reinforcement learning, especially in sparse-reward tasks. In cooperative multi-agent reinforcement learning~(MARL), designing a suitable exploration strategy is much more challenging due to the large state space and the complex interaction among agents. Currently, mainstream exploration methods in MARL either contribute to exploring the unfamiliar states which are large and sparse, or measuring the interaction among agents with high computational costs. We found an interesting phenomenon that different kinds of exploration plays a different role in different MARL scenarios, and choosing a suitable one is often more effective than designing an exquisite algorithm. In this paper, we propose a exploration method that incorporate the \underline{C}uri\underline{O}sity-based and \underline{IN}fluence-based exploration~(COIN) which is simple but effective in various situations. First, COIN measures the influence of each agent on the other agents based on mutual information theory and designs it as intrinsic rewards which are applied to each individual value function. Moreover, COIN computes the curiosity-based intrinsic rewards via prediction errors which are added to the extrinsic reward. For integrating the two kinds of intrinsic rewards, COIN utilizes a novel framework in which they complement each other and lead to a sufficient and effective exploration on cooperative MARL tasks. We perform extensive experiments on different challenging benchmarks, and results across different scenarios show the superiority of our method.

IJCAI Conference 2022 Conference Paper

A Speech-driven Sign Language Avatar Animation System for Hearing Impaired Applications

  • Li Hu
  • Jiahui Li
  • Jiashuo Zhang
  • Qi Wang
  • Bang Zhang
  • Ping Tan

Sign language is the communication language used in hearing impaired community. Recently, the research of sign language production has made great progress but still need to cope with some critical challenges. In this paper, we propose a system-level scheme and push forward the implementation of sign language production for practical usage. We build a system capable of translating speech into sign language avatar. Different from previous approach only focusing on single technology, we systematically combine algorithms of language translation, body gesture animation and facial avatar generation. We also develop two applications: Sign Language Interpretation APP and Virtual Sign Language Anchor, to facilitate easy and clear communication for hearing impaired people.

JBHI Journal 2021 Journal Article

Left Ventricle Quantification Challenge: A Comprehensive Comparison and Evaluation of Segmentation and Regression for Mid-Ventricular Short-Axis Cardiac MR Data

  • Wufeng Xue
  • Jiahui Li
  • Zhiqiang Hu
  • Eric Kerfoot
  • James Clough
  • Ilkay Oksuz
  • Hao Xu
  • Vicente Grau

Automatic quantification of the left ventricle (LV) from cardiac magnetic resonance (CMR) images plays an important role in making the diagnosis procedure efficient, reliable, and alleviating the laborious reading work for physicians. Considerable efforts have been devoted to LV quantification using different strategies that include segmentation-based (SG) methods and the recent direct regression (DR) methods. Although both SG and DR methods have obtained great success for the task, a systematic platform to benchmark them remains absent because of differences in label information during model learning. In this paper, we conducted an unbiased evaluation and comparison of cardiac LV quantification methods that were submitted to the Left Ventricle Quantification (LVQuan) challenge, which was held in conjunction with the Statistical Atlases and Computational Modeling of the Heart (STACOM) workshop at the MICCAI 2018. The challenge was targeted at the quantification of 1) areas of LV cavity and myocardium, 2) dimensions of the LV cavity, 3) regional wall thicknesses (RWT), and 4) the cardiac phase, from mid-ventricle short-axis CMR images. First, we constructed a public quantification dataset Cardiac-DIG with ground truth labels for both the myocardium mask and these quantification targets across the entire cardiac cycle. Then, the key techniques employed by each submission were described. Next, quantitative validation of these submissions were conducted with the constructed dataset. The evaluation results revealed that both SG and DR methods can offer good LV quantification performance, even though DR methods do not require densely labeled masks for supervision. Among the 12 submissions, the DR method LDAMT offered the best performance, with a mean estimation error of 301 mm $^2$ for the two areas, 2. 15 mm for the cavity dimensions, 2. 03 mm for RWTs, and a 9. 5% error rate for the cardiac phase classification. Three of the SG methods also delivered comparable performances. Finally, we discussed the advantages and disadvantages of SG and DR methods, as well as the unsolved problems in automatic cardiac quantification for clinical practice applications.

v2026.09.13