Arrow Research search

Author name cluster

Li Jin

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
2 author rows

Possible papers

12

AAAI Conference 2026 Conference Paper

Rectify Evaluation Preference: Improving LLMs’ Critique on Math Reasoning via Perplexity-aware Reinforcement Learning

  • Changyuan Tian
  • Zhicong Lu
  • Shuang Qian
  • Nayu Liu
  • Peiguang Li
  • Li Jin
  • Leiyi Hu
  • Zhizhao Zeng

To improve Multi-step Mathematical Reasoning (MsMR) of Large Language Models (LLMs), it is crucial to obtain scalable supervision from the corpus by automatically critiquing mistakes in the reasoning process of MsMR and rendering a final verdict of the problem-solution. Most existing methods rely on crafting high-quality supervised fine-tuning demonstrations for critiquing capability enhancement and pay little attention to delving into the underlying reason for the poor critiquing performance of LLMs. In this paper, we orthogonally quantify and investigate the potential reason — imbalanced evaluation preference, and conduct a statistical preference analysis. Motivated by the analysis of the reason, a novel perplexity-aware reinforcement learning algorithm is proposed to rectify the evaluation preference, elevating the critiquing capability. Specifically, to probe into LLMs' critiquing characteristics, a One-to-many Problem-Solution (OPS) benchmark is meticulously constructed to quantify the behavior difference of LLMs when evaluating the problem solutions generated by itself and others. Then, to investigate the behavior difference in depth, we conduct a statistical preference analysis oriented on perplexity and find an intriguing phenomenon — "LLMs incline to judge solutions with lower perplexity as correct", which is dubbed as imbalanced evaluation preference. To rectify this preference, we regard perplexity as the baton in the algorithm of Group Relative Policy Optimization, supporting the LLMs to explore trajectories that judge lower perplexity as wrong and higher perplexity as correct. Extensive experimental results on our built OPS and existing available critic benchmarks demonstrate the validity of our method.

ICRA Conference 2025 Conference Paper

A Visual Servo System for Robotic on-Orbit Servicing Based on 3D Perception of Non-Cooperative Satellite

  • Panpan Zhao
  • Li Jin
  • Yeheng Chen
  • Jiachen Li
  • Xiuqiang Song
  • Wenxuan Chen
  • Nan Li
  • Wenjuan Du

The 3D perception of satellites, including both their shape and pose, is a key foundation for robotic on-orbit servicing. However, the demanding space environment-such as intense and dim illumination-presents significant challenges. Previous non-cooperative methods focus on specific geometric features like solar panel brackets or docking rings, overlooking the satellite's overall shape and increasing the risk of collisions during grasping. Additionally, satellites are often weakly textured, limiting the accuracy of 3D perception. To address these issues, we propose, for the first time, a 3D perceptionbased visual servo system of non-cooperative satellites. This system combines reconstruction and tracking to enhance shape perception and pose estimation accuracy in orbital conditions. Specifically, we employ an alternating iterative strategy to simultaneously reconstruct and track the satellite and introduce a novel constraint to fuse different cues under extreme conditions. Further, we develop a simulation environment platform, a dualarm microgravity grasping system, and an online monitoring module to enhance system capabilities for on-orbit servicing. Synthetic and real-world datasets from the simulation environment are also created for experimental validation. Results show that each module of our system achieves state-of-the-art performance.

AAAI Conference 2025 Conference Paper

HyperMixer: Specializable Hypergraph Channel Mixing for Long-term Multivariate Time Series Forecasting

  • Changyuan Tian
  • Zhicong Lu
  • Zequn Zhang
  • Heming Yang
  • Wei Cao
  • Zhi Guo
  • Xian Sun
  • Li Jin

Long-term Multivariate Time Series (LMTS) forecasting aims to predict extended future trends based on channel-interrelated historical data. Considering the elusive channel correlations, most existing methods compromise by treating channels as independent or tentatively modeling pairwise channel interactions, making it challenging to handle the characteristics of both higher-order interactions and time variation in channel correlations. In this paper, we propose HyperMixer, a novel specializable hypergraph channel mixing plugin which introduces versatile hypergraph structures to capture group channel interactions and time-varying patterns for long-term multivariate time series forecasting. Specifically, to encode the higher-order channel interactions, we structure multiple channels into a hypergraph, achieving a two-phase message-passing mechanism: channel-to-group and group-to-channel. Moreover, the functionally specializable hypergraph structures are presented to boost the capability of hypergraph to capture the time-varying patterns across periods, further refining modeling of channel correlations. Extensive experimental results on seven available benchmark datasets demonstrate the effectiveness and generalization of our plugin in LMTS forecasting. The visual analysis further illustrates that HyperMixer with specializable hypergraphs tailors channel interactions specific to certain periods.

AAAI Conference 2024 Conference Paper

CAMEL: Capturing Metaphorical Alignment with Context Disentangling for Multimodal Emotion Recognition

  • Linhao Zhang
  • Li Jin
  • Guangluan Xu
  • Xiaoyu Li
  • Cai Xu
  • Kaiwen Wei
  • Nayu Liu
  • Haonan Liu

Understanding the emotional polarity of multimodal content with metaphorical characteristics, such as memes, poses a significant challenge in Multimodal Emotion Recognition (MER). Previous MER researches have overlooked the phenomenon of metaphorical alignment in multimedia content, which involves non-literal associations between concepts to convey implicit emotional tones. Metaphor-agnostic MER methods may be misinformed by the isolated unimodal emotions, which are distinct from the real emotions blended in multimodal metaphors. Moreover, contextual semantics can further affect the emotions associated with similar metaphors, leading to the challenge of maintaining contextual compatibility. To address the issue of metaphorical alignment in MER, we propose to leverage a conditional generative approach for capturing metaphorical analogies. Our approach formulates schematic prompts and corresponding references based on theoretical foundations, which allows the model to better grasp metaphorical nuances. In order to maintain contextual sensitivity, we incorporate a disentangled contrastive matching mechanism, which undergoes curricular adjustment to regulate its intensity during the learning process. The automatic and human evaluation experiments on two benchmarks prove that, our model provides considerable and stable improvements in recognizing multimodal emotion with metaphor attributes.

ICRA Conference 2024 Conference Paper

Implicit Coarse-to-Fine 3D Perception for Category-level Object Pose Estimation from Monocular RGB Image

  • Jia Li
  • Li Jin
  • Xibin Song
  • Yeheng Chen
  • Nan Li
  • Xueying Qin

Category-level object pose estimation demonstrates robust generalization capabilities that benefit robotics applications. However, exclusive reliance on RGB images without leveraging any 3D information introduces ambiguity in the translation and size of objects, leading to suboptimal performance. In this paper, we propose a framework for category-level pose estimation from a single RGB image in an end-to-end manner, i. e. , Feature Auxiliary Perception Network (FAP-Net). To address inaccurate pose estimation caused by the inherent ambiguity of RGB images, we design a coarse-to-fine approach that first harnesses geometry supervision to facilitate coarse 3D feature perception and subsequently refines the features based on pose and size constraints. Experimental results on REAL275 and CAMERA25 demonstrate that FAP-Net achieves significant improvements (14. 7% on 10°10cm and 11. 4% on IoU50 on the real-scene REAL275 dataset) over the state-of-the-art and real-time inference (42 FPS).

AAAI Conference 2024 Conference Paper

Video Event Extraction with Multi-View Interaction Knowledge Distillation

  • Kaiwen Wei
  • Runyan Du
  • Li Jin
  • Jian Liu
  • Jianhua Yin
  • Linhao Zhang
  • Jintao Liu
  • Nayu Liu

Video event extraction (VEE) aims to extract key events and generate the event arguments for their semantic roles from the video. Despite promising results have been achieved by existing methods, they still lack an elaborate learning strategy to adequately consider: (1) inter-object interaction, which reflects the relation between objects; (2) inter-modality interaction, which aligns the features from text and video modality. In this paper, we propose a Multi-view Interaction with knowledge Distillation (MID) framework to solve the above problems with the Knowledge Distillation (KD) mechanism. Specifically, we propose the self-Relational KD (self-RKD) to enhance the inter-object interaction, where the relation between objects is measured by distance metric, and the high-level relational knowledge from the deeper layer is taken as the guidance for boosting the shallow layer in the video encoder. Meanwhile, to improve the inter-modality interaction, the Layer-to-layer KD (LKD) is proposed, which integrates additional cross-modal supervisions (i.e., the results of cross-attention) with the textual supervising signal for training each transformer decoder layer. Extensive experiments show that without any additional parameters, MID achieves the state-of-the-art performance compared to other strong methods in VEE.

ICRA Conference 2023 Conference Paper

Online Hand-Eye Calibration with Decoupling by 3D Textureless Object Tracking

  • Li Jin
  • Kang Xie
  • Wenxuan Chen
  • Xin Cao
  • Yuehua Li
  • Jiachen Li
  • Jiankai Qian
  • Xueying Qin

Hand-eye calibration estimates the pose of a camera relative to a robot, which is a fundamental problem for visually guided robots, especially for dynamic object grasping. Most methods use 2D fiducial markers with distinctive visual features and require pre-calibration for accurate calibration, which can not work online. In this paper, we propose a novel hand-eye calibration method based on the natural 3D object, which can work online and automatically even if the object is textureless or weakly textured. We first propose a Pose Refinement Network (PR-Net) to improve the accuracy of 3D object tracking. Then we build a 3D convergence point constraint based on the multi-view information with the accurate object pose to adjust the object position. Finally, we optimize the hand-eye pose by the closed-loop constraint with the optimized object position, solving the problem that is easy to fall into a local minimum. The experiments show that the average error of our hand-eye calibration method is 1. 20 degrees and 23. 18 mm. The results achieve state-of-the-art by using the working object to realize the online hand-eye calibration.

AAAI Conference 2023 Conference Paper

TOT:Topology-Aware Optimal Transport for Multimodal Hate Detection

  • Linhao Zhang
  • Li Jin
  • Xian Sun
  • Guangluan Xu
  • Zequn Zhang
  • Xiaoyu Li
  • Nayu Liu
  • Qing Liu

Multimodal hate detection, which aims to identify the harmful content online such as memes, is crucial for building a wholesome internet environment. Previous work has made enlightening exploration in detecting explicit hate remarks. However, most of their approaches neglect the analysis of implicit harm, which is particularly challenging as explicit text markers and demographic visual cues are often twisted or missing. The leveraged cross-modal attention mechanisms also suffer from the distributional modality gap and lack logical interpretability. To address these semantic gap issues, we propose TOT: a topology-aware optimal transport framework to decipher the implicit harm in memes scenario, which formulates the cross-modal aligning problem as solutions for optimal transportation plans. Specifically, we leverage an optimal transport kernel method to capture complementary information from multiple modalities. The kernel embedding provides a non-linear transformation ability to reproduce a kernel Hilbert space (RKHS), which reflects significance for eliminating the distributional modality gap. Moreover, we perceive the topology information based on aligned representations to conduct bipartite graph path reasoning. The newly achieved state-of-the-art performance on two publicly available benchmark datasets, together with further visual analysis, demonstrate the superiority of TOT in capturing implicit cross-modal alignment.

JBHI Journal 2021 Journal Article

User Independent Estimations of Gait Events With Minimal Sensor Data

  • Seth R. Donahue
  • Li Jin
  • Michael E. Hahn

Goal: The purpose of this study was to provide an initial examination of the utility of the Beta Process - Auto Regressive - Hidden Markov Model (BP-AR-HMM) for the prior identification of gait events. A secondary objective was to determine whether the output of the model could be used for classification and prediction of locomotion states. Methods: In this study we utilized the output of the BP-AR-HMM to develop user-independent identification of gait events and gait classification from an idealized three-dimensional acceleration signal. The input acceleration data were collected from two walking (1. 4 and 1. 6 ms -1 ) and two running (2. 6 and 3. 0 ms -1 ) steady state speeds, and during two dynamic walk to run and run to walk transitions (1. 8-2. 4 and 2. 4-1. 8 ms -1 ) on an instrumented force treadmill. Results: The BP-AR-HMM identified 9 unique states. Of these, two states, 4 and 1, were utilized to estimate initial contact and toe off, respectively. The lead time from the first instance of state 4 to initial contact was 0. 13 ± 0. 02 s. Similarly, the first instance of state 1 occurred 0. 14 ± 0. 03 s before toe off. Two other states (3 and 7) were examined for possible utilization in a probabilistic model for the prediction of pending locomotion state transitions. Conclusion: The identification of gait events prior to their occurrence by the BP-AR-HMM appears to be an approach that can minimize the quantity of sensor data in an offline approach. Furthermore, there is evidence it could also be used as a basis to build a probabilistic model to estimate locomotion transitions.

YNICL Journal 2019 Journal Article

Deep/mixed cerebral microbleeds are associated with cognitive dysfunction through thalamocortical connectivity disruption: The Taizhou Imaging Study

  • Yingzhe Wang
  • Yanfeng Jiang
  • Chen Suo
  • Ziyu Yuan
  • Kelin Xu
  • Qi Yang
  • Weijun Tang
  • Kexun Zhang

BACKGROUND: Cerebral microbleeds (CMBs) are considered to be risk factors for cognitive dysfunction. The specific pathology and clinical manifestations of CMBs are different based on their locations. We investigated the association between CMBs at different locations and cognitive dysfunction and explored the potential underlying pathways in a rural Han Chinese population. METHODS: We used baseline data from 562 community-dwelling adults (55-65 years old) in the Taizhou Imaging Study between 2013 and 2015. All individuals underwent multimodal brain magnetic resonance imaging (MRI) and 444 subjects completed neuropsychological tests: the Mini-Mental Status Examination and the Montreal Cognitive Assessment. Multinomial logistic regression was used to estimate the association between CMBs and cognitive dysfunction. The volume of brain regions and white matter microstructure were analyzed using Freesurfer and tract-based spatial statistics, respectively. RESULTS: CMBs were detected in 104 individuals (18.5%) in our study. Multinomial logistic regression found deep/mixed CMBs were associated with global cognitive dysfunction (OR 3.52; 95% CI 1.21 to 10.26), whereas lobar CMBs (OR 1.76; 95% CI 0.56 to 5.53) were not. Quantification of multimodal brain MRI showed that deep/mixed CMBs were accompanied by decreased thalamic volume and loss of fractional anisotropy of bilateral anterior thalamic radiations. CONCLUSION: Deep/mixed CMBs were associated with cognitive dysfunction in this Chinese cross-sectional study. Disruption of thalamocortical connectivity might be a potential pathway underlying this relationship.

AIJ Journal 2011 Journal Article

Dynamics of argumentation systems: A division-based method

  • Beishui Liao
  • Li Jin
  • Robert C. Koons

The changing of arguments and their attack relation is an intrinsic property of a variety of argumentation systems. So, it is very important to efficiently figure out how the status of arguments in a system evolves when the system is updated. However, unlike other areas of argumentation that have been deeply explored, such as argumentation semantics, proof theories, and algorithms, etc. , dynamics of argumentation systems has been comparatively neglected. In this paper, we formulate a general theory (called a division-based method) to cope with this problem based on a new concept: the division of an argumentation framework. When an argumentation framework is updated, it is divided into three parts: an unaffected, an affected, and a conditioning part. The status of arguments in the unaffected sub-framework remains unchanged, while the status of the affected arguments is computed in a special argumentation framework (called a conditioned argumentation framework, or briefly CAF) that is composed of an affected part and a conditioning part. We have proved that under a certain semantics that satisfies the directionality criterion (complete, preferred, ideal, or grounded semantics), the extensions of the updated framework are equal to the result of a combination of the extensions of an unaffected sub-framework and sets of the extensions of a set of assigned CAFs. Due to the efficiency of the division-based method, it is expected to be very useful in various kinds of argumentation systems where arguments and attacks are dynamics.

AAAI Conference 2006 Short Paper

KDMAS: A Multi-Agent System for Knowledge Discovery via Planning

  • Li Jin

In the real world, there are some domain knowledge discovery problems that can be formulated into knowledge-based planning problems, such as chemical reaction process and biological pathway discovery problems. A view of these domain problems can be re-cast as a planning problem, such that initial and final states are known and processes can be captured as abstract operators that modify the environment. We believe that AI planning technology can provide a modeling formalism for this task such that hypotheses can be generated, tested, queried and qualitatively simulated to improve the domain knowledge and rules. Our approach is to build a general multi-agent system for knowledge discovery (KDMAS) via planning for any domain whose problems can be modeled as AI planning problems. The plans produced are hypotheses capturing relevant qualitative information regarding domain knowledge. We will use the biological pathway domain as a model to present our approach.

v2026.09.13