Arrow Research search

Author name cluster

Yaqian Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
2 author rows

Possible papers

12

EAAI Journal 2026 Journal Article

A semantic and geometric perception framework for safety evaluation in bulk cargo grab operations

  • Yikang Shi
  • Weipeng Rong
  • Yaqian Li
  • Haibin Li
  • Wenming Zhang
  • Zhongqiang Wu

Bulk cargo grab operations in ports are highly safety-critical, requiring reliable perception and interpretable monitoring under adverse environmental conditions. This study presents a Light Detection and Ranging (LiDAR) perception framework that constructs a consistent evidence chain from semantic segmentation to auditable safety indicators. A deep learning semantic segmentation module with a PointNet++ backbone is adopted to provide class-specific cues from point clouds. The model is further integrated with hatch-specific geometric priors and temporal state-space filters to obtain stabilized corner trajectories with explicit variance estimation. Building on these stabilized results, interpretable risk metrics — including minimum grab-to-hatch distance, Time-to-Collision (TTC), and swing angle — are derived, while grid-based material height mapping supports operational planning and throughput management. Uncertainty propagation is modeled throughout the perception–geometry–risk chain, enabling rational trade-offs between false alarms and missed detections. Extensive experiments on a multi-condition dataset of 78, 000 frames collected over 43 h demonstrate hatch geometry estimation within 20–30 centimeters (cm), TTC prediction with a mean absolute error of 0. 65 seconds (s), and balanced false positive and negative rates near five percent, alongside an online processing speed of approximately 1. 3 frames per second (FPS), with a peak graphics processing unit (GPU) memory usage of about 1. 2 gigabytes (GB). Overall, the proposed framework advances safety-oriented LiDAR perception in safety-critical port environments by integrating semantic segmentation cues with geometric reasoning and interpretable risk metrics, offering both methodological novelty and engineering feasibility for intelligent monitoring and collision avoidance.

JBHI Journal 2025 Journal Article

Characterization of Cortical Connectivity in the Deception State With a Data-Driven Network Model Based on EEG Signal

  • Qianruo Kang
  • Yaqian Li
  • Xiang Li
  • Min Tian
  • Yin Xiang
  • Feng Li
  • Siyu Peng
  • Yijun Xiong

This study investigates the pattern of information interaction at the cortical level during deception, aiming to reveal the cognitive processes involved in the deception task. Our study involves the 64-channel EEG signals of 28 subjects (14 for innocent and 14 for guilty groups) acquired under the guilty knowledge test (GKT) lie-detection protocol. Additionally, we establish the functional connectivity network at the cortical level considering volume conduction effects, use a data-driven approach to select the regions of interest (ROIs) on the subject's cortex based on scalp electrical activity, and perform cortical current density estimation on 15 ROIs. The nonlinear dependence between the cortical waveforms of the ROIs is quantified based on mutual information, and a network of cortical mutual information connections is constructed in four frequency bands: delta, theta, alpha, and beta. The feature extraction and classification process are performed in each frequency band, and the mutual information connections statistically different between the innocent and guilty groups are first selected as features using statistical tests. Moreover, the optimal feature subset (OFS) is found by combining the SVM classifier and the wrapper feature selection strategy. Furthermore, the most important mutual information connections (MIMICs) per frequency band are obtained by refining the OFS according to the classification performance curve. The average test accuracies of MIMICs in the delta, theta, alpha, and beta bands reached 99. 76%, 96. 42%, 84. 04%, and 97. 61%, respectively. Finally, the physiological significance of each frequency sub-band and the physiological function of MIMICs are combined to explore the cognitive mechanism of lies and provide new evidence for cognitive activity in lying states.

JBHI Journal 2025 Journal Article

NFFGRAM: Nonlinear Multi-Feature Fusion and Gated Recurrent Self-Attention Mechanism for Traditional Chinese Medicine Formula Recommendation

  • Hailong Hu
  • Yaqian Li
  • Zhong Li

Traditional Chinese Medicine (TCM) prescriptions are derived from the distinctive thought process and clinical experiences of Chinese medical theory. With the advent of artificial intelligence (AI), there is an enhanced ability to formulate these prescriptions by analyzing symptom data. However, the inherent sparseness of herb-symptom association data still limits the efficacy of such predictive methods. This study introduces an enhanced bipartite graph diffusion algorithm coupled with a gated recurrent self-attention mechanism for predicting herb and symptom associations. The initial phase involves the reconstruction of the herb-symptom association matrix, leveraging the fractal-weighted K-nearest neighbor algorithm. Subsequently, a method is conceived to extract analogous features between herbs and symptoms, which integrates linear neighborhood similarity with Gaussian kernel similarity, both based on fractal dimensions. The next stage employs a modified bipartite graph diffusion to deduce underlying herb-symptom relationships. This process culminates with the integration of the gated recurrent self-attention mechanism and a confidence scoring system to refine the herb-symptom association predictive matrix at a granular level. We benchmark our results against leading-edge algorithms to ascertain the precision and reliability of our model. Such as improvements of precision@20 by 21. 77%, recall@20 by 12. 46%, and F1-score@20 by 19. 28% compared with the best baseline for the TCM2 dataset. Additionally, comprehensive case studies are undertaken, evaluating recommended prescriptions using insights from contemporary medicine and network pharmacology. The proposed model provides a novel paradigm for enhancing herbal prescription methodologies and TCM herb-based treatments.

EAAI Journal 2025 Journal Article

Parallel segmentation network for real-time semantic segmentation

  • Guanke Chen
  • Haibin Li
  • Yaqian Li
  • Wenming Zhang
  • Tao Song

Real-time semantic segmentation holds extensive application prospects in autonomous driving and robot navigation. Recently, real-time semantic segmentation networks mainly adopt encoder-decoder architecture and multi-branch architecture. However, both approaches have their own advantages and limitations. Encoder-decoder models are generally better at extracting contextual information, but may face challenges in capturing fine details and local spatial information. On the other hand, the multi-branch structure excels at capturing boundary and spatial detail information, but it requires an efficient and flexible feature fusion strategy to prevent information redundancy. To leverage the strengths of both approaches, we propose a Parallel Segmentation Network (PaSeNet) which adopts the unsymmetrical encoder-decoder structure to introduce novel ideas for research and applications in real-time semantic segmentation. Specifically, we design a main branch with a spatial information enhancement path during the encoding phase and introduce mask autoencoder based on self-supervised learning as an auxiliary branch to supplement the main branch in extracting details as well as local spatial information. Additionally, we propose the Grouped Aggregation Pyramid Pooling Module to optimize the extraction of contextual information. In the decoding phase, we introduce the Coordinate-Attention-Guided Decoder to effectively integrate diverse information from different branches. A large number of experiments on the Cityscapes, Cambridge-driving Labeled Video database (CamVid), NightCity and instance Segmentation in Aerial Images Dataset demonstrate that our method achieves competitive results. Specifically, PaSeNet-Base obtains 79. 9% mean Intersection Over Union (mIOU) at 55. 6 Frames Per Second (FPS) on Cityscapes test dataset and 80. 2% mIOU at 96. 8 FPS on CamVid test dataset.

JBHI Journal 2025 Journal Article

Variability of Spatiotemporal-Rhythmic Network During Inhibitory Control in Repetitive Subconcussion

  • Xiang Li
  • Zhenghao Fu
  • Hui Zhou
  • Yin Xiang
  • Yaqian Li
  • Yida He
  • Jiaqi Zhang
  • Huanhuan Li

The inhibitory control dysfunction associated with the cognitive symptoms resulting from repetitive subconcussion (SC) is frequent. Implementing inhibitory control is temporally resolved and is likely related to the dynamic interactions in functional brain networks. However, investigations of the dynamic activity of these brain networks using electroencephalography (EEG) are often limited to specific frequency bands without entirely utilizing the spatiotemporal rhythmic information. Therefore, we proposed an innovative framework for constructing a large-scale spatiotemporal-rhythmic network (STRN) using the dynamic cross-frequency phase synchronization to track cognitive deficits induced by repetitive subconcussion during the inhibitory control. Seventeen parachuters with repeated subconcussive exposure and 17 healthy controls (HC) were subjected to a Stroop task while recording the continuous scalp EEG data. Our results indicated an STRN-specific activation pattern that achieved a high classification performance with an average accuracy of 90. 98%, which may serve as a biomarker for identifying the repetitive subconcussion inhibitory control dysfunction. In this STRN state, the SC exhibited mostly lower network rhythmic information interactions than the HC. These findings suggested that the STRN presented in this study could be an effective analytical method for understanding the cognitive dysfunction observed in the repetitive subconcussion and other related conditions.

AAAI Conference 2024 Conference Paper

Debiased Novel Category Discovering and Localization

  • Juexiao Feng
  • Yuhong Yang
  • Yanchun Xie
  • Yaqian Li
  • Yandong Guo
  • Yuchen Guo
  • Yuwei He
  • Liuyu Xiang

In recent years, object detection in deep learning has experienced rapid development. However, most existing object detection models perform well only on closed-set datasets, ignoring a large number of potential objects whose categories are not defined in the training set. These objects are often identified as background or incorrectly classified as pre-defined categories by the detectors. In this paper, we focus on the challenging problem of Novel Class Discovery and Localization (NCDL), aiming to train detectors that can detect the categories present in the training data, while also actively discover, localize, and cluster new categories. We analyze existing NCDL methods and identify the core issue: object detectors tend to be biased towards seen objects, and this leads to the neglect of unseen targets. To address this issue, we first propose an Debiased Region Mining (DRM) approach that combines class-agnostic Region Proposal Network (RPN) and class-aware RPN in a complementary manner. Additionally, we suggest to improve the representation network through semi-supervised contrastive learning by leveraging unlabeled data. Finally, we adopt a simple and efficient mini-batch K-means clustering method for novel class discovery. We conduct extensive experiments on the NCDL benchmark, and the results demonstrate that the proposed DRM approach significantly outperforms previous methods, establishing a new state-of-the-art.

ICLR Conference 2024 Conference Paper

Tag2Text: Guiding Vision-Language Model via Image Tagging

  • Xinyu Huang
  • Youcai Zhang
  • Jinyu Ma
  • Weiwei Tian
  • Rui Feng 0001
  • Yuejie Zhang
  • Yaqian Li
  • Yandong Guo

This paper presents Tag2Text, a vision language pre-training (VLP) framework, which introduces image tagging into vision-language models to guide the learning of visual-linguistic features. In contrast to prior works which utilize object tags either manually labeled or automatically detected with a limited detector, our approach utilizes tags parsed from its paired text to learn an image tagger and meanwhile provides guidance to vision-language models. Given that, Tag2Text can utilize large-scale annotation-free image tags in accordance with image-text pairs, and provides more diverse tag categories beyond objects. Strikingly, Tag2Text showcases the ability of a foundational image tagging model, with superior zero-shot performance even comparable to full supervision manner. Moreover, by leveraging tagging guidance, Tag2Text effectively enhances the performance of vision-language models on both generation-based and alignment-based tasks. Across a wide range of downstream benchmarks, Tag2Text achieves state-of-the-art results with similar model sizes and data scales, demonstrating the efficacy of the proposed tagging guidance.

ECAI Conference 2024 Conference Paper

u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model

  • Jinjin Xu
  • Liwu Xu
  • Yuzhe Yang 0001
  • Xiang Li 0179
  • Fanyi Wang
  • Yanchun Xie
  • Yi-Jie Huang
  • Yaqian Li

Recent advancements in multi-modal large language models (MLLMs) have led to substantial improvements in visual understanding, primarily driven by sophisticated modality alignment strategies. However, predominant approaches prioritize global or regional comprehension, with less focus on fine-grained, pixel-level tasks. To address this gap, we introduce u-LLaVA, an innovative unifying multi-task framework that integrates pixel, regional, and global features to refine the perceptual faculties of MLLMs. We commence by leveraging an efficient modality alignment approach, harnessing both image and video datasets to bolster the model’s foundational understanding across diverse visual contexts. Subsequently, a joint instruction tuning method with task-specific projectors and decoders for end-to-end downstream training is presented. Furthermore, this work contributes a novel mask-based multi-task dataset comprising 277K samples, crafted to challenge and assess the fine-grained perception capabilities of MLLMs. The overall framework is simple, effective, and achieves state-of-the-art performance across multiple benchmarks. We make model, data, and code publicly accessible at https: //github. com/OPPOMKLab/u-LLaVA.

JBHI Journal 2023 Journal Article

Analysis of Weight-Directed Functional Brain Networks in the Deception State Based on EEG Signal

  • Sihong Wei
  • Junfeng Gao
  • Yong Yang
  • Neal Xiong
  • Jiaqi Zhang
  • Jian Song
  • Qianruo Kang
  • Yaqian Li

Although analyzing the brain's functional and structural network has revealed that numerous brain networks are necessary to collaborate during deception, the directionality of these functional networks is still unknown. This study investigated the effective connectivity of the brain networks during deception and uncovers the information-interaction patterns of lying neural oscillations. The electroencephalography (EEG) data of 40 lying persons and 40 honest persons were used to create the weight- directed functional brain networks (WDFBN). Specifically, the connecting edge weight was defined based on the normalized phase transfer entropy (dPTE) between each electrode pair, where the network nodes involved 30 electrode channels. Additionally, the signal connectivity matrices were constructed in four frequency bands: delta, theta, alpha, and beta and were subjected to a difference analysis of entropy values between the groups. Statistical analysis of the classification results revealed that all frequency bands correctly detect deception and innocence with an accuracy of 92. 83%, 94. 17%, 85. 93%, and 92. 25%, respectively. Therefore, dPTE can be considered a valuable feature for identifying lying. According to WDFBN analysis, deception has stronger information flow in the frontoparietal, frontotemporal and temporoparietal networks compare to honest people. Furthermore, the prefrontal cortex was also found to be activated in all frequency ranges. This study examined the critical pathways of brain information interaction during deception, providing new insights into the underlying neural mechanisms. Our analysis offers significant evidence for the development of brain networks that could potentially be used for lie detection.

IJCAI Conference 2023 Conference Paper

Matting Moments: A Unified Data-Driven Matting Engine for Mobile AIGC in Photo Gallery

  • Yanhao Zhang
  • Fanyi Wang
  • Weixuan Sun
  • Jingwen Su
  • Peng Liu
  • Yaqian Li
  • Xinjie Feng
  • Zhengxia Zou

Image matting is a fundamental technique in visual understanding and has become one of the most significant capabilities in mobile phones. Despite the development of mobile storage and computing power, achieving diverse mobile Artificial Intelligence Generated Content (AIGC) applications remains a great challenge. To address this issue, we present an innovative demonstration of an automatic system called "Matting Moments" that enables automatic image editing based on matting models in different scenarios. Coupled with accurate and refined matting subjects, our system provides visual element editing abilities and backend services for distribution and recommendation that respond to emotional expressions. Our system comprises three components: 1) photo content structuring, 2) data-driven matting engine, and 3) AIGC functions for generation, which automatically achieve diverse photo beautification in the gallery. This system offers a unified framework that guides consumers to obtain intelligent recommendations with beautifully generated contents, helping them enjoy the moments and memories of their present life.

ICLR Conference 2023 Conference Paper

Mosaic Representation Learning for Self-supervised Visual Pre-training

  • Zhaoqing Wang
  • Ziyu Chen
  • Yaqian Li
  • Yandong Guo
  • Jun Yu 0001
  • Mingming Gong
  • Tongliang Liu

Self-supervised learning has achieved significant success in learning visual representations without the need for manual annotation. To obtain generalizable representations, a meticulously designed data augmentation strategy is one of the most crucial parts. Recently, multi-crop strategies utilizing a set of small crops as positive samples have been shown to learn spatially structured features. However, it overlooks the diverse contextual backgrounds, which reduces the variance of the input views and degenerates the performance. To address this problem, we propose a mosaic representation learning framework (MosRep), consisting of a new data augmentation strategy that enriches the backgrounds of each small crop and improves the quality of visual representations. Specifically, we randomly sample numbers of small crops from different input images and compose them into a mosaic view, which is equivalent to introducing different background information for each small crop. Additionally, we further jitter the mosaic view to prevent memorizing the spatial locations of each crop. Along with optimization, our MosRep gradually extracts more discriminative features. Extensive experimental results demonstrate that our method improves the performance far greater than the multi-crop strategy on a series of downstream tasks, e.g., +7.4% and +4.9% than the multi-crop strategy on ImageNet-1K with 1% label and 10% label, respectively. Code is available at https://github.com/DerrickWang005/MosRep.git.

AAAI Conference 2022 Conference Paper

On the Efficacy of Small Self-Supervised Contrastive Models without Distillation Signals

  • Haizhou Shi
  • Youcai Zhang
  • Siliang Tang
  • Wenjie Zhu
  • Yaqian Li
  • Yandong Guo
  • Yueting Zhuang

It is a consensus that small models perform quite poorly under the paradigm of self-supervised contrastive learning. Existing methods usually adopt a large off-the-shelf model to transfer knowledge to the small one via distillation. Despite their effectiveness, distillation-based methods may not be suitable for some resource-restricted scenarios due to the huge computational expenses of deploying a large model. In this paper, we study the issue of training self-supervised small models without distillation signals. We first evaluate the representation spaces of the small models and make two non-negligible observations: (i) the small models can complete the pretext task without overfitting despite their limited capacity and (ii) they universally suffer the problem of over clustering. Then we verify multiple assumptions that are considered to alleviate the over-clustering phenomenon. Finally, we combine the validated techniques and improve the baseline performances of five small architectures with considerable margins, which indicates that training small self-supervised contrastive models is feasible even without distillation signals. The code is available at https: //github. com/WOWNICE/ssl-small.

v2026.09.13