Arrow Research search

Author name cluster

Jun Peng

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
1 author row

Possible papers

6

EAAI Journal 2026 Journal Article

Automated crack measurement in slab tracks using deformable instance segmentation and boundary augmentation with unsupervised style transfer

  • Wenbo Hu
  • Zheng Wu
  • Weidong Wang
  • Xianhua Liu
  • Jun Peng

Crack detection and measurement in slab tracks are critical for maintenance decision-making. Pre-trained deep learning segmentation models often struggle with cracking instances due to domain adaptation and data scarcity. This study proposes an instance segmentation framework incorporating dynamic snake convolution (DSConv) modules, combined with an unsupervised style transfer-based boundary augmentation strategy. The DSConv-enhanced architecture prioritizes linear crack features in cluttered backgrounds, while the augmentation introduces controlled perturbations to global pixels and local crack boundaries, generating structurally consistent diversified training samples. The results demonstrate that the deformable DSConv enhanced architecture achieves optimal mean average precision (mAP), improving segmentation performance by nearly 13 % compared to "fine-tuned Segment Anything Model". Its segmentation capability surpasses eight state-of-the-art models, especially for multiple intermittent microcracks. Furthermore, unsupervised style transfer-generated augmented data enhances crack instance segmentation performance by 10 % compared to non-augmented baselines, surpassing conventional methods including horizontal flipping and color jittering. Quantitative crack width distributions from segmentation-quantification analysis provide more comprehensive structural health insights than manual discrete-point measurements, facilitating precise maintenance decisions for railway infrastructure.

AAAI Conference 2026 Conference Paper

KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs

  • Baiyang Song
  • Jun Peng
  • Yuxin Zhang
  • Guangyao Chen
  • Feidiao Yang
  • Jianyuan Guo

Training-free video understanding methods leverage the strong image comprehension capabilities of pre-trained vision language models (VLMs) by treating videos as a sequences of static frames, thus obviating the need for costly video-specific training. However, this paradigm often suffers from severe visual redundancy and high computational overhead, especially when processing long videos. Crucially, existing keyframe selection strategies, especially those based on CLIP similarity, are prone to biases and may inadvertently overlook critical frames, resulting in suboptimal video comprehension. To address these significant challenges, we propose KTV, a novel two-stage framework for efficient and effective training-free video understanding. In the first stage, KTV performs question-agnostic keyframe selection by clustering frame-level visual features, yielding a compact, diverse, and representative subset of frames that mitigates temporal redundancy. In the second stage, KTV applies key visual token selection, pruning redundant or less informative tokens from each selected keyframe based on token importance and redundancy, which significantly reduces the number of tokens fed into the LLM. Extensive experiments on the Multiple-Choice VideoQA task demonstrate that KTV outperforms state-of-the-art training-free baselines while using significantly fewer visual tokens, e.g., only 504 tokens for a 60 min video with 10800 frames, achieving 44.8% accuracy on the MLVU-Test benchmark. In particular, KTV also exceeds several training-based approaches on certain benchmarks.

AIIM Journal 2025 Journal Article

Hybrid approach for drug-target interaction predictions in ischemic stroke models

  • Jing-Jie Peng
  • Yi-Yue Zhang
  • Rui-Feng Li
  • Wen-Jun Zhu
  • Hong-Rui Liu
  • Hui-Yin Li
  • Bin Liu
  • Dong-Sheng Cao

Multiple cell death mechanisms are triggered during ischemic stroke and they are interconnected in a complex network with extensive crosstalk, complicating the development of targeted therapies. We therefore propose a novel framework for identifying disease-specific drug-target interaction (DTI), named strokeDTI, to extract key nodes within an interconnected graph network of activated pathways via leveraging transcriptomic sequencing data. Our findings reveal that the drugs a model can predict are highly representative of the characteristics of the database the model is trained on. However, models with comparable performance yield diametrically opposite predictions in real testing scenarios. Our analysis reveals a correlation between the reported literature on drug-target pairs and their binding scores. Leveraging this correlation, we introduced an additional module to assess the predictive validity of our model for each unique target, thereby improving the reliability of the framework's predictions. Our framework identified Cerdulatinib as a potential anti-stroke drug via targeting multiple cell death pathways, particularly necroptosis and apoptosis. Experimental validation in in vitro and in vivo models demonstrated that Cerdulatinib significantly attenuated stroke-induced brain injury via inhibiting multiple cell death pathways, improving neurological function, and reducing infarct volume. This highlights strokeDTI's potential for disease-specific drug-target identification and Cerdulatinib's potential as a potent anti-stroke drug.

AAAI Conference 2025 Conference Paper

TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning

  • Jingjing Xie
  • Yuxin Zhang
  • Jun Peng
  • Zhaohong Huang
  • Liujuan Cao

Despite the efficiency of prompt learning in transferring vision-language models (VLMs) to downstream tasks, existing methods mainly learn the prompts in a coarse-grained manner where the learned prompt vectors are shared across all categories. Consequently, the tailored prompts often fail to discern class-specific visual concepts, thereby hindering the transferred performance for classes that share similar or complex visual attributes. Recent advances mitigate this challenge by leveraging external knowledge from Large Language Models (LLMs) to furnish class descriptions, yet incurring notable inference costs. In this paper, we introduce TextRefiner, a plug-and-play method to refine the text prompts of existing methods by leveraging the internal knowledge of VLMs. Particularly, TextRefiner builds a novel local cache module to encapsulate fine-grained visual concepts derived from local tokens within the image branch. By aggregating and aligning the cached visual descriptions with the original output of the text branch, TextRefiner can efficiently refine and enrich the learned prompts from existing methods without relying on any external expertise. For example, it improves the performance of CoOp from 71.66% to 76.96% on 11 benchmarks, surpassing CoCoOp which introduced instance-wise feature for text prompts. Equipped with TextRefiner, PromptKD achieves state-of-the-art performance while keep inference efficient.

EAAI Journal 2024 Journal Article

Automated detection and quantification of pavement cracking around manhole

  • Jun Peng
  • Weidong Wang
  • Wenbo Hu
  • Chengbo Ai
  • Xinyue Xu
  • Youyin Shi
  • Jin Wang
  • Zhifa Ran

Damage detection plays an important role in pavement health monitoring and inspection. Unfortunately, research about damage detection of the special component of pavement structures, such as pavement manhole covers, is relatively few. A new pipeline for the detection and quantification of damage around the pavement manhole covers is proposed in this research. In this pipeline, the Attention-enhanced Manhole Detection Model (AMDM) is proposed to detect manhole covers. AMDM achieves an ideal balance between accuracy and speed by eliminating redundant structures and incorporating an attention mechanism. The BCSM (Boundary-enhanced Crack Segmentation Model) is proposed to segment the damage around the manhole cover, and the boundary loss function is used to enhance the fine segmentation ability of the model on the boundary. The MAP (Mean Average Precision) of the manhole covers detection model is 96. 68%, and the MIOU (Mean Intersection Over Union) of the crack segmentation model is 89. 73%. Attributing to this efficient and accurate pipeline, a reasonable damage evaluation method is proposed in the end, which is based on statistical data and engineering experience. Overall, not only this research will contribute an automatic and cost-effective method to the detection and evaluation of damage around the manhole cover, but also inspire the detection and evaluation of other special components of civil engineering structures.

IJCAI Conference 2024 Conference Paper

Efficient Event Stream Super-Resolution with Recursive Multi-Branch Fusion

  • Quanmin Liang
  • Zhilin Huang
  • Xiawu Zheng
  • Feidiao Yang
  • Jun Peng
  • Kai Huang
  • Yonghong Tian

Current Event Stream Super-Resolution (ESR) methods overlook the redundant and complementary information present in positive and negative events within the event stream, employing a direct mixing approach for super-resolution, which may lead to detail loss and inefficiency. To address these issues, we propose an efficient Recursive Multi-Branch Information Fusion Network (RMFNet) that separates positive and negative events for complementary information extraction, followed by mutual supplementation and refinement. Particularly, we introduce Feature Fusion Modules (FFM) and Feature Exchange Modules (FEM). FFM is designed for the fusion of contextual information within neighboring event streams, leveraging the coupling relationship between positive and negative events to alleviate the misleading of noises in the respective branches. FEM efficiently promotes the fusion and exchange of information between positive and negative branches, enabling superior local information enhancement and global information complementation. Experimental results demonstrate that our approach achieves over 17% and 31% improvement on synthetic and real datasets, accompanied by a 2. 3x acceleration. Furthermore, we evaluate our method on two downstream event-driven applications, i. e. , object recognition and video reconstruction, achieving remarkable results that outperform existing methods. Our code and Supplementary Material are available at https: //github. com/Lqm26/RMFNet.

v2026.09.13