Arrow Research search

Author name cluster

Jindong Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
1 author row

Possible papers

6

AAAI Conference 2026 Conference Paper

Beyond Retraining: Training-Free Unknown Class Filtering for Source-Free Open Set Domain Adaptation of Vision–Language Models

  • Yongguang Li
  • Jindong Li
  • Qi Wang
  • QianLi Xing
  • Runliang Niu
  • Shengsheng Wang
  • Menglin Yang

Vision-language models (VLMs) have gained widespread attention for their strong zero-shot capabilities across numerous downstream tasks. However, these models assume that each test image’s class label is drawn from a predefined label set and lack a reliable mechanism to reject samples from emerging unknown classes when only unlabeled data are available. To address this gap, open-set domain adaptation methods retrain models to push potential unknowns away from known clusters. Yet, some unknown samples remain stably anchored to specific known classes in the VLM feature space due to semantic relevance, which is termed as Semantic Affinity Anchoring (SAA). Forcibly repelling these samples unavoidably distorts the native geometry of VLMs and degrades performance. Meanwhile, existing score‑based unknown detectors use simplistic thresholds and suffer from threshold sensitivity, resulting in sub‑optimal performance. To address aforementioned issues, we propose VLM-OpenXpert, which comprises two training‑free, plug‑and‑play inference modules. SUFF performs SVD on high-confidence unknowns to extract a low-rank "unknown subspace". Each sample’s projection onto this subspace is weighted and softly removed from its feature, suppressing unknown components while preserving semantics. BGAT corrects score skewness via a Box–Cox transform, then fits a bimodal Gaussian mixture to adaptively estimate the optimal threshold balancing known-class recognition and unknown-class rejection. Experiments on 9 benchmarks and three backbones (CLIP, SigLIP, ALIGN) under Source-Free OSDA settings show that our training-free pipeline matches or outperforms retraining-heavy state-of-the-art methods, establishing a powerful lightweight inference calibration paradigm for open-set VLM deployment.

NeurIPS Conference 2025 Conference Paper

STEP: A Unified Spiking Transformer Evaluation Platform for Fair and Reproducible Benchmarking

  • Sicheng Shen
  • Dongcheng Zhao
  • Linghao Feng
  • Zeyang Yue
  • Jindong Li
  • Tenglong Li
  • Guobin Shen
  • Yi Zeng

Spiking Transformers have recently emerged as promising architectures for combining the efficiency of spiking neural networks with the representational power of self-attention. However, the lack of standardized implementations, evaluation pipelines, and consistent design choices has hindered fair comparison and principled analysis. In this paper, we introduce \textbf{STEP}, a unified benchmark framework for Spiking Transformers that supports a wide range of tasks, including classification, segmentation, and detection across static, event-based, and sequential datasets. STEP provides modular support for diverse components such as spiking neurons, input encodings, surrogate gradients, and multiple backends (e. g. , SpikingJelly, BrainCog). Using STEP, we reproduce and evaluate several representative models, and conduct systematic ablation studies on attention design, neuron types, encoding schemes, and temporal modeling capabilities. We also propose a unified analytical model for energy estimation, accounting for spike sparsity, bitwidth, and memory access, and show that quantized ANNs may offer comparable or better energy efficiency. Our results suggest that current Spiking Transformers rely heavily on convolutional frontends and lack strong temporal modeling, underscoring the need for spike-native architectural innovations. The full code is available at: https: //github. com/Fancyssc/STEP.

NeurIPS Conference 2025 Conference Paper

Stratify or Die: Rethinking Data Splits in Image Segmentation

  • Naga Venkata Sai Jitin Jami
  • Thomas Altstidl
  • Jonas Mueller
  • Jindong Li
  • Dario Zanca
  • Bjoern Eskofier
  • Heike Leutheuser

Random splitting of datasets in image segmentation often leads to unrepresentative test sets, resulting in biased evaluations and poor model generalization. While stratified sampling has proven effective for addressing label distribution imbalance in classification tasks, extending these ideas to segmentation remains challenging due to the multi-label structure and class imbalance typically present in such data. Building on existing stratification concepts, we introduce Iterative Pixel Stratification (IPS), a straightforward, label-aware sampling method tailored for segmentation tasks. Additionally, we present Wasserstein-Driven Evolutionary Stratification (WDES), a novel genetic algorithm designed to minimize the Wasserstein distance, thereby optimizing the similarity of label distributions across dataset splits. We prove that WDES is globally optimal given enough generations. Using newly proposed statistical heterogeneity metrics, we evaluate both methods against random sampling and find that WDES consistently produces more representative splits. Applying WDES across diverse segmentation tasks, including street scenes, medical imaging, and satellite imagery, leads to lower performance variance and improved model evaluation. Our results also highlight the particular value of WDES in handling small, imbalanced, and low-diversity datasets, where conventional splitting strategies are most prone to bias.

IJCAI Conference 2024 Conference Paper

ScreenAgent: A Vision Language Model-driven Computer Control Agent

  • Runliang Niu
  • Jindong Li
  • Shiqi Wang
  • Yali Fu
  • Xiyu Hu
  • Xueyuan Leng
  • He Kong
  • Yi Chang

Large Language Models (LLM) can invoke a variety of tools and APIs to complete complex tasks. The computer, as the most powerful and universal tool, could potentially be controlled by a trained LLM agent. Powered by the computer, we can hopefully build a more generalized agent to assist humans in various daily digital works. In this paper, we construct an environment for a Vision Language Model (VLM) agent to interact with a real computer screen. Within this environment, the agent can observe screenshots and manipulate the Graphical User Interface (GUI) by outputting mouse and keyboard actions. We also design an automated control pipeline that includes planning, acting, and reflecting phases, guiding the agent to continuously interact with the environment and complete multi-step tasks. Additionally, we construct the ScreenAgent Dataset, which collects screenshots and action sequences when completing daily computer tasks. Finally, we train a model, ScreenAgent, which achieves comparable computer control capabilities to GPT-4V and demonstrated more precise UI positioning capabilities. Our attempts could inspire further research on building a generalist LLM agent. The code and more detailed information are at https: //github. com/niuzaisheng/ScreenAgent.

EAAI Journal 2024 Journal Article

Unsupervised video forecasting with flow parsing mechanism of human visual system

  • Beibei Jin
  • Xiaohui Song
  • Jindong Li
  • Pengfei Zhang

Video forecasting aims to predict future video frames based on past observed video frames, and unlike object recognition or object classification, it does not require manual labeling of the data set. The explosive growth of Internet video data provides a huge space for its development. At present, it has become a research hotspot in the field of computer vision, and has broad application prospects in the field of automatic driving or robot navigation. However, due to the high dimensional characteristics and the complex spatial–temporal logic of video data, current methods still face the challenges of blurry and inconsistent prediction. The cognitive ability of “flow parsing mechanism” helps humans adapt to new situations systematically. Inspired by this, a deep flow parsing network for future video forecasting is proposed in this paper, which is designed to predict future scenes by parsing optical flow into rigid flow and residual flow. The rigid flow represents the scene dynamics due to observer’s ego-motion, while the residual flow corresponds to the movement of the other objects in the scene. With this procedure, the model exhibits much more comprehensive understanding over the environment and achieves top performance on competitive driving datasets, demonstrating its effectiveness and generalizability.

EAAI Journal 2020 Journal Article

A novel bat algorithm with double mutation operators and its application to low-velocity impact localization problem

  • Qi Liu
  • Jindong Li
  • Lei Wu
  • Fengde Wang
  • Wensheng Xiao

The low-velocity impact localization in the plate structure of the ship is a critical problem which can be considered as a nonlinear optimization problem. The bat algorithm (BA) has been widely used to solve nonlinear optimization problems. However, the standard BA exhibits poor performance on complex problems because of its premature convergence. In this study, a novel bat algorithm with double mutation operators (TMBA), in which a modified time factor and two mutation operators are integrated, is proposed to enhance BA’s performance on nonlinear optimization problems. Classical benchmark functions are employed to analyze the contributions of the three modifications and demonstrate the significant improvement of TMBA. For the low-velocity impact localization problem, the low-velocity impact localization system based on fiber Bragg grating (FBG) sensors is utilized to receive the impact signals. The wavelet threshold de-noising method and the generalized cross-correlation method are both applied to the extraction of time differences between the impact signals. Then, the proposed algorithm and several well-known optimization algorithms are adopted to solve the minimization fitness function which is established using the triangulation method. The statistical results indicate that TMBA is more feasible and effective for solving the low-velocity impact localization problem.

v2026.09.13