Arrow Research search

Author name cluster

Shijie Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

13 papers
2 author rows

Possible papers

13

AAAI Conference 2026 Conference Paper

CogStream: Context-guided Streaming Video Question Answering

  • Zicheng Zhao
  • Kangyu Wang
  • Shijie Li
  • Rui Qian
  • Weiyao Lin
  • Huabin Liu

Despite advancements in Video Large Language Models (Vid-LLMs) improving multimodal understanding, challenges persist in streaming video reasoning due to its reliance on contextual information. Existing paradigms feed all available historical contextual information into Vid-LLMs, resulting in a significant computational burden for visual data processing. Furthermore, the inclusion of irrelevant context distracts models from key details. This paper introduces a challenging task called Context-guided Streaming Video Reasoning (CogStream), which simulates real-world streaming video scenarios, requiring models to identify the most relevant historical contextual information to deduce answers for questions about the current stream. To support CogStream, we present a densely annotated dataset featuring extensive and hierarchical question-answer pairs, generated by a semi-automatic pipeline. Additionally, we present CogReasoner as a baseline model. It effectively tackles this task by leveraging visual stream compression and historical dialogue retrieval. Extensive experiments prove the effectiveness of this method.

AAAI Conference 2026 Conference Paper

SSR-SAM: Retrieval-Style Segment Anything Model for Semi-Supervised Ultra-High-Resolution Image Segmentation

  • Shijie Li
  • Yiming Chen
  • Zhineng Chen
  • Kai Hu
  • Xieping Gao

Accurate segmentation of ultra-high-resolution (UHR) images, which often exceed tens of millions of pixels, is critically important in domains such as remote sensing and biomedical imaging. However, acquiring pixel-level annotations for such high-resolution images is prohibitively expensive and labor-intensive. While semi-supervised semantic segmentation can significantly reduce the annotation burden, its extension to UHR images holds great potential for addressing the unique challenges posed by sparse supervision. To this end, we propose SSR-SAM, a retrieval-style semi-supervised segmentation framework tailored for UHR images. Leveraging the promptable paradigm of the Segment Anything Model (SAM), SSR-SAM treats locally annotated regions as prompts to retrieve semantically consistent pixels across the entire image. Building upon this retrieval-style segmentation paradigm, we further introduce prompt-level perturbation, a novel trail to deploy consistency regularization for semi-supervised segmentation. It encourages the model to learn consistency across predictions guided by diverse visual-semantic prompts, thereby enhancing generalization on unlabeled data. We evaluate SSR-SAM on three UHR datasets: Inria Aerial, BCSS, and URUR. Experimental results show that SSR-SAM achieves clear performance gains over the labeled-only supervision, with average mIoU improvements of 4.9%, 4.15%, and 2.5%, respectively. Additionally, SSR-SAM possesses zero-shot segmentation capability, exhibiting potential for general retrieval-style segmentation tasks.

EAAI Journal 2025 Journal Article

A pre-trained multi-step prediction informer for ship motion prediction with a mechanism-data dual-driven framework

  • Wenhe Shen
  • Xinjue Hu
  • Jialun Liu
  • Shijie Li
  • Hongdong Wang

The advancement of autonomous maritime surface ships has increased the need for accurate and rapid multi-step prediction of ship motion for decision-making, motion planning, and real-time control tasks. This paper proposes a multi-step prediction method based on Informer with a pre-trained strategy to achieve accurate and fast motion prediction for ships, which substitutes generative inference for rolling prediction to avoid the cumulative error caused by the increasing time horizon. Due to the difference in temporal features from long-term control actions and short-term state sequences, heterogeneous inputs of encoder and decoder are designed to respectively capture their information without information redundancy. To address the bottleneck between the high cost of real data acquisition and the high demand for deep learning methods for data, we propose a mechanism-data dual-driven framework. This framework utilizes a prior mechanism model to generate virtual data incorporating a range of excitation signals designed in accordance with the results of free-running model tests. To reduce the need for real data and increase interpretability, the improved Informer is pre-trained by virtual data from the mechanism model before being trained by real data. Our experiments for multi-step ship motion prediction demonstrate that the proposed method respectively reduces the error and time to 41. 36% and 13. 20% on average compared to state-of-the-art and classical methods.

ICML Conference 2025 Conference Paper

Average Sensitivity of Hierarchical k-Median Clustering

  • Shijie Li
  • Weiqiang He
  • Ruobing Bai
  • Pan Peng 0001

Hierarchical clustering is a widely used method for unsupervised learning with numerous applications. However, in the application of modern algorithms, the datasets studied are usually large and dynamic. If the hierarchical clustering is sensitive to small perturbations of the dataset, the usability of the algorithm will be greatly reduced. In this paper, we focus on the hierarchical $k$ -median clustering problem, which bridges hierarchical and centroid-based clustering while offering theoretical appeal, practical utility, and improved interpretability. We analyze the average sensitivity of algorithms for this problem by measuring the expected change in the output when a random data point is deleted. We propose an efficient algorithm for hierarchical $k$-median clustering and theoretically prove its low average sensitivity and high clustering quality. Additionally, we show that single linkage clustering and a deterministic variant of the CLNSS algorithm exhibit high average sensitivity, making them less stable. Finally, we validate the robustness and effectiveness of our algorithm through experiments.

ECAI Conference 2025 Conference Paper

DAT-CRF: Improving the Directed Acyclic Transformer with CRF Integration

  • Shijie Li
  • Inigo Jauregi Unanue
  • Massimo Piccardi

Non-autoregressive transformers (NATs) have demonstrated significant potential in reducing decoding latency for language generation tasks. However, “vanilla” NATs often struggle to effectively capture the sequential structure of generated text. To address this limitation, the recently proposed Directed Acyclic Transformer (DAT) introduces a graph structure into the decoder, explicitly modeling state transitions within the graph. Although DAT has achieved impressive performance, it typically requires large graph sizes to attain optimal results, which are highly memory-intensive, thereby limiting its applicability to tasks involving lengthy output sentences such as document-level machine translation. To mitigate this issue, we propose reducing DAT’s reliance on large graph sizes by coupling its state transition model with a more robust observation model—a Conditional Random Field (CRF). The CRF inherently models pairwise transitions between output tokens, enabling the model to capture dependencies without relying on large graphs. In addition, unlike other NAT-CRF models where the NAT and CRF modules operate independently, our approach is the first to jointly decode both modules, permitting joint optimality of the inferred graph and the output tokens. Experimental results on both sentence-level and document-level machine translation show that this modification substantially improves the baseline DAT in both lexical and semantic metrics, while retaining near-parity of decoding speed.

IROS Conference 2025 Conference Paper

Design and Dynamic Modeling Analysis of Undulatory Propulsion Underwater Robot with Rotational Passive Degrees of Freedom in Fin Rays

  • Tangjia Zhang
  • Qiao Hu
  • Shijie Li
  • Yangbin Zeng
  • Siyu Zu
  • Liangjie Sun

Current research on undulatory propulsion robots has predominantly centered on hydrodynamic performance simulations. However, challenges such as limited mobility and difficulties in parameter identification during underwater bio-mimetic motion remain unresolved. To address these issues, this study proposes a novel undulating fin robot featuring passive rotational joints, aiming to enhance motion capabilities and facilitate more accurate modeling. These joints enhance both the agility and stability of the robot's movements. Initially, the research develops models for the undulatory motion of the undulating fin and the rotational passive degrees of freedom in the fin rays. Based on fluid drag theory, a hydrodynamic model for undulating fin propulsion is constructed to analyze the thrust, lateral force, and lift generated at varying frequencies. Furthermore, a comprehensive dynamics model for the underwater motion of the biomimetic undulating fin robot is developed. Numerical simulations of the robot's non-steady-state motion are conducted to identify the hydrodynamic parameters of the model, thereby enabling the solution of the dynamic model. The experimental results demonstrate that the robot achieves an underwater straight-line motion speed exceeding 0. 5m/s, a turning speed of approximately 45°/s, and an inclined upward motion speed of 0. 21 m/s. This study provides a novel approach for the design of underwater undulating fin robots and the resolution of kinematic models for underwater robots. It is hoped that this research can contribute to the further development of undulatory propulsion robot technology.

AAAI Conference 2025 Conference Paper

Unsupervised Photometric-Consistent Depth Estimation from Endoscopic Monocular Video

  • Shijie Li
  • Weijun Lin
  • Qingyuan Xiang
  • Yunbin Tu
  • Shitan Asu
  • Zheng Li

Recent advancements in unsupervised monocular depth estimation typically rely on an assumption that image photometry remains consistent across consecutive frames. However, this assumption often fails in endoscopic scenes due to: 1) local photometric inconsistency caused by specular reflections creating highlights; and 2) global photometric inconsistency resulting from the simultaneous movement of the light source and the camera. Since unsupervised depth estimation methods rely on appearance discrepancies between frames as a supervisory signal, these photometric inconsistencies inevitably deteriorate loss function calculation. In this paper, our goal is to obtain a strong and reliable supervisory signal for achieving photometric-consistent depth estimation. To this end, for local photometric inconsistency, we utilize the specular reflection model to introduce a Highlight Loss for handling the estimation of highlight regions. For global photometric inconsistency, we design a Photometric Match module, which utilizes the spotlight illumination model to derive an analytical expression, achieving photometric alignment across different frames. Unlike previous works that introduce additional optical flow or networks, our method is simpler and more efficient. Extensive experiments demonstrate our method achieves the state-of-the-art results on C3VD, SCARED and SERV-CT datasets.

ECAI Conference 2024 Conference Paper

Dual Attention Encoder with Joint Preservation for Medical Image Segmentation

  • Shijie Li
  • Yunbin Tu
  • Yu Gong
  • Bowen Zhong
  • Zheng Li

Transformers have recently gained considerable popularity for capturing long-range dependencies in the medical image segmentation. However, most transformer-based segmentation methods primarily focus on modeling global dependencies and fail to fully explore the complementary nature of different dimensional dependencies within features. These methods simply treat the aggregation of multi-dimensional dependencies as auxiliary modules for incorporating context into the Transformer architecture, thereby limiting the model’s capability to learn rich feature representations. To address this issue, we introduce the Dual Attention Encoder with Joint Preservation (DANIE) for medical image segmentation, which synergistically aggregates spatial-channel dependencies across both local and global areas through attention learning. Additionally, we design a lightweight aggregation mechanism, termed Joint Preservation, which learns a composite feature representation, allowing different dependencies to complement each other. Without bells and whistles, our DANIE significantly improves the performance of previous state-of-the-art methods on five popular medical image segmentation benchmarks, including Synapse, ACDC, ISIC 2017, ISIC 2018 and GlaS.

YNIMG Journal 2023 Journal Article

Distinct brain state dynamics of native and second language processing during narrative listening in late bilinguals

  • Xiangrong Tang
  • Juan Zhang
  • Lanfang Liu
  • Menghan Yang
  • Shijie Li
  • Jie Chen
  • Yumeng Ma
  • Jia Zhang

The process of complex cognition, which includes language processing, is dynamic in nature and involves various network modes or cognitive modes. This dynamic process can be manifested by a set of brain states and transitions between them. Previous neuroimaging studies have shed light on how bilingual brains support native language (L1) and second language (L2) through a shared network. However, the mechanism through which this shared brain network enables L1 and L2 processing remains unknown. This study examined this issue by testing the hypothesis that L1 and L2 processing is associated with distinct brain state dynamics in terms of brain state integration and transition flexibility. A group of late Chinese-English bilinguals was scanned using functional magnetic resonance imaging (fMRI) while listening to eight short narratives in Chinese (L1) and English (L2). Brain state dynamics were modeled using the leading eigenvector dynamic analysis framework. The results show that L1 processing involves more integrated states and frequent transitions between integrated and segregated states, while L2 processing involves more segregated states and fewer transitions. Our work provides insight into the dynamic process of narrative listening comprehension in late bilinguals and sheds new light on the neural representation of language processing and related disorders.

ICRA Conference 2022 Conference Paper

Enhanced Spatial Attention Graph for Motion Planning in Crowded, Partially Observable Environments

  • Weixian Shi
  • Yanying Zhou
  • Xiangyu Zeng 0002
  • Shijie Li
  • Maren Bennewitz

Collision-free navigation while moving amongst static and dynamic obstacles with a limited sensor range is still a great challenge for modern mobile robots. Therefore, the ability to avoid collisions with obstacles in crowded, partially observable environments is one of the most important indicators to measure the navigation performance of a mobile robot. In this paper, we propose a novel deep reinforcement learning architecture that combines a spatial graph and attention rea-soning to tackle this problem. We take the relative positions and velocities of observed humans as nodes of the spatial graph and robot-human pairs as nodes of the attention graph to capture the spatial relations between the robot and the humans. In this way, our approach enhances the modeling of the relationship between the moving robot, static obstacles, and the people in the surrounding. As a result, our proposed navigation framework significantly outperforms state-of-the-art approaches [1], [2] in crowded scenarios when the robot has only a limited sensor range in terms of a reduced collision rate. Furthermore, we realize a seriously decreased training time by applying parallel Double Deep Q-Learning.

AAAI Conference 2021 Conference Paper

MiniSeg: An Extremely Minimum Network for Efficient COVID-19 Segmentation

  • Yu Qiu
  • Yun Liu
  • Shijie Li
  • Jing Xu

The rapid spread of the new pandemic, i. e. , COVID-19, has severely threatened global health. Deep-learning-based computer-aided screening, e. g. , COVID-19 infected CT area segmentation, has attracted much attention. However, the publicly available COVID-19 training data are limited, easily causing overfitting for traditional deep learning methods that are usually data-hungry with millions of parameters. On the other hand, fast training/testing and low computational cost are also necessary for quick deployment and development of COVID-19 screening systems, but traditional deep learning methods are usually computationally intensive. To address the above problems, we propose MiniSeg, a lightweight deep learning model for efficient COVID-19 segmentation. Compared with traditional segmentation methods, MiniSeg has several significant strengths: i) it only has 83K parameters and is thus not easy to overfit; ii) it has high computational efficiency and is thus convenient for practical deployment; iii) it can be fast retrained by other users using their private COVID-19 data for further improving performance. In addition, we build a comprehensive COVID-19 segmentation benchmark for comparing MiniSeg to traditional methods.

TIST Journal 2019 Journal Article

A Visual Analysis Approach for Understanding Durability Test Data of Automotive Products

  • Ying Zhao
  • Lei Wang
  • Shijie Li
  • Fangfang Zhou
  • Xiaoru Lin
  • Qiang Lu
  • Lei Ren

People face data-rich manufacturing environments in Industry 4.0. As an important technology for explaining and understanding complex data, visual analytics has been increasingly introduced into industrial data analysis scenarios. With the durability test of automotive starters as background, this study proposes a visual analysis approach for understanding large-scale and long-term durability test data. Guided by detailed scenario and requirement analyses, we first propose a migration-adapted clustering algorithm that utilizes a segmentation strategy and a group of matching-updating operations to achieve an efficient and accurate clustering analysis of the data for starting mode identification and abnormal test detection. We then design and implement a visual analysis system that provides a set of user-friendly visual designs and lightweight interactions to help people gain data insights into the test process overview, test data patterns, and durability performance dynamics. Finally, we conduct a quantitative algorithm evaluation, case study, and user interview by using real-world starter durability test datasets. The results demonstrate the effectiveness of the approach and its possible inspiration for the durability test data analysis of other similar industrial products.

EAAI Journal 2016 Journal Article

Distributed constraint optimization for addressing vessel rotation planning problems

  • Shijie Li
  • Rudy R. Negenborn
  • Gabriel Lodewijks

A distributed constraint optimization problem (DCOP) is a description of constraint optimization problem where variables and constraints are distributed among a group of agents, and where each agent can only interact with agents that share constraints. Even though DCOPs have been studied since the 1990s, there are only a few attempts to address real world problems using this formalism, mainly because of the complexity of the solution algorithms. In this paper, we compare 4 state-of-the-art DCOP approaches to solve the vessel rotation planning problem (VRPP), which concerns deciding on the optimal sequence of vessel visits to different terminals in a large seaport. We hereby also consider two agent structures: a single layer and a multi-layer structure. For each of the structures, we compare the four different algorithms for solving DCOPs, aiming at studying how the algorithms perform in VRPPs of increasing sizes. We assess the methods based on the size and quantity of messages exchanged, computation time, and quality of solutions.

v2026.09.13