Arrow Research search

Author name cluster

Xinliang Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

5 papers
2 author rows

Possible papers

5

JBHI Journal 2026 Journal Article

Multi-level Asymmetric Contrastive Learning for Medical Image Segmentation Pre-training

  • Shuang Zeng
  • Lei Zhu
  • Xinliang Zhang
  • Qian Chen
  • Hangzhou He
  • Lujia Jin
  • Zifeng Tian
  • Zhaoheng Xie

Medical image segmentation is a fundamental yet challenging task due to the arduous process of acquiring large volumes of high-quality labeled data from experts. Contrastive learning offers a promising but still problematic solution to this dilemma. Firstly existing medical contrastive learning strategies focus on extracting image-level representation, which ignores abundant multi-level representations. Furthermore they underutilize the decoder either by random initialization or separate pre-training from the encoder, thereby neglecting the potential collaboration between the encoder and decoder. To address these issues, we propose a novel multi-level asymmetric contrastive learning framework named MACL for enhancing medical image segmentation. Specifically, we design an asymmetric contrastive learning structure to pre-train encoder and decoder simultaneously to provide better initialization for segmentation models. Moreover, we develop a multi-level contrastive learning strategy that integrates correspondences across feature-level, image-level, and pixel-level representations to ensure the encoder and decoder capture comprehensive details from representations of varying scales and granularities during the pre-training phase. Finally, experiments on 8 medical image datasets indicate our MACL framework outperforms existing 11 contrastive learning strategies. i. e. Our MACL achieves a superior performance with more precise predictions from visualization figures and 1. 72%, 7. 87%, 2. 49% and 1. 48% Dice higher than previous best results on ACDC, MMWHS, HVSMR and CHAOS with 10% labeled data, respectively. And our MACL also has a strong generalization ability among 5 variant U-Net backbones.

EAAI Journal 2025 Journal Article

An improved you only look once model for the multi-scale steel surface defect detection with multi-level alignment and cross-layer redistribution features

  • Jianhang Huang
  • Xinliang Zhang
  • Lijie Jia
  • Yitian Zhou

Steel surface defects involve a wide variety of sizes and irregular shapes. The performance of the detection models depends greatly on the effective extraction of the cross-scale features. The sequential feature fusion in the traditional Feature Pyramid Network (FPN) is subject to an information loss during the aggregation and transition across shallow and deep features. The classification and localization of steel surface defects using common You Only Look Once (YOLO) detection model may lead to sub-optimal performance when they are involved with the multi-scale objects. To realize a complete shallow-deep feature representation for different scale steel surface defects, an aggregation-redistribution network is introduced into the YOLO detection model to aggregate and refine features across different levels. In the aggregation sub-network, a Multi-level Alignment Module (MAM) is adopted to address the feature misalignment in FPN by aligning the feature maps from level-wise extractors. Therein, the scale deviation at pixels is compensated by the multiple parallel dilated convolutions. Meanwhile, in the redistribution sub-network, a Fusion-Redistribution Module (FRM) is constructed to impose the global information on the fused multi-level features to steer cross-layerly the generation of prediction feature maps for YOLO. The global feature provides a semantic adaptive weight and serves as a complement on the multi-level features through an attention mechanism. Finally, by incorporating the aggregation-redistribution network into YOLOv5, an improved YOLO detection model, i. e. , Aggregation-Redistribution YOLO (ARYOLO), is derived for the steel surface defects. The validation results indicate that ARYOLO has achieved a rather satisfactory detection performance for the defects with large variations in scale and shape. It provides a mean average precision of 80. 7% on the North Eastern University Surface Defect Dataset (NEU-DET) and 71. 2% on the Global 10 Class Metallic Surface Defect Dataset (GC10-DET), giving a great potential in detection and localization tasks.

AAAI Conference 2025 Conference Paper

V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept Tokenizer

  • Hangzhou He
  • Lei Zhu
  • Xinliang Zhang
  • Shuang Zeng
  • Qian Chen
  • Yanye Lu

Concept Bottleneck Models (CBMs) offer inherent interpretability by initially translating images into human-comprehensible concepts, followed by a linear combination of these concepts for classification. However, the annotation of concepts for visual recognition tasks requires extensive expert knowledge and labor, constraining the broad adoption of CBMs. Recent approaches have leveraged the knowledge of large language models to construct concept bottlenecks, with multimodal models like CLIP subsequently mapping image features into the concept feature space for classification. Despite this, the concepts produced by language models can be verbose and may introduce non-visual attributes, which hurts accuracy and interpretability. In this study, we investigate to avoid these issues by constructing CBMs directly from multimodal models. To this end, we adopt common words as base concept vocabulary and leverage auxiliary unlabeled images to construct a Vision-to-Concept (V2C) tokenizer that can explicitly quantize images into their most relevant visual concepts, thus creating a vision-oriented concept bottleneck tightly coupled with the multimodal model. This leads to our V2C-CBM which is training efficient and interpretable with high accuracy. Our V2C-CBM has matched or outperformed LLM-supervised CBMs on various visual classification benchmarks, validating the efficacy of our approach.

AAAI Conference 2024 Conference Paper

Scribble Hides Class: Promoting Scribble-Based Weakly-Supervised Semantic Segmentation with Its Class Label

  • Xinliang Zhang
  • Lei Zhu
  • Hangzhou He
  • Lujia Jin
  • Yanye Lu

Scribble-based weakly-supervised semantic segmentation using sparse scribble supervision is gaining traction as it reduces annotation costs when compared to fully annotated alternatives. Existing methods primarily generate pseudo-labels by diffusing labeled pixels to unlabeled ones with local cues for supervision. However, this diffusion process fails to exploit global semantics and class-specific cues, which are important for semantic segmentation. In this study, we propose a class-driven scribble promotion network, which utilizes both scribble annotations and pseudo-labels informed by image-level classes and global semantics for supervision. Directly adopting pseudo-labels might misguide the segmentation model, thus we design a localization rectification module to correct foreground representations in the feature space. To further combine the advantages of both supervisions, we also introduce a distance entropy loss for uncertainty reduction, which adapts per-pixel confidence weights according to the reliable region determined by the scribble and pseudo-label's boundary. Experiments on the ScribbleSup dataset with different qualities of scribble annotations outperform all the previous methods, demonstrating the superiority and robustness of our method. The code is available at https://github.com/Zxl19990529/Class-driven-Scribble-Promotion-Network.

ICRA Conference 2022 Conference Paper

Brick Yourself within 3 Minutes

  • Guyue Zhou
  • Liyi Luo
  • Hao Xu
  • Xinliang Zhang
  • Haole Guo
  • Hao Zhao 0002

This paper presents an intelligent machine which can automatically convert the captured portrait into a physical gadget made up of LEGO bricks. On the contrary to synthesising a 2D image or a virtual 3D object, generating physical 3D assembly object needs to take physical properties and assembly process into consideration, leading to more challenges. To generate brick models for arbitrary portraits, we formulate the transformation between the attribute space (extracted from 2D images) and the brick model space as a constraint integer programming problem which can be solved with a heuristic search method. Furthermore, as the bricks are physically scattered, we propose an algorithm to generate corresponding assembly instructions for customized figure-featured-bricks to facilitate users' assembly. Meanwhile, we deploy the proposed algorithms on an automatic machine which integrates a camera, a printer, a laptop, and a brick operation unit. Finally, the generated brick models and assembly instructions are evaluated by a large number of users. It is worth noting that the whole system works as an intelligent vending machine, producing a 150-brick-model within 3 minutes.

v2026.09.13