Arrow Research search
Back to JBHI

JBHI 2026

Text-Driven Weakly Supervised OCT Lesion Segmentation With Structural Guidance

Journal Article journal-article Artificial Intelligence ยท Biomedical and Health Informatics

Abstract

Accurate segmentation of Optical Coherence Tomography (OCT) images is crucial for diagnosing and monitoring retinal diseases. However, the labor-intensive nature of pixel-level annotation limits the scalability of supervised learning for large datasets. Weakly Supervised Semantic Segmentation (WSSS) offers a promising alternative by using weaker forms of supervision, such as image-level labels, to reduce the annotation burden. Despite its advantages, weak supervision inherently carries limited information. We propose a novel WSSS framework with only image-level labels for OCT lesion segmentation that integrates structural and text-driven guidance to produce high-quality, pixel-level pseudo labels. The framework employs two visual processing modules: one that processes the original OCT images and another that operates on layer segmentations augmented with anomalous signals, enabling the model to associate lesions with their corresponding anatomical layers. Complementing these visual cues, we leverage large-scale pretrained models to provide two forms of textual guidance: label-derived descriptions that encode local semantics, and domain-agnostic synthetic descriptions that, although expressed in natural image terms, capture spatial and relational semantics useful for generating globally consistent representations. By fusing these visual and textual features in a multi-modal framework, our method aligns semantic meaning with structural relevance, thereby improving lesion localization and segmentation performance. Experiments on three OCT datasets demonstrate state-of-the-art results, highlighting its potential to advance diagnostic accuracy and efficiency in medical imaging.

Authors

Keywords

  • Lesions
  • Biomedical imaging
  • Retina
  • Visualization
  • Semantics
  • Annotations
  • Optical coherence tomography
  • Accuracy
  • Semantic segmentation
  • Location awareness
  • Lesion Segmentation
  • Structural Guidance
  • Medical Imaging
  • Visual Features
  • Natural Images
  • Segmentation Accuracy
  • Visual Modality
  • Optical Coherence Tomography Images
  • Segmentation Performance
  • Textual Features
  • Layer Segmentation
  • Pseudo Labels
  • Weak Supervision
  • Pixel-level Annotations
  • Image-level Labels
  • Structural Information
  • Feature Maps
  • Similarity Score
  • Localization Accuracy
  • Class Activation Maps
  • Subretinal Fluid
  • Pigment Epithelial Detachment
  • Global Max Pooling
  • Text Generation
  • Retinal Layer
  • Text Encoder
  • Lesion Classification
  • Encoding Stage
  • Text Labels
  • Multimodal learning
  • retinal OCT lesion segmentation
  • vision-language models
  • weakly supervised semantic segmentation
  • Tomography, Optical Coherence
  • Humans
  • Supervised Machine Learning
  • Image Interpretation, Computer-Assisted
  • Retinal Diseases
  • Algorithms
  • Databases, Factual

Context

Venue
IEEE Journal of Biomedical and Health Informatics
Archive span
2013-2026
Indexed papers
6337
Paper id
472938256458442506
v2026.09.13