Arrow Research search
Back to JBHI

JBHI 2023

Multi-Scale Efficient Graph-Transformer for Whole Slide Image Classification

Journal Article journal-article Artificial Intelligence ยท Biomedical and Health Informatics

Abstract

The multi-scale information among the whole slide images (WSIs) is essential for cancer diagnosis. Although the existing multi-scale vision Transformer has shown its effectiveness for learning multi-scale image representation, it still cannot work well on the gigapixel WSIs due to their extremely large image sizes. To this end, we propose a novel Multi-scale Efficient Graph-Transformer (MEGT) framework for WSI classification. The key idea of MEGT is to adopt two independent efficient Graph-based Transformer (EGT) branches to process the low-resolution and high-resolution patch embeddings (i. e. , tokens in a Transformer) of WSIs, respectively, and then fuse these tokens via a multi-scale feature fusion module (MFFM). Specifically, we design an EGT to efficiently learn the local-global information of patch tokens, which integrates the graph representation into Transformer to capture spatial-related information of WSIs. Meanwhile, we propose a novel MFFM to alleviate the semantic gap among different resolution patches during feature fusion, which creates a non-patch token for each branch as an agent to exchange information with another branch by cross-attention mechanism. In addition, to expedite network training, a new token pruning module is developed in EGT to reduce the redundant tokens. Extensive experiments on both TCGA-RCC and CAMELYON16 datasets demonstrate the effectiveness of the proposed MEGT.

Authors

Keywords

  • Transformers
  • Cancer
  • Feature extraction
  • Spatial resolution
  • Fuses
  • Semantics
  • Task analysis
  • Medical diagnosis
  • Slide Images
  • Cancer Diagnosis
  • Feature Fusion
  • Efficient Learning
  • Multi-scale Features
  • Digital Pathology
  • Multi-scale Information
  • Multi-scale Representation
  • Vision Transformer
  • Multi-scale Feature Fusion
  • Semantic Gap
  • Local Information
  • Feature Representation
  • Input Features
  • Global Information
  • Transformer Model
  • Graph Convolutional Network
  • Computer-aided Diagnosis
  • Self-supervised Learning
  • Feature Pyramid
  • Multiple Instance Learning
  • Transformer Encoder
  • Attention Scores
  • Representation For Classification
  • Heterogeneous Graph
  • Linear Projection
  • Transformer Layers
  • Hierarchical Representation
  • Multi-head Self-attention
  • Adjacent Patches
  • cross-attention
  • graph-Transformer
  • Whole slide images
  • Humans
  • Electric Power Supplies

Context

Venue
IEEE Journal of Biomedical and Health Informatics
Archive span
2013-2026
Indexed papers
6337
Paper id
488895389343081904
v2026.09.13