EAAI Journal 2026 Journal Article
Cross-Granularity Fusion Vision Mamba UNet for medical image segmentation
- Tuersunjiang Baidi
- Zitong Ren
- Kurban Ubul
- Alimjan Aysa
- Boyuan Li
- Shihao Wang
Recently, state space models (SSMs), represented by Mamba, have shown significant potential in medical image segmentation. However, the inherent axial sequential scanning mechanism limits the modeling of complex spatial relationships, and single-scale processing hinders effective cross-granularity feature interaction. Therefore, this paper proposes a novel hybrid network, named the Cross-Granularity Fusion Vision Mamba UNet (CGFM-UNet), whose core is the Cross-Granularity Fusion Vision State Space (CGF-VSS) block. Specifically, within CGF-VSS, the Multi-Scale Focal Enhancement (MFE) module decomposes the input features into a fine-grained structure-aware branch and a coarse-grained semantics-aware branch. Subsequently, the novel Chess-trajectory Stepwise Selective Scan (CTSt-SS) block leverages this dual-stream information to guide Mamba along spatially interleaved paths, enabling the deep modeling of structured long-range dependencies. Finally, the Dynamic Gating Fusion (DGF) module adaptively aggregates the enhanced features to form a discriminative unified cross-granularity representation. Extensive experiments on three datasets from different imaging modalities demonstrate that CGFM-UNet achieves highly competitive performance, providing an effective solution for complex medical image segmentation and advancing medical artificial intelligence.