EAAI 2026
Multimodal and multiscale learning network for efficient polyp segmentation
Abstract
Colon cancer has become the second leading cause of cancer-related deaths worldwide, resulting in more than 900 000 deaths each year. Accurate segmentation of intestinal polyps is a key step in preventing the progression of colorectal cancer; however, the robustness of existing methods in complex scenarios still needs improvement. To address this issue, we propose an efficient analysis network that integrates multi-channel and multiscale features, termed the Multimodal and Multiscale Learning Network (MML-Net). Built on the encoder architecture of the Pyramid Vision Transformer (PVT), MML-Net employs a Multi-Channel Feature Fusion Module (MCF) and a Multiscale Parallel Attention Module (MPA) to achieve cross-level feature fusion and deep, fine-grained representation learning, and introduces a Local Attention (LA) mechanism to enhance local context modeling. Experiments show strong learning and generalization capabilities, yielding Dice coefficients of 0. 934 and 0. 925 on the internal datasets Colonoscopy Vision Clinic Database and Kvasir Dataset, and 0. 830 and 0. 813 on the challenging external datasets Colonoscopy Vision Colon Database and ETIS Larib Polyp Database, respectively. The model contains only 34. 4 million (M) parameters and requires 13. 9 giga floating-point operations (GFLOPs), achieving an excellent accuracy–efficiency balance. Implemented artificial-intelligence techniques include PVT, MCF, MPA and LA. This work demonstrates the application of artificial intelligence to medical image analysis for intestinal-polyp segmentation in colonoscopy.
Authors
Keywords
Context
- Venue
- Engineering Applications of Artificial Intelligence
- Archive span
- 1988-2026
- Indexed papers
- 13269
- Paper id
- 923023858195653418