EAAI Journal 2026 Journal Article
Integrating whole-slide images and transcriptomic data for survival analysis using multimodal attention networks
- Chunfeng Shao
- Yuanshen Zhao
- Yinsheng Chen
- Jingxian Duan
- Zeyu Zhang
- Rongpin Wang
- Dong Liang
- Dongyue Chen
The integration of multimodal data holds great promise for tumor prognosis prediction by providing a more comprehensive view of the tumor. However, current fusion algorithms, which either merge feature levels or use alignment mechanisms for complex integration, often overlook the relationships between different modalities. To address this issue, we propose a novel multimodal deep learning algorithm that integrates pathological images and transcriptomic data for tumor prognosis prediction. Our model not only uses information from each data modality but also mines the correlations between the two modalities through a Hierarchical Cross-Modal Attention (HCMA) mechanism. In the pathological branch, we designed a self-attention encoder and decoder to extract various levels of contextual features from the pathological images. In the transcriptomic branch, we employed a selective state space module specifically designed to capture the complex dependencies among genetic features. To capture the correlations between the two modalities, we used an HCMA transformer module to generate cross-modal features, including gene-guided pathological features and pathology-guided gene features. Finally, we apply a specific feature alignment mechanism to constrain these features before fusing them to predict survival outcomes. We validated the proposed algorithm on four diverse cancer datasets from The Cancer Genome Atlas (TCGA). For testing the performance of the proposed survival prediction model, we also validated it on an independent cohort of 86 glioma patients from Sun Yat-sen University Cancer Center. Results show the model performs excellently in cancer survival analysis, which outperforms comparative single-modal models and multimodal fusion models.