Versatile Vision-Language Model for 3D Computed Tomography

Jiayu Lei; Ziqing Fan; Yanyong Zhang; Weidi Xie; Ya Zhang; Yanfeng Wang

doi:10.1609/aaai.v40i8.37517

Back to AAAI

AAAI 2026

Versatile Vision-Language Model for 3D Computed Tomography

Conference Paper AAAI Technical Track on Computer Vision V Artificial Intelligence

PDF Details DOI

Abstract

Representation learning serves as a foundational component of medical vision-language models (MVLMs), enabling cross-modal alignment, semantic consistency, and enhanced generalization capabilities for downstream tasks. As generalist models rapidly evolve, there is a pressing need to unify diverse downstream tasks, such as diagnosis, segmentation, report generation, and multiple choice within a cohesive framework, demanding more efficient and versatile visual representation learning. However, current MVLMs predominately follow CLIP-style vision pretraining, failing to leverage heterogeneous data resources with multi-dimensional imaging and diverse annotation forms. And there lacks systematic analysis of efficient vision encoder design across varied downstream applications, including diagnosis, segmentation, and text generation tasks, particularly for volumetric imaging like Computed Tomography (CT). Besides, current MVLMs exhibit constrained voxel-level capabilities, lacking effective multi-task instruction tuning framework capable of achieving robust performance across various downstream tasks. To address these challenges, we propose CTInstruct, a novel MVLM employing a hybrid ResNet-ViT encoder with multi-granular vision-language pretraining for efficient heterogeneous data modeling, and unified instruction tuning that jointly optimizes discriminative, generative, and voxel-level reasoning for volumetric medical imaging. CTInstruct achieves SOTA performance across 8 CT benchmarks, setting a new standard for data-efficient multimodal learning in medical imaging.

Versatile Vision-Language Model for 3D Computed Tomography

Abstract

Authors

Keywords

Context