CIM-VTP:基于关联引导的图像建模与视觉-文本任务提示的通用医学图像配准

CIM-VTP: Correlation-Guided Image Modeling with Visual-Textual Task Prompt for Universal Medical Image Registration

Housheng Xie, Xiaoru Gao, Guoyan Zheng

IEEE TRANSACTIONS ON MEDICAL IMAGING

Abstract

Universal medical image registration through a single model handling various registration tasks has attracted increasing interest. However, existing deep learning-based methods face two major challenges in adapting to universal registration tasks: 1) they lack generalizable feature representation capabilities for cross-task registration; 2) they rely solely on model architectures with fixed parameters, which limits their flexibility to dynamically adapt to different registration tasks and inherently compromises their generalization capability for zero-shot performance onunseen tasks.To address these limitations,we propose CIM-VTP, a novel two-stage universal registration framework. In the first stage, our proposed Correlation guided Image Modeling (CIM)-based pretraining strategy leverages cross-image correlation to guide the masked modeling process,which facilitates spatial correspondence capturing that is essential for registration and provides universal representation capabilities as a foundation for registration learning. In the second stage, we introduce a registration task classifier to identify the type of a given input task, which explicitly quantifies the similarity between current inputs and previously seen tasks. The obtained task similarity scores are then fed as prior information into our carefully designed multi-resolution Visual-Textual Task Prompt (VTP) modules, which integrate task-relevant knowledge through prompt learning to adaptively adjust decoder parameters for different input domains. Extensive experiments across six different registration tasks demonstrate that the proposed CIM-VTP exhibits superior universal image registration performance.

Fig. 1. The overall framework of the proposed method and its zero-shot performance. (A) a schematic illustration of the overall framework of the proposed method, which consists of two stages. In particular, stage I employs correlation-guided image modeling for self-supervised pretraining while stage II performs registration learning based on prompt learning; (B) the comparison of zero-shot performance between our method (Ours) and the state-of-the-art (SOTA) in terms of Dice Similarity Coefficient (DSC) scores. The zero-shot performance of each method is evaluated using a leave-one-domain-out setting where each target domain is excluded from the training set and then tested for zero-shot generalization capability.

 

Fig.2. The Correlation-guided Image Modeling(CIM)framework.

 

 

DOI 10.1109/TMI.2026.3690772

Copyright © 2025上海交通大学医疗机器人研究院 版权所有 沪交ICP备20190057   流量统计