Multi-modal prompt tuning in vision-language models like CLIP has significantly enhanced machine comprehension and interaction with multi-modal content. However, prior works have primarily focused on unidirectional semantic alignment across modalities, limiting their effectiveness in dynamic and open-world visual concept comprehension due to loosely constrained mappings. To address this limitation, we introduce a novel Bidirectional Multi-modal Spatial-consistent Prompt Tuning (BMSPT) framework to ensure a dynamic and fine-grained alignment between text and image modalities. Specifically, our proposed framework facilitates bidirectional interactions between modalities through two novel Vision-Language (V-L) and Language-Vision (L-V) transfer functions. This bidirectional alignment is supported by a modality-alignment loss using the Wasserstein distance to enhance semantic coherence and a content-alignment loss to ensure information integrity across transformations. Extensive experiments on 11 diverse image recognition benchmarks demonstrate that BMSPT can significantly improve adaptability and accuracy and achieve an adaptive, multi-modal learning systems.



