2608004833
  • Open Access
  • Article

Bidirectional Multi-Modal Spatial-Consistent Prompt Tuning for Enhanced Vision-Language Alignment

  • Sikai Bai †,   
  • Haoxi Li †,   
  • Jie Zhang 1,*

Received: 07 May 2026 | Revised: 30 Jul 2026 | Accepted: 06 Aug 2026 | Published: 24 Aug 2026

Abstract

Multi-modal prompt tuning in vision-language models like CLIP has significantly enhanced machine comprehension and interaction with multi-modal content. However, prior works have primarily focused on unidirectional semantic alignment across modalities, limiting their effectiveness in dynamic and open-world visual concept comprehension due to loosely constrained mappings. To address this limitation, we introduce a novel Bidirectional Multi-modal Spatial-consistent Prompt Tuning (BMSPT) framework to ensure a dynamic and fine-grained alignment between text and image modalities. Specifically, our proposed framework facilitates bidirectional interactions between modalities through two novel Vision-Language (V-L) and Language-Vision (L-V) transfer functions. This bidirectional alignment is supported by a modality-alignment loss using the Wasserstein distance to enhance semantic coherence and a content-alignment loss to ensure information integrity across transformations. Extensive experiments on 11 diverse image recognition benchmarks demonstrate that BMSPT can significantly improve adaptability and accuracy and achieve an adaptive, multi-modal learning systems.

References 

  • 1.

    Radford, A.; Kim, J.W.; Hallacy, C.; et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning; PMLR: New York, NY, USA, 2021, pp. 8748–8763.

  • 2.

    Liu, P.; Yuan, W.; Fu, J.; et al. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Comput. Surv. 2023, 55, 1–35.

  • 3.

    Wang, Z.; Zhang, Z.; Lee, C.-Y.; et al. Learning to prompt for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 139–149.

  • 4.

    Brown, T.; Mann, B.; Ryder, N.; et al. Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 2020, 33, 1877–1901.

  • 5.

    Jiang, Z.; Xu, F.F.; Araki, J.; et al. How can we know what language models know? Trans. Assoc. Comput. Linguist. 2020, 8, 423–438.

  • 6.

    Bai, S.; Li, H.; Zhang, J.; et al. Diep: Adaptive mixture-of-experts compression through differentiable expert pruning. Adv. Neural Inf. Process. Syst. 2026, 38, 56090–56115.

  • 7.

    Bai, S.; Li, H.; Zhang, J.; et al. Ttvs: Boosting self-exploring reinforcement learning via test-time variational synthesis. In Findings of the Association for Computational Linguistics; ACL: San Diego, CA, USA, 2026; pp. 32651–32665.

  • 8.

    Xiong, S.; Zhao, Y.; Zhang, J.; et al. Dual prompt tuning based contrastive learning for hierarchical text classification. In Findings of the Association for Computational Linguistics; ACL: Bangkok, Thailand, 2024; pp. 12146–12158.

  • 9.

    Zhou, K.; Yang, J.; Loy, C.C.; et al. Learning to prompt for vision-language models. Int. J. Comput. Vis. 2022, 130, 2337–2348.

  • 10.

    Zhou, K.; Yang, J.; Loy, C.C.; et al. Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 16816–16825.

  • 11.

    Long, Y.; Han, J.; Huang, R.; et al. Fine-grained visual-text prompt-driven self-training for open-vocabulary object detection. IEEE Trans. Neural Netw. Learn. Syst. 2023, 35, 16277–16287.

  • 12.

    Zhao, C.; Wang, Y.; Jiang, X.; et al. Learning domain invariant prompt for vision-language models. IEEE Trans. Image Process. 2024, 33, 1348–1360.

  • 13.

    Jia, M.; Tang, L.; Chen, B.-C.; et al. Visual prompt tuning. In European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 709–727.

  • 14.

    Bai, S.; Li, S.; Zhuang, W.; et al. Combating data imbalances in federated semi-supervised learning with dual regulators. Proc. AAAI Conf. Artif. Intell. 2024, 38, 10989–10997.

  • 15.

    Yang, J.; Li, Z.; Zheng, F.; et al. Prompting for multi-modal tracking. In Proceedings of the 30th ACM International Conference on Multimedia, Lisbon, Portugal, 10–14 October 2022; pp. 3492–3500.

  • 16.

    Tsimpoukelli, M.; Menick, J.L.; Cabi, S.; et al. Multimodal few-shot learning with frozen language models. Adv. Neural Inf. Process. Syst. 2021, 34, 200–212.

  • 17.

    Yao, H., Zhang, R. and Xu, C., Visual-language prompt tuning with knowledge-guided context optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 6757–6767.

  • 18.

    Wang, D.; Li, M.; Liu, X.; et al. Tuning multi-mode token-level prompt alignment across modalities. Adv. Neural Inf. Process. Syst. 2023, 36, 52792–52810.

  • 19.

    Yao, H.; Zhang, R.; Xu, C. Tcp: Textual-based class-aware prompt tuning for visual-language model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 23438–23448.

  • 20.

    Huang, M.; Yang, C.; Yu, X. Cutcp: Custom text generation-based class-aware prompt tuning for visual-language models. Sci. Rep. 2025, 15, 2681.

  • 21.

    Khattak, M.U.; Rasheed, H.; Maaz, M.; et al. Maple: Multi-modal prompt learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–19 June 2023.

  • 22.

    Khattak, M.U.; Wasim, S.T.; Naseer, M.; et al. Self-regulating prompts: Foundational model adaptation without forgetting. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 15190–15200.

  • 23.

    Cai, R.; Yang, G.; Averbuch-Elor, H.; et al. Learning gradient fields for shape generation. In Proceedings of the Computer Vision—ECCV 2020: 16th European Conference, Glasgow, UK, 23–28 August 2020; Proceedings, Part III 16; Springer: Berlin/Heidelberg, Germany, 2020; pp. 364–381.

  • 24.

    Gao, Y.; Shi, X.; Zhu, Y.; et al. Visual prompt tuning for test-time domain adaptation. arXiv 2022, arXiv:2210.04831.

  • 25.

    Sohn, K.; Chang, H.; Lezama, J.; et al. Visual prompt tuning for generative transfer learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 19840–19851.

  • 26.

    Cao, G.; Shi, K.; Fu, H.; et al. Aple: Token-wise adaptive for multi-modal prompt learning. arXiv 2024, arXiv:2401.06827.

  • 27.

    Bai, S.; Zhang, J.; Guo, S.; et al. Diprompt: Disentangled prompt tuning for multiple latent domain generalization in federated learning. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 27274–27283.

  • 28.

    Xin, Y.; Du, J.; Wang, Q.; et al. Mmap: Multi-modal alignment prompt for cross-domain multi-task learning. Proc. AAAI Conf. Artif. Intell. 2024, 38, 16076–16084.

  • 29.

    Deng, J.; Dong, W.; Socher, R.; et al. Imagenet: A large-scale hierarchical image database. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, 20–25 June 2009; pp. 248–255.

  • 30.

    Fei-Fei, L.; Fergus, R.; Perona, P. Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In Proceedings of the Computer Vision and Pattern Recognition Workshop, Washington, DC, USA, 27 June–2 July 2004.

  • 31.

    Parkhi, O.M.; Vedaldi, A.; Zisserman, A.; et al. Cats and dogs. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA, 16–21 June 2012; pp. 3498–3505.

  • 32.

    Krause, J.; Stark, M.; Deng, J.; et al. 3d object representations for fine-grained categorization. In Proceedings of the IEEE International Conference on Computer Vision Workshops, Sydney, Australia, 2–8 December 2013; pp. 554–561.

  • 33.

    Nilsback, M.-E.; Zisserman, A. Automated flower classification over a large number of classes. In Proceedings of the 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing. Bhubaneswar, India, 16–19 December 2008; pp. 722–729.

  • 34.

    Bossard, L.; Guillaumin, M.; Van Gool, L. Food-101—Mining discriminative components with random forests. In Proceedings of the Computer Vision—ECCV 2014: 13th European Conference, Zurich, Switzerland, 6–12 September 2014; Proceedings, Part VI 13. Springer: Berlin/Heidelberg, Germany, 2014; pp. 446–461.

  • 35.

    Maji, S.; Rahtu, E.; Kannala, J.; Blaschko, M.; Vedaldi, A. Fine-grained visual classification of aircraft. arXiv 2013, arXiv:1306.5151.

  • 36.

    Helber, P.; Bischke, B.; Dengel, A.; et al. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2019, 12, 2217–2226.

  • 37.

    Soomro, K.; Zamir, A.R.; Shah, M. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv 2012, arXiv:1212.0402.

  • 38.

    Cimpoi, M.; Maji, S.; Kokkinos, I.; et al. Describing textures in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 3606–3613.

  • 39.

    Xiao, J.; Hays, J.; Ehinger, K.A.; et al. Sun database: Large-scale scene recognition from abbey to zoo. In Proceedings of the 2010 IEEE computer society conference on computer vision and pattern recognition, San Francisco, CA, USA, 13–18 June 2010; pp. 3485–3492.

  • 40.

    Recht, B.; Roelofs, R.; Schmidt, L.; et al. Do imagenet classifiers generalize to imagenet? In Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA, 9–15 June 2019; pp. 5389–5400.

  • 41.

    Wang, H.; Ge, S.; Lipton, Z.; et al. Learning robust global representations by penalizing local predictive power. Adv. Neural Inf. Process. Syst. 2019, 32, 10506–10518.

  • 42.

    Hendrycks, D.; Zhao, K.; Basart, S.; et al. Natural adversarial examples. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 15262–15271.

  • 43.

    Hendrycks, D.; Basart, S.; Mu, N.; et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021; pp. 8340–8349.

  • 44.

    Ju, C.; Han, T.; Zheng, K.; et al. Prompting visual-language models for efficient video understanding. In European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 105–124.

  • 45.

    Kuehne, H.; Jhuang, H.; Garrote, E.; et al. Hmdb: A large video database for human motion recognition. In Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain, 6–13 November 2011; pp. 2556–2563.

  • 46.

    Jeong, J.; Zou, Y.; Kim, T.; et al. Winclip: Zero-/few-shot anomaly classification and segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 19606–19616.

  • 47.

    Zou, Y.; Jeong, J.; Pemula, L.; et al. Spot-the-difference self-supervised pre-training for anomaly detection and segmentation. In European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 392–408.

  • 48.

    Sun, J.; Yin, H.; Tian, Y.; et al. Two-level multimodal fusion for sentiment analysis in public security. Secur. Commun. Netw. 2021, 2021, 6662337.

  • 49.

    Bezirganyan, G.; Sellami, S.; Berti-Equille, L.; et al. Multimodal learning with uncertainty quantification based on discounted belief fusion. arXiv 2024, arXiv:2412.18024.

  • 50.

    Balderas-Díaz, S.; Guerrero-Contreras, G.; Bueno-Crespo, A.; et al. Fuzzy logic-driven confidence aggregation for multimodal sentiment classification. Multimed. Tools Appl. 2026, 85, 289.

  • 51.

    Feng, X.; Lin, Y.; He, L.; et al. Knowledge-guided dynamic modality attention fusion framework for multimodal sentiment analysis. In Findings of the Association for Computational Linguistics; EMNLP: Miami, FL, USA, 2024; pp. 14755–14766.

  • 52.

    Majumder, N.; Hazarika, D.; Gelbukh, A.; et al. Multimodal sentiment analysis using hierarchical fusion with context modeling. Knowl. -Based Syst. 2018, 161, 124–133.

Share this article:
How to Cite
Bai, S.; Li, H.; Zhang, J. Bidirectional Multi-Modal Spatial-Consistent Prompt Tuning for Enhanced Vision-Language Alignment. Edge Intelligence and Systems 2026, 1 (1), 6.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.