2606004315
  • Open Access
  • Article

Edge-Efficient Compositional Recognition via Disentangled Prompt Tuning of Frozen Vision-Language Models

  • Xiaocheng Lu 1,   
  • Chuan He 2,   
  • Ziming Liu 2,   
  • Fushuo Huo 2,   
  • Sikai Bai 1,   
  • Tao Han 1,   
  • Jie Zhang 1,*

Received: 11 May 2026 | Revised: 29 May 2026 | Accepted: 18 Jun 2026 | Published: 17 Jul 2026

Abstract

Deploying vision-language models (VLMs) for fine-grained compositional recognition on edge devices presents a fundamental tension: full fine-tuning is prohibitive for resource-constrained nodes, while frozen backbones with generic prompts fail on the long-tailed state-object compositions encountered in real-world edge scenarios. Soft prompt tuning offers a deployable middle ground because it leaves the backbone frozen and updates only a compact set of prompt parameters per task. However, existing prompt tuning methods for compositional zero-shot learning (CZSL) jointly optimize state and object embeddings, causing prompts to drift toward locally suboptimal regions under uneven entanglement distributions. We propose Progressively Disentangled and Recurrent Prompt Tuning (PDRPT), an edge-efficient framework that decouples object and state updates before joint refinement, suppresses traction force from highly-entangled prompts, and preserves alignment with the natural language space of CLIP. Experiments show consistent improvements over prior state-of-the-art methods while maintaining compact task adapters, providing a deployable path for VLM-based compositional perception on consumer-grade edge hardware. In this revision, we explicitly separate the experimental settings: accuracy results are obtained with CLIP ViT-L/14, while measured edge profiling is conducted with a deployable CLIP ViT-B/32 configuration on an Apple M1 (Apple Inc., Cupertino, CA, USA); ViT-L/14 edge numbers are reported as analytical FLOPs-scaled estimates rather than direct on-device measurements.

References 

  • 1.

    Deng, S.; Zhao, H.; Fang, W.; et al. Edge Intelligence: The Confluence of Edge Computing and Artificial Intelligence. IEEE Internet Things J. 2020, 7, 7457–7469.

  • 2.

    Shi, W.; Cao, J.; Zhang, Q.; et al. Edge Computing: Vision and Challenges. IEEE Internet Things J. 2016, 3, 637–646.

  • 3.

    Radford, A.; Kim, J.W.; Hallacy, C.; et al. Learning Transferable Visual Models from Natural Language Supervision. Computer Vision and Pattern Recognition. arXiv 2021, arXiv:2103.00020.

  • 4.

    Nayak, N.V.; Yu, P.; Bach, S.H. Learning to Compose Soft Prompts for Compositional Zero-Shot Learning. arXiv 2022, arXiv:2204.03574.

  • 5.

    Lu, X.; Guo, S.; Liu, Z.; et al. Decomposed Soft Prompt Guided Fusion Enhancing for Compositional Zero-Shot Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, Canada, 18–22 June 2023; pp. 23560–23569.

  • 6.

    Li, Y.L.; Xu, Y.; Mao, X.; et al. Symmetry and Group in Attribute-Object Compositions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020.

  • 7.

    Misra, I.; Gupta, A.; Hebert, M. From Red Wine to Red Tomato: Composition with Context. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA 21–26 July 2017.

  • 8.

    Mancini, M.; Naeem, M.F.; Xian, Y.; et al. Learning Graph Embeddings for Open World Compositional Zero-Shot Learning. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 1545–1560.

  • 9.

    Karthik, S.; Mancini, M.; Akata, Z. KG-SP: Knowledge Guided Simple Primitives for Open World Compositional Zero-Shot Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA 19–24 June 2022; pp. 9336–9345.

  • 10.

    Li, X.; Yang, X.; Wei, K.; et al. Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA 19–24 June 2022; pp. 9326–9335.

  • 11.

    Yang, Y.; Pan, R.; Li, X.; et al. Dual-Stream Contrastive Learning for Compositional Zero-Shot Recognition. IEEE Trans. Multimed. 2023, 26, 1909–1919.

  • 12.

    Atzmon, Y.; Kreuk, F.; Shalit, U.; et al. A Causal View of Compositional Zero-Shot Recognition. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Online, 6–12 December 2020; pp. 1462–1473.

  • 13.

    Purushwalkam, S.; Nickel, M.; Gupta, A.; et al. Task-Driven Modular Networks for Zero-Shot Compositional Learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea 27 October–2 November 2019; pp. 3593–3602.

  • 14.

    Wang, X.; Yu, F.; Darrell, T.; et al. Task-Aware Feature Generation for Zero-Shot Compositional Learning. Computer Vision and Pattern Recognition. arXiv 2019, arXiv:1906.04854.

  • 15.

    Nagarajan, T.; Grauman, K. Attributes as Operators: Factorizing Unseen Attribute-Object Compositions. Computer Vision and Pattern Recognition. arXiv 2018, arXiv:1803.09851.

  • 16.

    Nan, Z.; Liu, Y.; Zheng, N.; et al. Recognizing Unseen Attribute-Object Pair with Generative Model. Proc. AAAI Conf. Artif. Intell. 2019, 33, 8811–8818.

  • 17.

    Yang, M.; Deng, C.; Yan, J.; et al. Learning Unseen Concepts via Hierarchical Decomposition and Composition. In Proceedings of the Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020.

  • 18.

    Biederman, I. Recognition-by-Components: A Theory of Human Image Understanding. Psychol. Rev. 1987, 94, 115.

  • 19.

    Hoffman, D.D.; Richards, W.A. Parts of Recognition. Cognition 1984, 18, 65–96.

  • 20.

    Bach, S.H.; Sanh, V.; Yong, Z.-X.; et al. Promptsource: An Integrated Development Environment and Repository for Natural Language Prompts. arXiv 2022, arXiv:2202.01279.

  • 21.

    Bommasani, R.; Hudson, D.A.; Adeli, E.; et al. On the Opportunities and Risks of Foundation Models. arXiv 2021, arXiv:2108.07258.

  • 22.

    Brown, T.; Mann, B.; Ryder, N.; et al. Language Models Are Few-Shot Learners. Adv. Neural Inf. Process. Syst. 2020, 33, 1877–1901.

  • 23.

    Sanh, V.; Webson, A.; Raffel, C.; et al. Multitask Prompted Training Enables Zero-Shot Task Generalization. arXiv 2021, arXiv:2110.08207.

  • 24.

    Vu, T.; Lester, B.; Constant, N.; et al. Spot: Better Frozen Model Adaptation Through Soft Prompt Transfer. arXiv 2021, arXiv:2110.07904.

  • 25.

    Zhou, K.; Yang, J.; Loy, C.C.; et al. Learning to Prompt for Vision-Language Models. Int. J. Comput. Vis. 2022, 130, 2337–2348.

  • 26.

    Qin, G.; Eisner, J. Learning How to Ask: Querying LMs with Mixtures of Soft Prompts. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), Online, 6–11 June 2021; pp. 5203–5212.

  • 27.

    Hu, E.J.; Shen, Y.; Wallis, P.; et al. LoRA: Low-Rank Adaptation of Large Language Models. In International conference on learning representations (ICLR). arXiv 2022, arXiv:2106.09685.

  • 28.

    Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; et al. Parameter-Efficient Transfer Learning for NLP. In Proceedings of the 36th International Conference on Machine Learning (ICML), Long Beach, CA, USA, 9–15 June 2019; pp. 2790–2799.

  • 29.

    Li, X.L.; Liang, P. Prefix-Tuning: Optimizing Continuous Prompts for Generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL), Online 1–6 August 2021; pp. 4582–4597.

  • 30.

    Li, E.; Zeng, L.; Zhou, Z.; et al. On-Demand Accelerating Deep Neural Network Inference via Edge Computing. IEEE Trans. Wirel. Commun. 2019, 19, 447–457.

  • 31.

    Vasu, P.K.A.; Pouransari, H.; Faghri, F.; et al. MobileCLIP: Fast Image-Text Models Through Multi-Modal Reinforced Training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 15963–15974.

  • 32.

    Wu, K.; Peng, H.; Zhou, Z.; et al. TinyCLIP: CLIP Distillation via Affinity Mimicking and Weight Inheritance. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 21970–21980.

  • 33.

    Guo, T.; Guo, S.; Wang, J.; et al. PromptFL: Let Federated Participants Cooperatively Learn Prompts Instead of Models—Federated Learning in Age of Foundation Model. IEEE Trans. Mob. Comput. 2023, 23, 1–14.

  • 34.

    Xian, Y.; Lampert, C.H.; Schiele, B.; et al. Zero-Shot Learning—A Comprehensive Evaluation of the Good, the Bad and the Ugly. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 41, 2251–2265.

  • 35.

    Naeem, M.F.; Xian, Y.; Tombari, F.; et al. Learning Graph Embeddings for Compositional Zero-Shot Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 953–962.

  • 36.

    Isola, P.; Lim, J.J.; Adelson, E.H. Discovering States and Transformations in Image Collections. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2015; pp. 1383–1391.

  • 37.

    Yu, A.; Grauman, K. Fine-Grained Visual Comparisons with Local Learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 192–199.

  • 38.

    Mancini, M.; Naeem, M.F.; Xian, Y.; et al. Open World Compositional Zero-Shot Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; Nashville, TN, USA, 20–25 June 2021; pp. 5222–5230.

  • 39.

    Paszke, A.; Gross, S.; Massa, F.; et al. Pytorch: An Imperative Style, High-Performance Deep Learning Library. Adv. Neural Inf. Process. Syst. 2019, 32, 8026–8037.

  • 40.

    Nagarajan, T.; Grauman, K. Attributes as Operators: Factorizing Unseen Attribute-Object Compositions. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 169–185.

  • 41.

    Zhang, T.; Liang, K.; Du, R.; et al. Learning Invariant Visual Representations for Compositional Zero-Shot Learning. In Proceedings of the Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, 23–27 October 2022; proceedings, part XXIV.; Springer: Berlin/Heidelberg, Germany, 2022; pp. 339–355.

  • 42.

    Reddi, V.J.; Kanter, D.; Mattson, P.; et al. MLPerf Mobile Inference Benchmark: An Industry-Standard Open-Source Machine Learning Benchmark for on-Device AI. In Proceedings of the Machine Learning and Systems (MLSys), Santa Clara, CA, USA, 29 August–1 September 2022.

Share this article:
How to Cite
Lu, X.; He, C.; Liu, Z.; Huo, F.; Bai, S.; Han, T.; Zhang, J. Edge-Efficient Compositional Recognition via Disentangled Prompt Tuning of Frozen Vision-Language Models. Edge Intelligence and Systems 2026, 1 (1), 3.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.