2606004289
  • Open Access
  • Review

On the Element-Wise Representation and Reasoning in Zero-Shot Image Recognition: A Systematic Survey

  • Jingcai Guo 1,2,*,   
  • Zhijie Rao 1,   
  • Zhi Chen 3,   
  • Jingren Zhou 4,   
  • Dacheng Tao 5,   
  • Song Guo 6

Received: 04 Apr 2026 | Revised: 14 Jun 2026 | Accepted: 17 Jun 2026 | Published: 20 Jul 2026

Abstract

Zero-shot image recognition (ZSIR) aims to recognize and reason in unseen domains by learning generalized knowledge from limited data in the seen domain. In resource-constrained scenarios such as edge intelligence, zero-shot learning can significantly reduce data requirements, thereby alleviating the burden of data annotation, storage and communication. The gist of ZSIR is constructing a well-aligned mapping between the input visual space and the target semantic space, which is a bottom-up paradigm inspired by the process by which humans observe the world. In recent years, ZSIR has witnessed significant progress on a broad spectrum, from theory to algorithm design, as well as widespread applications. However, to the best of our knowledge, there remains a lack of a systematic review of ZSIR from an element-wise perspective, i.e., learning fine-grained elements of data and their inferential associations. As the element-wise paradigm shows great potential for adapting to dynamic edge environments with limited memory and low latency, this paper thoroughly investigates recent advances in element-wise ZSIR and provides a sound basis for its future development. Concretely, we first integrate three basic ZSIR tasks, i.e., object recognition, compositional recognition, and foundation model-based open-world recognition, into a unified element-wise paradigm and provide a detailed taxonomy and analysis of the main approaches. Next, we summarize the benchmarks, covering technical implementations, standardized datasets, and some more details as a library. Last, we sketch out related applications, discuss vital challenges, and suggest potential future directions.

References 

  • 1.

    Deng, J.; Dong, W.; Socher, R.; et al. ImageNet: A large-scale hierarchical image database. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, 20–25 June 2009; pp. 248–255.

  • 2.

    Chen, C.-Y.; Grauman, K. Inferring analogous attributes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 200–207.

  • 3.

    Misra, I.; Gupta, A.; Hebert, M. From red wine to red tomato: Composition with context. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 1792–1801.

  • 4.

    Yang, Z.; Luo, T.; Wang, D.; et al. Learning to navigate for fine-grained classification. In Proceedings of the 15th European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 420–435.

  • 5.

    Akata, Z.; Reed, S.; Walter, D.; et al. Evaluation of output embeddings for fine-grained image classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; pp. 2927–2936.

  • 6.

    Nagarajan, T.; Grauman, K. Attributes as operators: factorizing unseen attribute-object compositions. In Proceedings of the 15th European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 169–185.

  • 7.

    Wah, C.; Branson, S.; Welinder, P.; et al. The Caltech-Ucsd Birds-200-2011 Dataset; California Institute of Technology: Pasadena, CA, USA, 2011.

  • 8.

    Patterson, G.; Hays, J. Sun attribute database: Discovering, annotating, and recognizing scene attributes. In 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA, 16–21 June 2012; pp. 2751–2758.

  • 9.

    Wei, X.-S.; Song, Y.-Z.; Aodha, O.M.; et al. Fine-grained image analysis with deep learning: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 44, 8927–8948.

  • 10.

    Wang, J.; Cheng, Y.; Feris, R.S. Walk and learn: Facial attribute representation learning from egocentric video and contextual data. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 2295–2304.

  • 11.

    Kovashka, A.; Parikh, D.; Grauman, K. Whittlesearch: Image search with relative attribute feedback. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA, 16–21 June 2012; pp. 2973–2980.

  • 12.

    Wu, Q.; Shen, C.; Wang, P.; et al. Image captioning and visual question answering based on attributes and external knowledge. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 40, 1367–1381.

  • 13.

    Cruz, R.S.; Fernando, B.; Cherian, A.; et al. Neural algebra of classifiers. In Proceedings of the 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Tahoe, NV, USA, 12–15 March 2018; pp. 729–737.

  • 14.

    Biederman, I. Recognition-by-components: a theory of human image understanding. Psychol. Rev. 1987, 94, 115.

  • 15.

    Hoffman, D.D.; Richards, W.A. Parts of recognition. In Readings in Computer Vision; Elsevier: Amsterdam, The Netherlands, 1987; pp. 227–242.

  • 16.

    Felzenszwalb, P.; McAllester, D.; Ramanan, D. A discriminatively trained, multiscale, deformable part model. In Proceedings of the 2008 IEEE Conference on Computer Vision and Pattern Recognition, Anchorage, AK, USA, 23–28 June 2008; pp. 1–8.

  • 17.

    Battaglia, P.W.; Hamrick, J.B.; Bapst, V.; et al. Relational inductive biases, deep learning, and graph networks. arXiv 2018, arXiv:1806.01261.

  • 18.

    Lampert, C.H.; Nickisch, H.; Harmeling, S. Learning to detect unseen object classes by between-class attribute transfer. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, 20–25 June 2009; pp. 951–958.

  • 19.

    Wang, Y.; Yao, Q.; Kwok, J.T.; et al. Generalizing from a few examples: A survey on few-shot learning. ACM Comput. Surv. 2020, 53, 63.

  • 20.

    O’Mahony, N.; Campbell, S.; Carvalho, A.; et al. One-shot learning for custom identification tasks; a review. Procedia Manuf. 2019, 38, 186–193.

  • 21.

    Yang, J.; Zhou, K.; Li, Y.; et al. Generalized out-of-distribution detection: A survey. Int. J. Comput. Vis. 2024, 132, 5635–5662.

  • 22.

    Xian, Y.; Lampert, C.H.; Schiele, B.; et al. Zero-shot learning—A comprehensive evaluation of the good, the bad and the ugly. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 41, 2251–2265.

  • 23.

    Wang, W.; Zheng, V.W.; Yu, H.; et al. A survey of zero-shot learning: Settings, methods, and applications. ACM Trans. Intell. Syst. Technol. 2019, 10, 13.

  • 24.

    Pourpanah, F.; Abdar, M.; Luo, Y.; et al. A review of generalized zero-shot learning methods. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 4051–4070.

  • 25.

    Lake, B.M.; Ullman, T.D.; Tenenbaum, J.B.; et al. Building machines that learn and think like people. Behav. Brain Sci. 2017, 40, e253.

  • 26.

    Purushwalkam, S.; Nickel, M.; Gupta, A.; et al. Task-driven modular networks for zero-shot compositional learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 3593–3602.

  • 27.

    Mancini, M.; Naeem, M.F.; Xian, Y.; et al. Open world compositional zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 5222–5230.

  • 28.

    Radford, A.; Kim, J.W.; Hallacy, C.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, Virtual, 18–24 July 2021; pp. 8748–8763.

  • 29.

    Jia, C.; Yang, Y.; Xia, Y.; et al. Scaling up visual and vision-language representation learning with noisy text supervision. In Proceedings of the International Conference on Machine Learning, Virtual, 18–24 July 2021; pp. 4904–4916.

  • 30.

    Zhou, K.; Yang, J.; Loy, C.C.; et al. Learning to prompt for vision-language models. Int. J. Comput. Vis. 2022, 130, 2337–2348.

  • 31.

    Zhou, K.; Yang, J.; Loy, C.C.; et al. Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 16816–16825.

  • 32.

    Menon, S.; Vondrick, C. Visual classification via description from large language models. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual, 25–29 April 2022.

  • 33.

    Guo, Z.; Zhang, R.; Qiu, L.; et al. CALIP: Zero-shot enhancement of clip with parameter-free attention. In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023; Volume 37, pp. 746–754.

  • 34.

    Zhu, Y.; Xie, J.; Tang, Z.; et al. Semantic-guided multi-attention localization for zero-shot learning. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 8–14 December 2019; Volume 32.

  • 35.

    Xie, G.-S.; Liu, L.; Jin, X.; et al. Attentive region embedding network for zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 9384–9393.

  • 36.

    Xie, G.-S.; Liu, L.; Zhu, F.; et al. Region graph embedding network for zero-shot learning. In Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK, 23–28 August 2020; pp. 562–580.

  • 37.

    Guo, J.; Guo, S.; Zhou, Q.; et al. Graph knows unknowns: Reformulate zero-shot learning as sample-level graph recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023; Volume 37, pp. 7775–7783.

  • 38.

    Huynh, D.; Elhamifar, E. Fine-grained generalized zero-shot learning via dense attribute-based attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 4483–4493.

  • 39.

    Chen, S.; Hong, Z.; Liu, Y.; et al. TRANSZERO: Attribute-guided transformer for zero-shot learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 22 February–1 March 2022; Volume 36, pp. 330–338.

  • 40.

    Liu, Z.; Luo, P.; Qiu, S.; et al. DeepFashion: Powering robust clothes recognition and retrieval with rich annotations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 1096–1104.

  • 41.

    Atzmon, Y.; Kreuk, F.; Shalit, U.; et al. A causal view of compositional zero-shot recognition. In Proceedings of the Advances in Neural Information Processing Systems, Virtual, 6–12 December 2020; Volume 33, pp. 1462–1473.

  • 42.

    Fu, Y.; Hospedales, T.M.; Xiang, T.; et al. Transductive multi-view zero-shot learning. IEEE Trans. Pattern Anal. Mach. Intell. 2015, 37, 2332–2345.

  • 43.

    Jiang, H.; Wang, R.; Shan, S.; et al. Transferable contrastive network for generalized zero-shot learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 9765–9774.

  • 44.

    Zhou, K.; Liu, Z.; Qiao, Y.; et al. Domain generalization: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 4396–4415.

  • 45.

    Ganin, Y.; Lempitsky, V. Unsupervised domain adaptation by backpropagation. In Proceedings of the International Conference on Machine Learning, Lille, France, 6–11 July 2015; pp. 1180–1189.

  • 46.

    Rao, Z.; Guo, J.; Tang, L.; et al. SRCD: Semantic reasoning with compound domains for single-domain generalized object detection. IEEE Trans. Neural Netw. Learn. Syst. 2025, 36, 12497–12506.

  • 47.

    Ge, J.; Xie, H.; Min, S.; et al. Dual part discovery network for zero-shot learning. In Proceedings of the 30th ACM International Conference on Multimedia, Lisboa, Portugal, 10–14 October 2022; pp. 3244–3252.

  • 48.

    Wei, K.; Yang, M.; Wang, H.; et al. Adversarial fine-grained composition learning for unseen attribute-object recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 3741–3749.

  • 49.

    Romera-Paredes, B.; Torr, P. An embarrassingly simple approach to zero-shot learning. In Proceedings of the International Conference on Machine Learning, Lille, France, 6–11 July 2015; pp. 2152–2161.

  • 50.

    Xian, Y.; Lorenz, T.; Schiele, B.; et al. Feature generating networks for zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 5542–5551.

  • 51.

    Verma, V.K.; Arora, G.; Mishra, A.; et al. Generalized zero-shot learning via synthesized examples. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 4281–4289.

  • 52.

    Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; et al. Generative adversarial networks. Commun. ACM 2020, 63, 139–144.

  • 53.

    Chen, Z.; Li, J.; Luo, Y.; et al. CANZSL: Cycle-consistent adversarial networks for zero-shot learning from natural language. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Snowmass Village, CO, USA, 1–5 March 2020; pp. 874–883.

  • 54.

    Chen, Z.; Wang, S.; Li, J.; et al. Rethinking generative zero-shot learning: An ensemble learning perspective for recognising visual patches. In Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA, 12–16 October 2020; pp. 3413–3421.

  • 55.

    Kingma, D.P.; Welling, M. Auto-encoding variational bayes. In Proceedings of the International Conference on Learning Representations (ICLR), Banff, AB, Canada, 14–16 April 2014.

  • 56.

    Chen, Z.; Luo, Y.; Qiu, R.; et al. Semantics disentangling for generalized zero-shot learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021; pp. 8712–8720.

  • 57.

    Chen, Z.; Luo, Y.; Wang, S.; et al. GSMFlow: Generation shifts mitigating flow for generalized zero-shot learning. IEEE Trans. Multimed. 2022, 25, 5374–5385.

  • 58.

    Chen, Z.; Luo, Y.; Wang, S.; et al. Mitigating generation shifts for generalized zero-shot learning. In Proceedings of the 29th ACM International Conference on Multimedia, Virtual, 20–24 October 2021; pp. 844–852.

  • 59.

    Liu, S.; Long, M.; Wang, J.; et al. Generalized zero-shot learning with deep calibration network. In Proceedings of the Advances in Neural Information Processing Systems, Montréal, QC, Canada, 3–8 December 2018; Volume 31.

  • 60.

    Wang, Z.; Gou, Y.; Li, J.; et al. Region semantically aligned network for zero-shot learning. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, Virtual, 1–5 November 2021; pp. 2080–2090.

  • 61.

    Ji, Z.; Fu, Y.; Guo, J.; et al. Stacked semantics-guided attention model for fine-grained zero-shot learning. In Proceedings of the Advances in Neural Information Processing Systems, Montréal, QC, Canada, 3–8 December 2018; Volume 31.

  • 62.

    Elhoseiny, M.; Zhu, Y.; Zhang, H.; et al. Link the head to the “beak”: Zero shot learning from noisy text description at part precision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 5640–5649.

  • 63.

    Zhu, Y.; Elhoseiny, M.; Liu, B.; et al. A generative adversarial approach for zero-shot learning from noisy texts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 1004–1013.

  • 64.

    Li, Y.; Zhang, J.; Zhang, J.; et al. Discriminative learning of latent features for zero-shot recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 7463–7471.

  • 65.

    Ge, J.; Xie, H.; Min, S.; et al. Semantic-guided reinforced region embedding for generalized zero-shot learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 2–9 February 2021; Volume 35, pp. 1406–1414.

  • 66.

    Li, Y.; Liu, Z.; Yao, L.; et al. An entropy-guided reinforced partial convolutional network for zero-shot learning. IEEE Trans. Circuits Syst. Video Technol. 2022, 32, 5175–5186.

  • 67.

    Liu, M.; Zhang, C.; Bai, H.; et al. Part-object progressive refinement network for zero-shot learning. IEEE Trans. Image Process. 2024, 33, 2032–2043.

  • 68.

    Chen, S.; Hou, W.; Khan, S.; et al. Progressive semantic-guided vision transformer for zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 23964–23974.

  • 69.

    Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; et al. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria, 3–7 May 2021.

  • 70.

    Xu, W.; Xian, Y.; Wang, J.; et al. VGSE: Visually-grounded semantic embeddings for zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 9316–9325.

  • 71.

    Chen, Z.; Zhao, Z.; Guo, J.; et al. SVIP: Semantically contextualized visual patches for zero-shot learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–20 October 2025.

  • 72.

    Kipf, T.N.; Welling, M. Semi-supervised classification with graph convolutional networks. In Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France, 26–24 April 2017.

  • 73.

    Hu, Z.; Zhao, H.; Peng, J.; et al. Region interaction and attribute embedding for zero-shot learning. Inf. Sci. 2022, 609, 984–995.

  • 74.

    Chen, S.; Hong, Z.; Xie, G.; et al. GNDAN: Graph navigated dual attention network for zero-shot learning. IEEE Trans. Neural Netw. Learn. Syst. 2022, 35, 4516–4529.

  • 75.

    Vaswani, A.; Shazeer, N.; Parmar, N.; et al. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Volume 30.

  • 76.

    Liu, Y.; Zhou, L.; Bai, X.; et al. Goal-oriented gaze estimation for zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 3794–3803.

  • 77.

    Huynh, D.; Elhamifar, E. Compositional zero-shot learning via fine-grained dense feature composition. In Proceedings of the Advances in Neural Information Processing Systems, Virtual, 6–12 December 2020; Volume 33, pp. 19849–19860.

  • 78.

    Liu, Y.; Dang, Y.; Gao, X.; et al. Zero-shot learning with attentive region embedding and enhanced semantics. IEEE Trans. Neural Netw. Learn. Syst. 2022, 35, 4220–4231.

  • 79.

    Naeem, M.F.; Xian, Y.; Gool, L.V.; et al. I2DFORMER: Learning image to document attention for zero-shot image classification. In Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA, 28 November–9 December 2022; Volume 35, pp. 12283–12294.

  • 80.

    Chen, Z.; Huang, Y.; Chen, J.; et al. DUET: Cross-modal semantic grounding for contrastive zero-shot learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023; Volume 37, pp. 405–413.

  • 81.

    Liu, M.; Li, F.; Zhang, C.; et al. Progressive semantic-visual mutual adaption for generalized zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 15337–15346.

  • 82.

    Chen, S.; Hong, Z.; Xie, G.-S.; et al. MSDN: Mutually semantic distillation network for zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 7612–7621.

  • 83.

    Chen, S.; Chen, S.; Xie, G.-S.; et al. Mutually causal semantic distillation network for zero-shot learning, Int. J. Comput. Vis. 2026, 134, 246.

  • 84.

    Jiang, H.; Li, Z.; Yu, X.; et al. Visual and semantic prompt collaboration for generalized zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025; pp. 20275–20285.

  • 85.

    Hou, W.; Fu, D.; Li, K.; et al. Zeromamba: Exploring visual state space model for zero-shot learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA, 25 February–4 March 2025; Volume 39, pp. 3527–3535.

  • 86.

    Rao, Z.; Guo, J. Balancing cross-modal attention for generalized zero-shot learning. In Proceedings of the 33rd ACM International Conference on Multimedia, Dublin, Ireland, 27–31 October 2025; pp. 3360–3369.

  • 87.

    Yang, H.-M.; Zhang, X.-Y.; Yin, F.; et al. Robust classification with convolutional prototype learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 3474–3482.

  • 88.

    Liu, J.; Ren, Y.; Li, W.; et al. Improving out-of-distribution detection with margin-based prototype learning. In Proceedings of the International Conference on Neural Information Processing, Changsha, China, 20–23 November 2023; pp. 149–160.

  • 89.

    Wang, K.; Liew, J.H.; Zou, Y.; et al. PANET: Few-shot image semantic segmentation with prototype alignment. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 9197–9206.

  • 90.

    Xu, W.; Xian, Y.; Wang, J.; et al. Attribute prototype network for zero-shot learning. In Proceedings of the Advances in Neural Information Processing Systems, Virtual, 6–12 December 2020; Volume 33, pp. 21969–21980.

  • 91.

    Jayaraman, D.; Sha, F.; Grauman, K. Decorrelating semantic visual attributes by resisting the urge to share. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 1629– 1636.

  • 92.

    Wang, C.; Min, S.; Chen, X.; et al. Dual progressive prototype network for generalized zero-shot learning. In Proceedings of the Advances in Neural Information Processing Systems, Virtual, 6–14 December 2021; Volume 34, pp. 2936–2948.

  • 93.

    Yue, Q.; Cui, J.; Liang, J.; et al. Class semantic attribute perception guided zero-shot learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA, 25 February–4 March 2025; Volume 39, pp. 22281–22289.

  • 94.

    Guo, T.; Liang, J.; Xie, G.-S. Group-wise interactive region learning for zero-shot recognition. Inf. Sci. 2023, 642, 119135.

  • 95.

    Cheng, D.; Wang, G.; Wang, N.; et al. Discriminative and robust attribute alignment for zero-shot learning. IEEE Trans. Circuits Syst. Video Technol. 2023, 33, 4244–4256.

  • 96.

    Du, Y.; Shi, M.; Wei, F.; et al. Boosting zero-shot learning via contrastive optimization of attribute representations. IEEE Trans. Neural Netw. Learn. Syst. 2023, 35, 16706–16719.

  • 97.

    Chen, Z.; Zhang, P.; Li, J.; et al. Zero-shot learning by harnessing adversarial samples. In Proceedings of the 31st ACM International Conference on Multimedia, Ottawa, ON, Canada, 29 October–3 November 2023; pp. 4138–4146.

  • 98.

    Liu, Y.; Guo, J.; Cai, D.; et al. Attribute attention for semantic disambiguation in zero-shot learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 6698–6707.

  • 99.

    Liu, L.; Zhou, T.; Long, G.; et al. Attribute propagation network for graph zero-shot learning. In Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA, 7–12 February 2020; Volume 34, pp. 4868–4875.

  • 100.

    Akata, Z.; Malinowski, M.; Fritz, M.; et al. Multi-cue zero-shot learning with strong supervision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 59–68.

  • 101.

    Rao, Z.; Guo, J.; Lu, X.; et al. Dual expert distillation network for generalized zero-shot learning. In Proceedings of the International Joint Conference on Artificial Intelligence, Jeju, Republic of Korea, 3–9 August 2024.

  • 102.

    Nan, Z.; Liu, Y.; Zheng, N.; et al. Recognizing unseen attribute-object pair with generative model. In Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA, 27 January–1 February 2019; Volume 33, pp. 8811–8818.

  • 103.

    Yuan, Z.; Wang, Z.; Pan, Y.; et al. I2CD: An invertible causal framework for compositional zero-shot learning via disentangle-compose-disentangle. In Proceedings of the AAAI Conference on Artificial Intelligence, Singapore, 20–27 January 2026; Volume 40, pp. 12295–12303.

  • 104.

    Ruis, F.; Burghouts, G.; Bucur, D. Independent prototype propagation for zero-shot compositionality. In Proceedings of the Advances in Neural Information Processing Systems, Virtual, 6–14 December 2021; Volume 34, pp 10641–10653.

  • 105.

    Hu, X.; Wang, Z. Leveraging sub-class discrimination for compositional zero-shot learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023, Volume 37, pp. 890–898.

  • 106.

    Zheng, Z.; Zhu, H.; Nevatia, R. CAILA: Concept-aware intra-layer adapters for compositional zero-shot learning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA, 3–8 January 2024; pp. 1721–1731.

  • 107.

    Yang, M.; Xu, C.; Wu, A.; et al. A decomposable causal view of compositional zero-shot learning. IEEE Trans. Multimed. 2022, 25, 5892–5902.

  • 108.

    Rao, Z.; Guo, J.; Li, M.; et al. Exploring transferable homogeneous groups for compositional zero-shot learning. In Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence (IJCAI-25), Montreal, QC, Canada, 16–22 August, 2025.

  • 109.

    Saini, N.; Pham, K.; Shrivastava, A. Disentangling visual embeddings for attributes and objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 13658–13667.

  • 110.

    Hao, S.; Han, K.; Wong, K.-Y.K. Learning attention as disentangler for compositional zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 15 315– 15 324.

  • 111.

    Jing, C.; Li, Y.; Chen, H.; et al. Retrieval-augmented primitive representations for compositional zero-shot learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 26–27 February 2024; Volume 38, pp. 2652–2660.

  • 112.

    Zhang, T.; Liang, K.; Du, R.; et al. Learning invariant visual representations for compositional zero-shot learning. In Proceedings of the European Conference on Computer Vision (ECCV), Tel Aviv, Israel, 23–27 October 2022; pp. 339–355.

  • 113.

    Lu, X.; Guo, S.; Liu, Z.; et al. Decomposed soft prompt guided fusion enhancing for compositional zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 23 560–23 569.

  • 114.

    Huang, S.; Gong, B.; Feng, Y.; et al. TROIKA: Multi-path cross-modal traction for compositional zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 24 005–24 014.

  • 115.

    Panda, A.; Mukherjee, D.P. Compositional zero-shot learning using multi-branch graph convolution and cross-layer knowledge sharing. Pattern Recognit. 2024, 145, 109916.

  • 116.

    Kim, H.; Lee, J.; Park, S.; et al. Hierarchical visual primitive experts for compositional zero-shot learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–3 October 2023; pp. 5675–5685.

  • 117.

    Li, Y.-L.; Xu, Y.; Mao, X.; et al. Symmetry and group in attribute- object compositions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 11316–11325.

  • 118.

    Wang, Q.; Liu, L.; Jing, C.; et al. Learning conditional attributes for compositional zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 11197–11206.

  • 119.

    Li, L.; Chen, G.; Wang, Z.; et al. Compositional zero-shot learning via progressive language-based observations. In Proceedings of the 33rd ACM International Conference on Multimedia, Dublin, Ireland, 27–31 October 2025; pp. 3827–3836.

  • 120.

    Liu, Z.; Li, Y.; Yao, L.; et al. Simple primitives with feasibility-and contextuality-dependence for open-world compositional zero-shot learning. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 46, 543–560.

  • 121.

    Huo, F.; Xu, W.; Guo, S.; et al. PROCC: Progressive cross-primitive compatibility for open-world compositional zero-shot learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 26–27 February 2024; Volume 38, pp. 12689–12697.

  • 122.

    Xu, Z.; Wang, G.; Wong, Y.; et al. Relation-aware compositional zero-shot learning for attribute-object pair recognition. IEEE Trans. Multimed. 2021, 24, 3652–3664.

  • 123.

    Khan, M.G.Z.A.; Naeem, M.F.; Van Gool, L.; et al. Learning attention propagation for compositional zero-shot learning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA, 2–7 January 2023; pp. 3828–3837.

  • 124.

    Naeem, M.F.; Xian, Y.; Tombari, F.; et al. Learning graph embeddings for compositional zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 953–962.

  • 125.

    Anwaar, M.U.; Pan, Z.; Kleinsteuber, M. On leveraging variational graph embeddings for open world compositional zero-shot learning. In Proceedings of the 30th ACM International Conference on Multimedia, Lisboa, Portugal, 10–14 October 2022; pp. 4645–4654.

  • 126.

    Xu, G.; Chai, J.; Kordjamshidi, P. GIPCOL: Graph-injected soft prompting for compositional zero-shot learning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA, 3–8 January 2024; pp. 5774–5783.

  • 127.

    Jaiswal, A.; Babu, A.R.; Zadeh, M.Z.; et al. A survey on contrastive self-supervised learning. Technologies 2020, 9, 2.

  • 128.

    Li, X.; Yang, X.; Wei, K.; et al. Siamese contrastive embedding network for compositional zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 9326–9335.

  • 129.

    Yang, Y.; Pan, R.; Li, X.; et al. Dual-stream contrastive learning for compositional zero-shot recognition. IEEE Trans. Multimed. 2023, 26, 1909–1919.

  • 130.

    Wang, X.; Yu, F.; Wang, R.; et al. Tafe-Net: Task-aware feature embeddings for low shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 1831–1840.

  • 131.

    Jiang, C.; Ye, Q.; Wang, S.; et al. Mutual balancing in state-object components for compositional zero-shot learning. Pattern Recognit. 2024, 152, 110451.

  • 132.

    Jiang, C.; Zhang, H. Revealing the proximate long-tail distribution in compositional zero-shot learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 26–27 February 2024; Volume 38, pp. 2498–2506.

  • 133.

    Li, Y.; Liu, Z.; Chen, H.; et al. Context-based and diversity-driven specificity in compositional zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 17037–17046.

  • 134.

    Li, M.; Guo, J.; Da Xu, R.Y.; et al. TSCA: On the semantic consistency alignment via conditional transport for compositional zero-shot learning. In Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence (IJCAI-25), Montreal, QC, Canada, 16–22 August 2025.

  • 135.

    Wu, P.; Lu, X.; Hu, H.; et al. Logiczsl: Exploring logic-induced representation for compositional zero-shot learning. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 30301–30311.

  • 136.

    Karthik, S.; Mancini, M.; Akata, Z. KG-SP: Knowledge guided simple primitives for open world compositional zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 9336–9345.

  • 137.

    Hu, X.; Wang, Z. A dynamic learning method towards realistic compositional zero-shot learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 26–27 February 2024; Volume 38, pp. 2265–2273.

  • 138.

    Zhang, J.; Huang, J.; Jin, S.; et al. Vision-language models for vision tasks: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 5625–5644.

  • 139.

    Cui, Q.; Zhou, B.; Guo, Y.; et al. Contrastive vision-language pre-training with limited resources. In Proceedings of the European Conference on Computer Vision (ECCV), Tel Aviv, Israel, 23–27 October 2022; pp. 236–253.

  • 140.

    Singh, A.; Hu, R.; Goswami, V.; et al. FLAVA: A foundational language and vision alignment model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 15638–15650.

  • 141.

    Huang, R.; Long, Y.; Han, J.; et al. NLIP: Noise-robust language-image pre-training. In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023; Volume 37, pp. 926–934.

  • 142.

    Esmaeilpour, S.; Liu, B.; Robertson, E.; et al. Zero-shot out-of-distribution detection based on the pre-trained model clip. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 22 February–1 March 2022; Volume 36, pp. 6568–6576.

  • 143.

    Wang, H.; Li, Y.; Yao, H.; et al. CLIPN for zero-shot OOD detection: Teaching clip to say no. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–3 October 2023; pp. 1802–1812.

  • 144.

    Brown, T.; Mann, B.; Ryder, N.; et al. Language models are few-shot learners. In Proceedings of the Advances in Neural Information Processing Systems, Virtual, 6–12 December 2020; Volume 33, pp. 1877–1901.

  • 145.

    Pratt, S.; Covert, I.; Liu, R.; et al. What does a platypus look like? generating customized prompts for zero-shot image classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–3 October 2023; pp. 15691–15701.

  • 146.

    Novack, Z.; McAuley, J.; Lipton, Z.C.; et al. CHILS: Zero-shot image classification with hierarchical label sets. In Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA, 23–29 July 2023; pp. 26342–26362.

  • 147.

    Chen, Z. Don’t paint everyone with the same brush: Adaptive prompt prototype learning for vision-language models. In Proceedings of the Twelfth International Conference on Learning Representations, Vienna, Austria, 7–11 May 2024.

  • 148.

    Udandarao, V.; Gupta, A.; Albanie, S. SUS-X: Training-free name-only transfer of vision-language models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–3 October 2023; pp. 2725–2736.

  • 149.

    Rombach, R.; Blattmann, A.; Lorenz, D.; et al. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 10 684–10 695.

  • 150.

    Chen, Y.; Zhang, Q.; Shi, X.; et al. PromptMix: Llm-aided prompt learning for generalizing vision-language models. Inf. Fusion 2026, 131, 104186.

  • 151.

    Shu, M.; Nie, W.; Huang, D.-A.; et al. Test-time prompt tuning for zero-shot generalization in vision-language models. In Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA, 28 November–9 December 2022; Volume 35, pp. 14274–14289.

  • 152.

    Feng, C.-M.; Yu, K.; Liu, Y.; et al. Diverse data augmentation with diffusions for effective test-time prompt tuning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–3 October 2023; pp. 2704–2714.

  • 153.

    Samadh, J.A.; Gani, M.H.; Hussein, N.; et al. Align your prompts: Test- time prompting with distribution alignment for zero-shot generalization. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 10–15 December 2024; Volume 36.

  • 154.

    Ma, X.; Zhang, J.; Guo, S.; et al. Swapprompt: Test-time prompt adaptation for vision-language models. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 10–15 December 2024; Volume 36.

  • 155.

    Liu, Z.; Sun, H.; Peng, Y.; et al. DART: Dual-modal adaptive online prompting and knowledge retention for test-time adaptation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 26–27 February 2024; Volume 38, pp. 14106–14114.

  • 156.

    Zhang, Y.; Zhu, W.; Tang, H.; et al. Dual memory networks: A versatile adaptation approach for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 28718–28728.

  • 157.

    Karmanov, A.; Guan, D.; Lu, S.; et al. Efficient test-time adaptation of vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 14162–14171.

  • 158.

    Zhang, D.-C.; Zhou, Z.; Li, Y.-F. Robust test-time adaptation for zero-shot prompt tuning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 26–27 February 2024; Volume 38, pp. 16714–16722.

  • 159.

    Zanella, M.; Ayed, I.B. On the test-time zero-shot generalization of vision-language models: Do we really need prompt learning? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 23783–23793.

  • 160.

    Liang, Y.; Chen, H.; Xiong, Y.; et al. Advancing reliable test-time adaptation of vision-language models under visual variations. In Proceedings of the 33rd ACM International Conference on Multimedia, Dublin, Ireland, 27–31 October 2025; pp. 4788–4797.

  • 161.

    Zhu, X.; Wang, S.; Zhu, B.; et al. Dynamic multimodal prototype learning in vision-language models. In Proceedings of the IEEE/CVF international conference on computer vision, Honolulu, HI, USA, 19–20 October 2025; pp. 2501–2511.

  • 162.

    Huang, T.; Chu, J.; Wei, F. Unsupervised prompt learning for vision-language models. arXiv 2022, arXiv:2204.03649.

  • 163.

    Li, J.; Savarese, S.; Hoi, S.C. Masked unsupervised self-training for label-free image classification. In Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023.

  • 164.

    Qian, Q.; Xu, Y.; Hu, J. Intra-modal proxy learning for zero-shot visual categorization with clip. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 10–15 December 2024; Volume 36.

  • 165.

    Mirza, M.J.; Karlinsky, L.; Lin, W.; et al. LAFTER: Label-free tuning of zero-shot classifier using language and unlabeled image collections. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 10–15 December 2024; Volume 36.

  • 166.

    Kalantidis, Y.; Tolias, G. Label propagation for zero-shot classification with vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 23209–23218.

  • 167.

    Kahana, J.; Cohen, N.; Hoshen, Y. Improving zero-shot models with label distribution priors. arXiv 2022, arXiv:2212.00784.

  • 168.

    Lampert, C.H.; Nickisch, H.; Harmeling, S. Attribute-based classification for zero-shot visual object categorization. IEEE Trans. Pattern Anal. Mach. Intell. 2013, 36, 453–465.

  • 169.

    Nilsback, M.-E.; Zisserman, A. Automated flower classification over a large number of classes. In Proceedings of the 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing, Bhubaneswar, India, 16–19 December 2008; pp. 722–729.

  • 170.

    Van Horn, G.; Branson, S.; Farrell, R.; et al. Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; pp. 595–604.

  • 171.

    Helber, P.; Bischke, B.; Dengel, A.; et al. EUROSAT: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2019, 12, 2217–2226.

  • 172.

    Krizhevsky, A.; Hinton, G. Learning multiple layers of features from tiny images. 2009. Available online: https://cave.cs.toronto.edu/kriz/learning-features-2009-TR.pdf (accessed on 24 July 2024).

  • 173.

    Bossard, L.; Guillaumin, M.; Van Gool, L. Food-101–mining discriminative components with random forests. In Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland, 6–12 September 2014; pp. 446–461.

  • 174.

    Parkhi, O.M.; Vedaldi, A.; Zisserman, A.; et al. Cats and dogs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Providence, RI, USA, 16–21 June 2012; pp. 3498–3505.

  • 175.

    Krause, J.; Deng, J.; Stark, M.; et al. Collecting a Large-Scale Dataset of Fine-Grained Cars. 2013. Available online: https://ai.stanford.edu/~jkrause/papers/fgvc13.pdf (accessed on 23 July 2024).

  • 176.

    Maji, S.; Rahtu, E.; Kannala, J.; et al. Fine-grained visual classification of aircraft. arXiv 2013, arXiv:1306.5151.

  • 177.

    Cimpoi, M.; Maji, S.; Kokkinos, I.; et al. Describing textures in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 3606–3613.

  • 178.

    Farhadi, A.; Endres, I.; Hoiem, D.; et al. Describing objects by their attributes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Miami, FL, USA, 20–25 June 2009; pp. 1778–1785.

  • 179.

    Yu, A.; Grauman, K. Fine-grained visual comparisons with local learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 192–199.

  • 180.

    Isola, P.; Lim, J.J.; Adelson, E.H. Discovering states and transformations in image collections. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; pp. 1383–1391.

  • 181.

    Fei-Fei, L.; Fergus, R.; Perona, P. Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In Proceedings of the 2004 Conference on Computer Vision and Pattern Recognition Workshop, Washington, DC, USA, 27 June–2 July 2004; p. 178

  • 182.

    Griffin, G.; Holub, A.; Perona, P.; et al. Caltech-256 Object Category Dataset; Technical Report 7694; California Institute of Technology Pasadena: Pasadena CA, USA, 2007.

  • 183.

    Xiao, J.; Hays, J.; Ehinger, K.A.; et al. Sun database: Large-scale scene recognition from abbey to zoo. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), San Francisco, CA, USA, 13–18 June 2010; pp. 3485–3492.

  • 184.

    Berg, T.; Liu, J.; Lee, S.W.; et al. BIRDSNAP: Large-scale fine-grained visual categorization of birds. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 2011–2018.

  • 185.

    He, K.; Zhang, X.; Ren, S.; et al. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778.

  • 186.

    Simonyan, K.; Zisserman, A. Very deep convolutional networks for large-scale image recognition. In Proceedings of the International Conference on Learning Representations (ICLR), San Diego, CA, USA, 7–9 May 2015.

  • 187.

    Szegedy, C.; Liu, W.; Jia, Y.; et al. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; pp. 1–9.

  • 188.

    Krizhevsky, A.; Sutskever, I.; Hinton, G.E. Imagenet classification with deep convolutional neural networks. In Proceedings of the Advances in Neural Information Processing Systems, Lake Tahoe, NV, USA, 3–6 December 2012; Volume 25.

  • 189.

    Chen, X.; Deng, X.; Lan, Y.; et al. Explanatory object part aggregation for zero-shot learning. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 46, 851–868.

  • 190.

    Wang, Z.; Gou, Y.; Li, J.; et al. Language-augmented pixel embedding for generalized zero-shot learning. IEEE Trans. Circuits Syst. Video Technol. 2022, 33, 1019–1030.

  • 191.

    Cheng, D.; Wang, G.; Wang, B.; et al. Hybrid routing transformer for zero-shot learning. Pattern Recognit. 2023, 137, 109270.

  • 192.

    Li, Y.; Liu, Z.; Jha, S.; et al. Distilled reverse attention network for open-world compositional zero-shot learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–3 October 2023; pp. 1782–1791.

  • 193.

    Hou, Y.; Zhang, J.; Lin, Z.; et al. Large language models are zero-shot rankers for recommender systems. In Proceedings of the European Conference on Information Retrieval, Glasgow, UK, 24–28 March 2024; pp. 364–381.

  • 194.

    Kojima, T.; Gu, S.S.; Reid, M.; et al. Large language models are zero-shot reasoners. In Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA, 28 November–9 December 2022; Volume 35, pp. 22199–22213.

  • 195.

    Wei, J.; Bosma, M.; Zhao, V.Y.; et al. Finetuned language models are zero-shot learners. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual, 25–29 April 2022.

  • 196.

    Johnson, M.; Schuster, M.; Le, Q.V.; et al. Google’s multilingual neural machine translation system: Enabling zero-shot translation. Trans. Assoc. Comput. Linguist. 2017, 5, 339–351.

  • 197.

    Baek, J.; Aji, A.F.; Saffari, A. Knowledge-augmented language model prompting for zero-shot knowledge graph question answering. In Proceedings of the 1st Workshop on Natural Language Reasoning and Structured Explanations (NLRSE), Toronto, ON, Canada, 13 July 2023.

  • 198.

    Kumar, P.; Pathania, K.; Raman, B. Zero-shot learning based cross- lingual sentiment analysis for sanskrit text with insufficient labeled data. Appl. Intell. 2023, 53, 10096–10113.

  • 199.

    Chen, Q.; Wang, W.; Huang, K.; et al. Zero-shot text classification via knowledge graph embedding for social media data. IEEE Internet Things J. 2021, 9, 9205–9213.

  • 200.

    Bari, M.S.; Joty, S.; Jwalapuram, P. Zero-resource cross-lingual named entity recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA, 7–12 February 2020; Volume 34, pp. 7415–7423.

  • 201.

    Thakur, N.; Reimers, N.; Rücklé, A.; et al. BEIR: A heterogenous benchmark for zero-shot evaluation of information retrieval models. In Proceedings of the Thirty-Fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, Virtual, 6–14 December 2021.

  • 202.

    Huynh, D.; Elhamifar, E. A shared multi-attention framework for multi-label zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 8776–8786.

  • 203.

    Ren, S.; He, K.; Girshick, R.; et al. Faster r-CNN: Towards real- time object detection with region proposal networks. In Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada, 7–12 December 2015; Volume 28.

  • 204.

    Liu, W.; Anguelov, D.; Erhan, D.; et al. SSD: Single shot multibox detector. In Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands, 11–14 October 2016; pp. 21–37.

  • 205.

    Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; pp. 3431–3440.

  • 206.

    Kirillov, A.; Mintun, E.; Ravi, N.; et al. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–3 October 2023; pp. 4015–4026.

  • 207.

    Bansal, A.; Sikka, K.; Sharma, G.; et al. Zero-shot object detection. In Proceedings of the 15th European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 384–400.

  • 208.

    Rahman, S.; Khan, S.; Porikli, F. Zero-shot object detection: Learning to simultaneously recognize and localize novel concepts. in ACCV. Springer, 2018; pp. 547–563.

  • 209.

    Li, Z.; Yao, L.; Zhang, X.; et al. Zero-shot object detection with textual descriptions. In Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA, 27 January–1 February 2019; Volume 33, pp. 8690–8697.

  • 210.

    Gu, Z.; Zhou, S.; Niu, L.; et al. Context-aware feature generation for zero-shot semantic segmentation. In Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA, 12–16 October 2020; pp. 1921–1929.

  • 211.

    Huynh, D.; Kuen, J.; Lin, Z.; et al. Open-vocabulary instance segmentation via robust cross-modal pseudo-labeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 7020–7031.

  • 212.

    He, S.; Ding, H.; Jiang, W. Primitive generation and semantic-related alignment for universal zero-shot segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 11 238–11 247.

  • 213.

    Guo, Y.; Wang, H.; Hu, Q.; et al. Deep learning for 3D point clouds: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 43, 4338–4364.

  • 214.

    Naeem, M.F.; O(¨)rnek, E.P.; Xian, Y.; et al. 3D compositional zero-shot learning with decompositional consensus. In Proceedings of the European Conference on Computer Vision (ECCV), Tel Aviv, Israel, 23–27 October 2022; pp. 713–730.

  • 215.

    Cheraghian, A.; Rahman, S.; Chowdhury, T.F.; et al. Zero-shot learning on 3D point cloud objects and beyond. Int. J. Comput. Vis. 2022, 130, 2364–2384.

  • 216.

    Zhang, D.; Liang, D.; Yang, H.; et al. SAM3D: Zero-shot 3D object detection via segment anything model. Sci. China Inf. Sci. 2024, 67, 149101.

  • 217.

    Abdelreheem, A.; Skorokhodov, I.; Ovsjanikov, M.; et al. SATR: Zero-shot semantic segmentation of 3D shapes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–3 October 2023; pp. 15166–15179.

  • 218.

    Liu, R.; Wu, R.; Van Hoorick, B.; et al. Zero-1-to-3: Zero-shot one image to 3D object. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–3 October 2023.

  • 219.

    Liu, K.; Zhan, F.; Chen, Y.; et al. Stylerf: Zero-shot 3D style transfer of neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 8338–8348.

  • 220.

    Xu, J.; Wang, X.; Cheng, W.; et al. DREAM3D: Zero-shot text-to-3D synthesis using 3D shape prior and text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 20 908–20 918.

  • 221.

    Li, L.; Dai, A. Genzi: Zero-shot 3D human-scene interaction generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 20465–20474.

  • 222.

    Jiang, Z.; Zhou, Z.; Li, L.; et al. Back to optimization: Diffusion-based zero-shot 3D human pose estimation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA, 3–8 January 2024; pp. 6142–6152.

  • 223.

    Abdelreheem, A.; Eldesokey, A.; Ovsjanikov, M.; et al. Zero-shot 3D shape correspondence. In Proceedings of the SIGGRAPH Asia 2023 Conference Papers, Sydney, NSW, Australia, 12–15 December 2023; pp. 1–11.

  • 224.

    Javed, S.; Mahmood, A.; Ganapathi, I.I.; et al. CPLIP: Zero-shot learning for histopathology with comprehensive vision-language alignment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 11450–11459.

  • 225.

    Li, X.; Wen, C.; Hu, Y.; et al. RS-Clip: Zero shot remote sensing scene classification via contrastive vision-language supervision. Int. J. Appl. Earth Obs. Geoinf. 2023, 124, 103497.

  • 226.

    Li, A.; Qiu, C.; Kloft, M.; et al. Zero-shot anomaly detection via batch normalization. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 10–15 December 2024; Volume 36.

  • 227.

    Chen, S.; Huang, D. Elaborative rehearsal for zero-shot action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021; pp. 13638–13647.

  • 228.

    Song, L.; Shang, X.; Yang, C.; et al. Attribute-guided multiple instance hashing network for cross-modal zero-shot hashing. IEEE Trans. Multimed. 2022, 25, 5305–5318.

  • 229.

    Wang, S.; Chang, J.; Wang, Z.; et al. Content-aware rectified activation for zero-shot fine-grained image retrieval. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 4366–4380.

  • 230.

    Jiang, X.; Xu, X.; Zhou, Z.; et al. Zero-shot video moment retrieval with angular reconstructive text embeddings. IEEE Trans. Multimed. 2024, 26, 9657–9670.

  • 231.

    Wang, Z.; Hu, R.; Liang, C.; et al. Zero-shot person re-identification via cross-view consistency. IEEE Trans. Multimed. 2015, 18, 260–272.

  • 232.

    Hong, M.; Zhang, X.; Li, G.; et al. Fine-grained feature generation for generalized zero-shot video classification. IEEE Trans. Image Process. 2023, 32, 1599–1612.

  • 233.

    Wang, L.; Zhang, X.; Su, H.; et al. A comprehensive survey of continual learning: theory, method and application. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 5362–5383.

  • 234.

    Zhou, F.; Huang, S.; Liu, B.; et al. Multi-label image classification via category prototype compositional learning. IEEE Trans. Circuits Syst. Video Technol. 2021, 32, 4513–4525.

  • 235.

    Guo, J.; Zhou, Q.; Li, R.; et al. Parsnets: A parsimonious orthogonal and low-rank linear networks for zero-shot learning. In Proceedings of the International Joint Conference on Artificial Intelligence, Macao, China, 19–25 August 2023.

  • 236.

    Rahman, S.; Khan, S.; Porikli, F. A unified approach for conventional zero-shot, generalized zero-shot, and few-shot learning. IEEE Trans. Image Process. 2018, 27, 5652–5667.

  • 237.

    Shafiee, N.; Elhamifar, E. Zero-shot attribute attacks on fine- grained recognition models. In Proceedings of the European Conference on Computer Vision (ECCV), Tel Aviv, Israel, 23–27 October 2022; pp. 262–282.

Share this article:
How to Cite
Guo, J.; Rao, Z.; Chen, Z.; Zhou, J.; Tao, D.; Guo, S. On the Element-Wise Representation and Reasoning in Zero-Shot Image Recognition: A Systematic Survey. Edge Intelligence and Systems 2026, 1 (1), 4.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.