2608004957
  • Open Access
  • Article

Skeleton-Based Guidance and Human-in-the-Loop Supervision in VLM Adaptation for Active Vision Defect Inspection

  • Chenghao Huang 1,   
  • Rui Qian 1,   
  • Liqiao Xia 1,2,*,   
  • Pai Zheng 1,2

Received: 25 Apr 2026 | Revised: 02 Aug 2026 | Accepted: 17 Aug 2026 | Published: 24 Aug 2026

Abstract

Defect inspection is critical to industrial safety and reliability, where passive vision with full-area inspection remains limited by partial occlusion and restricted camera fields of view. Currently, active vision provides a promising alternative by enabling adaptive real-time viewpoint adjustment. However, existing active vision methods still suffer from large data requirements and limited generalization across materials, textures, and imaging conditions. Vision-language model (VLM) has shown strong potential for intelligent visual reasoning, motivating the development of a tailored fine-tuning strategy for active vision inspection. Accordingly, this paper proposes an active vision inspection framework through VLM adaptation. First, a skeleton-based guidance method is introduced to estimate the main propagation axis (MPA) of defects and provide a robust geometric prior for viewpoint planning. Subsequently, a human-in-the-loop supervision pipeline is developed to establish highquality expert inspection priors for model adaptation. Based on these priors, adversarial low-rank adaptation (AdLoRA) is employed to adapt a pre-trained VLM for active vision inspection. Experimental results show that the proposed method outperforms the compared methods in prediction accuracy and improves inspection efficiency relative to passive vision.

References 

  • 1.

    Yuan, Q.; Shi, Y.; Li, M. A Review of Computer Vision-Based Crack Detection Methods in Civil Infrastructure: Progress and Challenges. Remote Sens. 2024, 16, 2910. https://doi.org/10.3390/rs16162910.

  • 2.

    Deng, J.; Singh, A.; Zhou, Y.; et al. Review on Computer Vision-Based Crack Detection and Quantification Methodologies for Civil Structures. Constr. Build. Mater. 2022, 356, 129238. https://doi.org/10.1016/j.conbuildmat.2022.129238.

  • 3.

    Feng, J.; Qi, F.; Zhu, F.; et al. Uncertainty-Informed Cascaded Diagnosis of Compound Faults in Electromechanical Systems via a Physics-Informed Hypergraph Framework. Reliab. Eng. Syst. Saf. 2026, 272, 112685. https://doi.org/10.1016/j.ress.2026.112685.

  • 4.

    Fan, J.; Zheng, P.; Li, S. Vision-Based Holistic Scene Understanding towards Proactive Human–Robot Collaboration. Robot. Comput.-Integr. Manuf. 2022, 75, 102304. https://doi.org/10.1016/j.rcim.2021.102304.

  • 5.

    Tang, W.; Jahanshahi, M.R. Autonomous Robotic Inspection Based on Active Vision and Deep Reinforcement Learning. In Proceedings of the 14th International Workshop on Structural Health Monitoring, Stanford, CA, USA, 12–14 September 2023. https://doi.org/10.12783/shm2023/36973.

  • 6.

    Wang, Y.; Peng, T.; Wang, W.; et al. High-Efficient View Planning for Surface Inspection Based on Parallel Deep Reinforcement Learning. Adv. Eng. Inform. 2023, 55, 101849. https://doi.org/10.1016/j.aei.2022.101849.

  • 7.

    Tang, W.; Jahanshahi, M.R. Active Perception Based on Deep Reinforcement Learning for Autonomous Robotic Damage Inspection. Mach. Vis. Appl. 2024, 35, 110. https://doi.org/10.1007/s00138-024-01591-7.

  • 8.

    Cao, J.; He, H.; Zhang, Y.; et al. Crack Detection in Ultrahigh-Performance Concrete Using Robust Principal Component Analysis and Characteristic Evaluation in the Frequency Domain. Struct. Health Monit. 2023, 23, 1013–1024. https://doi.org/10.1177/14759217231178457.

  • 9.

    Jin, S.; Lee, S.E.; Hong, J.W. A Vision-Based Approach for Autonomous Crack Width Measurement with Flexible Kernel. Autom. Constr. 2020, 110, 103019. https://doi.org/10.1016/j.autcon.2019.103019.

  • 10.

    Payab, M.; Abbasina, R.; Khanzadi, M. A Brief Review and a New Graph-Based Image Analysis for Concrete Crack Quantification. Arch. Comput. Methods Eng. 2019, 26, 347–365. https://doi.org/10.1007/s11831-018-9263-6.

  • 11.

    Tang, W.; Jahanshahi, M.R. Uncertainty-Aware Autonomous Robotic Inspection Based on Active Vision and Deep Reinforcement Learning. In Proceedings of the 15th International Workshop on Structural Health Monitoring, Stanford, CA, USA, 9–11 September 2025. https://doi.org/10.12783/shm2025/37531.

  • 12.

    Li, X.; Gandhi, B.; Zhan, M.; et al. Fine-Tuning Vision-Language Models for Visual Navigation Assistance. arXiv 2025, arXiv:2509.07488.

  • 13.

    Wang, X.; Dengxiong, X.; Bai, S.; et al. VLAbot: A Human Vision–Language–Action Models Interaction Framework for Robotic Assembly. Robot. Comput.-Integr. Manuf. 2026, 100, 103268. https://doi.org/10.1016/j.rcim.2026.103268.

  • 14.

    Ren, M.; Zheng, P.; Colley, M. VLM-Driven Risk-Adaptive HUD Interactions for Trust Calibration in Automated Driving. In Proceedings of the Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems, Barcelona, Spain, 13–17 April 2026. https://doi.org/10.1145/3772363.3798997.

  • 15.

    Xu, H.; Zhou, L.; Zhou, X.; et al. An Augmented Reality-Enhanced Large Language Model-Based Interactive System for Human-Robot Collaboration. SSRN 2026, 6511489. https://doi.org/10.2139/ssrn.6511489.

  • 16.

    Xia, L.; Hu, Y.; Pang, J.; et al. Leveraging Large Language Models to Empower Bayesian Networks for Reliable Human-Robot Collaborative Disassembly Sequence Planning in Remanufacturing. IEEE Trans. Ind. Inform. 2025, 21, 3117–3126. https://doi.org/10.1109/tii.2024.3523551.

  • 17.

    Sung, Y.L.; Cho, J.; Bansal, M. VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language Tasks. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 5227–5237. https://doi.org/10.1109/cvpr52688.2022.00516.

  • 18.

    Zanella, M.; Ben Ayed, I. Low-Rank Few-Shot Adaptation of Vision-Language Models. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA, 17–18 June 2024. https://doi.org/10.1109/cvprw63382.2024.00166.

  • 19.

    Wang, Z.; Pang, J. Fast Measuring Pavement Crack Width by Cascading Principal Component Analysis. arXiv 2025, arXiv:2511.02144.

  • 20.

    Chen, C.; Seo, H.; Jun, C.H.; et al. Pavement Crack Detection and Classification Based on Fusion Feature of LBP and PCA with SVM. Int. J. Pavement Eng. 2022, 23, 3274–3283. https://doi.org/10.1080/10298436.2021.1888092.

  • 21.

    Weng, X.; Huang, Y.; Wang, W. Segment-Based Pavement Crack Quantification. Autom. Constr. 2019, 105, 102819. https://doi.org/10.1016/j.autcon.2019.04.014.

  • 22.

    Zeng, X.; Zaenker, T.; Bennewitz, M. Deep Reinforcement Learning for Next-Best-View Planning in Agricultural Applications. In Proceedings of the 2022 International Conference on Robotics and Automation (ICRA), Philadelphia, PA, USA, 23–27 May 2022; pp. 2323–2329.

  • 23.

    Qiao, Y.; Yu, Z.; Wu, Q. VLN-PETL: Parameter-Efficient Transfer Learning for Vision-and-Language Navigation. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 15397–15406. https://doi.org/10.1109/iccv51070.2023.01416.

  • 24.

    Hancock, A.J.; Wu, X.; Zha, L.; et al. Actions as Language: Fine-Tuning VLMs into VLAs without Catastrophic Forgetting. arXiv 2025, arXiv:2509.22195.

  • 25.

    Ghiasvand, S.; Oskouie, H.E.; Alizadeh, M.; et al. Few-Shot Adversarial Low-Rank Fine-Tuning of Vision-Language Models. arXiv 2025, arXiv:2505.15130.

  • 26.

    Ji, Y.; Liu, Y.; Zhang, Z.; et al. Enhancing Adversarial Robustness of Vision-Language Models through Low-Rank Adaptation. In Proceedings of the International Conference on Multimedia Retrieval, Chicago, IL, USA, 30 June–3 July 2025. https://doi.org/10.1145/3731715.3733324.

  • 27.

    Umrajkar, V. DAC-LoRA: Dynamic Adversarial Curriculum for Efficient and Robust Few-Shot Adaptation. arXiv 2025, arXiv:2509.20792.

  • 28.

    Fu, J.; Fang, J.; Sun, J. LoFT: LoRA-Based Efficient and Robust Fine-Tuning Framework for Adversarial Training. In Proceedings of the 2024 International Joint Conference on Neural Networks, Yokohama, Japan, 30 June–5 July 2024; pp. 1–8.

  • 29.

    Yuan, Z.; Zhang, J.; Shan, S. FullLoRA: Efficiently Boosting the Robustness of Pretrained Vision Transformers. IEEE Trans. Image Process. 2025, 34, 4580–4590. https://doi.org/10.1109/tip.2025.3587598.

  • 30.

    Dunn, E.; Frahm, J.M. Next Best View Planning for Active Model Improvement. In Proceedings of the British Machine Vision Conference, London, UK, 7–10 September 2009; pp. 1–11.

  • 31.

    Zeng, R.; Wen, Y.; Zhao, W.; et al. View Planning in Robot Active Vision: A Survey of Systems, Algorithms, and Applications. Comput. Vis. Media 2020, 6, 225–245. https://doi.org/10.1007/s41095-020-0179-3.

  • 32.

    Ultralytics. Crack Segmentation Dataset. Available online: https://docs.ultralytics.com/datasets/segment/crack-seg (accessed on 23 August 2026).

  • 33.

    Yang, A.; Li, A.; Yang, B.; et al. Qwen3 Technical Report. arXiv 2025, arXiv:2505.09388.

Share this article:
How to Cite
Huang, C.; Qian, R.; Xia, L.; Zheng, P. Skeleton-Based Guidance and Human-in-the-Loop Supervision in VLM Adaptation for Active Vision Defect Inspection. Human-Machine Interaction 2026, 1 (1), 2.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.