2607004775
  • Open Access
  • Article

Hierarchical Safe Reinforcement Learning Control for UAVs via Adaptive Control Barrier Function and Interacting Multiple Model

  • Jian Gao,   
  • Jiankun Sun *

Received: 30 Jun 2026 | Revised: 25 Jul 2026 | Accepted: 30 Jul 2026 | Published: 03 Sep 2026

Abstract

This paper considers the safety-critical control problem for UAVs in dynamic environments. To handle this, we propose a novel hierarchical safe reinforcement learning control via adaptive control barrier function (ACBF) and interacting multiple model (IMM). In the proposed hierarchical control method, the upper layer employs a proximal policy optimization (PPO) algorithm to generate reference control inputs that are capable of adapting to uncertain environments. While in the lower safety filter layer, we design a three-dimensional IMM to accurately predict the future finite-horizon trajectories of dynamic obstacles. Integrating these predictions, a novel ACBF with an adaptive coefficient is proposed to modify the reference inputs generated by the upper layer, thereby rigorously ensure the strict safety of UAVs in dynamic obstacle environments. The primary advantage of the proposed hierarchical control method lies in its dual capability to fully guarantee safety and well accommodate the variability of dynamic environments. The superiorities of the proposed control method are verified via simulation results.

References 

  • 1.

    Xiao, B.; Yin, S. A New Disturbance Attenuation Control Scheme for Quadrotor Unmanned Aerial Vehicles. IEEE Trans. Ind. Inform. 2017, 13, 2922–2932.

  • 2.

    Xian, B.; Yang, S. Robust Tracking Control of a Quadrotor Unmanned Aerial Vehicle-Suspended Payload System. IEEE/ASME Trans. Mechatron. 2021, 26, 2653–2663.

  • 3.

    Najm, A.A.; Ibraheem, I.K. Nonlinear PID controller design for a 6-DOF UAV quadrotor system. Eng. Sci. Technol. Int. J. 2019, 22, 1087–1097.

  • 4.

    Dong, J.; He, B. Novel fuzzy PID-type iterative learning control for quadrotor UAV. Sensors 2019, 19, 24.

  • 5.

    Lindqvist, B.; Mansouri, S.S.; Agha-mohammadi, A.; et al. Nonlinear MPC for Collision Avoidance and Control of UAVs With Dynamic Obstacles. IEEE Robot. Autom. Lett. 2020, 5, 6001–6008.

  • 6.

    Zheng, E.-H.; Xiong, J.-J.; Luo, J.-L. Second order sliding mode control for a quadrotor UAV. ISA Trans. 2014, 53, 1350–1356.

  • 7.

    Yu, D.; Ma, S.; Liu, Y.-J.; et al. Finite-time adaptive fuzzy backstepping control for quadrotor UAV with stochastic disturbance. IEEE Trans. Autom. Sci. Eng. 2024, 21, 1335–1345.

  • 8.

    Zhu, Q.; Liu, Z.; Golestani, M.; et al. Strong robust and large time-delay-independent control for unknown nonlinear nonaffine neutral-type systems. J. Frankl. Inst. 2026, 363, 108859.

  • 9.

    Zhu, Q.; Zhang,W.; Li, S.; et al. U-control—A universal platform for control system design with inversion/cancellation of nonlinearity, dynamic and coupling through model-based to model-free procedures. Int. J. Syst. Sci. 2025, 56, 484–501.

  • 10.

    Zhu, Q.; Lei, C.; Shi, B.; et al. Model-free robust underactuated control of inverted pendulums. J. Frankl. Inst. 2025, 362, 107911.

  • 11.

    Kaufmann, E.; Bauersfeld, L.; Loquercio, A.; et al. Championlevel drone racing using deep reinforcement learning. Nature 2023, 620, 982–987.

  • 12.

    Yu, Z.; Li, J.; Xu, Y.; et al. Reinforcement Learning-Based Fractional-Order Adaptive Fault-Tolerant Formation Control of Networked Fixed-Wing UAVs With Prescribed Performance. IEEE Trans. Neural Netw. Learn. Syst. 2024, 35, 3365–3379.

  • 13.

    Wang, X.; Wang, S.; Liang, X.; et al. Deep Reinforcement Learning: A Survey. IEEE Trans. Neural Netw. Learn. Syst. 2024, 35, 5064–5078.

  • 14.

    Brunke, L.; Greeff, M.; Hall, A.W.; et al. Safe learning in robotics: From learning-based control to safe reinforcement learning. Annu. Rev. Control Robot. Auton. Syst. 2022, 5, 411–444.

  • 15.

    Liu, J.; Yang, J.; Mao, J.; et al. Flexible Active Safety Motion Control for Robotic Obstacle Avoidance: A CBF-Guided MPC Approach. IEEE Robot. Autom. Lett. 2025, 10, 2686–2693.

  • 16.

    Tessler, C.; Mankowitz, D.J.; Mannor, S. Reward constrained policy optimization. arXiv 2018, arXiv:1805.11074.

  • 17.

    Achiam, J.; Held, D.; Tamar, A.; et al. Constrained Policy Optimization. In Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, 6–11 August 2017; Volume 70, pp. 22–31.

  • 18.

    Tan, X.; Dimarogonas, D.V. Distributed Implementation of Control Barrier Functions for Multi-agent Systems. IEEE Control Syst. Lett. 2022, 6, 1879–1884.

  • 19.

    Rao, J.; Xiang, C.; Xi, J.; et al. Path planning for dual UAVs cooperative suspension transport based on artificial potential field-A* algorithm. Knowl. Based Syst. 2023, 277, 110797.

  • 20.

    Hu, J.;Wang, M.; Zhao, C.; et al. Formation control and collision avoidance for multi-UAV systems based on Voronoi partition. Sci. China Technol. Sci. 2020, 63, 65–72.

  • 21.

    Zhou, M.; Wang, Z.; Wang, J.; et al. Multi-Robot Collaborative Hunting in Cluttered Environments with Obstacle-Avoiding Voronoi Cells. IEEE/CAA J. Autom. Sin. 2024, 11, 1643–1655.

  • 22.

    Ames, A.D.; Grizzle, J.W.; Tabuada, P. Control barrier function based quadratic programs with application to adaptive cruise control. In Proceedings of the 53rd IEEE Conference on Decision and Control, Los Angeles, CA, USA, 15–17 December 2014; pp. 6271–6278.

  • 23.

    Ames, A.D.; Coogan, S.; Egerstedt, M.; et al. Control Barrier Functions: Theory and Applications. In Proceedings of the 2019 18th European Control Conference (ECC), Naples, Italy, 25–28 June 2019; pp. 3420–3431.

  • 24.

    Sun, J.; Yang, J.; Zeng, Z. Safety-Critical Control with Control Barrier Function Based on Disturbance Observer. IEEE Trans. Autom. Control 2024, 69, 4750–4756.

  • 25.

    Ames, A.D.; Xu, X.; Grizzle, J.W.; et al. Control barrier function based quadratic programs for safety critical systems. IEEE Trans. Autom. Control 2017, 62, 3861–3876.

  • 26.

    Wang, P.; Yu, C.; Deng, F.; et al. Reinforcement Learning-Based Optimal Formation Tracking for UAVs with Safety Constraints. IEEE Trans. Neural Netw. Learn. Syst. 2026, 37, 2969–2982.

  • 27.

    Zhang, C.; Wang, S.; Meng, S.; et al. Safe Exploration of Reinforcement Learning with Data-Driven Control Barrier Function. In Proceedings of the 2022 China Automation Congress (CAC), Xiamen, China, 25–27 November 2022; pp. 1008–1013.

  • 28.

    Emam, Y.; Notomista, G.; Glotfelter, P.; et al. Safe Reinforcement Learning Using Robust Control Barrier Functions. IEEE Robot. Autom. Lett. 2025, 10, 2886–2893.

  • 29.

    Xiao, W.; Belta, C.; Cassandras, C.G. Adaptive Control Barrier Functions. IEEE Trans. Autom. Control 2022, 67, 2267–2281.

  • 30.

    Lopez, B.T.; Slotine, J.-J.E.; How, J.P. Robust Adaptive Control Barrier Functions: An Adaptive and Data-Driven Approach to Safety. IEEE Control Syst. Lett. 2021, 5, 1031–1036.

  • 31.

    Yuan, T.; Bar-Shalom, Y.; Willett, P.; et al. A multiple IMM estimation approach with unbiased mixing for thrusting projectiles. IEEE Trans. Aerosp. Electron. Syst. 2012, 48, 3250–3267.

  • 32.

    Hanumegowda, A.; Dewangan, S.; Bhupala, S.; et al. Extended Object Tracking with IMM Filter for Automotive Pre-Crash Safety Applications. In Proceedings of the 2021 18th European Radar Conference (EuRAD), London, UK, 5–7 April 2022; pp. 177–180.

  • 33.

    Mazor, E.; Averbuch, A.; Bar-Shalom, Y.; et al. Interacting multiple model methods in target tracking: A survey. IEEE Trans. Aerosp. Electron. Syst. 1998, 34, 103–123.

  • 34.

    Kim, B.; Yi, K.; Yoo, H.-J.; et al. An IMM/EKF Approach for Enhanced Multitarget State Estimation for Application to Integrated Risk Management System. IEEE Trans. Veh. Technol. 2015, 64, 876–889.

  • 35.

    Cui, R.; Yang, C.; Li, Y.; et al. Adaptive Neural Network Control of AUVs with Control Input Nonlinearities Using Reinforcement Learning. IEEE Trans. Syst. Man Cybern. Syst. 2017, 47, 1019–1029.

  • 36.

    Ding, L.; Li, S.; Gao, H.; et al. Adaptive Partial Reinforcement Learning Neural Network-Based Tracking Control for Wheeled Mobile Robotic Systems. IEEE Trans. Syst. Man Cybern. Syst. 2020, 50, 2512–2523.

  • 37.

    Xiong, Y.; Zhai, D.-H.; Tavakoli, M.; et al. Discrete-Time Control Barrier Function: High-Order Case and Adaptive Case. IEEE Trans. Cybern. 2023, 53, 3231–3239.

  • 38.

    Panerati, J.; Zheng, H.; Zhou, S.; et al. Learning to fly—A gym environment with pybullet physics for reinforcement learning of multi-agent quadcopter control. In Proceedings of the 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Prague, Czech Republic, 27 September–1 October 2021; pp. 7512–7519.

Share this article:
How to Cite
Gao, J.; Sun, J. Hierarchical Safe Reinforcement Learning Control for UAVs via Adaptive Control Barrier Function and Interacting Multiple Model. Advanced Mechatronics 2026, 1 (1), 3.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.