2605004035
  • Open Access
  • Review

Reinforcement Learning Based Optimal Control: A Survey of Adaptive Dynamic Programming for Manipulators and Wheeled Mobile Robots

  • Chi Pui Chan 1,   
  • Ningwei Bai 1,   
  • Qichen Yin 1,   
  • Junan Wang 2,   
  • Boming Hu 3,   
  • Yunda Yan 1,   
  • Zezhi Tang 1,*

Received: 26 Feb 2026 | Revised: 29 Apr 2026 | Accepted: 26 May 2026 | Published: 15 Jul 2026

Abstract

Adaptive Dynamic Programming (ADP) has emerged as a promising Reinforcement Learning (RL)-based optimal control approach for robotics, offering a middle ground between modelbased optimal control and purely data-driven deep RL.While previous surveys have organized the ADP literature by algorithm type or control problem, the significance of platform-specific physical constraints for ADP controller design is not widely addressed. This survey presents a roboticsoriented review of ADP for two representative platforms, namely robotic manipulators and Mobile Wheeled Robots (MWRs). The literature is reviewed on the basis of how platform characteristics shape controller architectures, modelling assumptions, and stability analysis. We compare studies employing typical ADP structures (actor–critic, single-critic, decentralized, and event-triggered formulations), approaches to robustness guarantees, hardware validation, and practical deployment limitations. Several recurring patterns are identified, including the predominance of Uniform Ultimate Boundedness (UUB) as the primary stability guarantee, the limited scale of hardware validation, and the tension between scalability, safety, and real-time implementation. Finally, open challenges and future directions are discussed, including certified safety, high-dimensional scalability, sample efficiency, sim-to-real transfer, and perception-driven ADP.

References 

  • 1.

    Truong, X.; Ngo, T. Toward socially aware robot navigation in dynamic and crowded environments: A proactive social motion model. IEEE Trans. Autom. Sci. Eng. 2017, 14, 1743–1760.

  • 2.

    Coelho, D.; Oliveira, M. A review of end-to-end autonomous driving in urban environments. IEEE Access 2022, 10, 75296–75311.

  • 3.

    Gehring, C.; Fankhauser, P.; Isler, L.; et al. Anymal in the field: Solving industrial inspection of an offshore HVDC platform with a quadrupedal robot. In Proceedings of the Field and Service Robotics: Results of The 12th International Conference, Tokyo, Japan, 29–31 August 2019; pp. 247–260.

  • 4.

    Bualat, M.; Barlow, J.; Fong, T.; et al. Astrobee: Developing a free-flying robot for the international space station. In Proceedings of the AIAA SPACE 2015 Conference and Exposition, Pasadena, CA, USA, 31 August–2 September 2015; p. 4643.

  • 5.

    Sutton, R.; Barto, A. Reinforcement Learning: An Introduction; MIT press: Cambridge, UK; 1998; Vol. 1.

  • 6.

    Kober, J.; Bagnell, A.; Peters, J. Reinforcement learning in robotics: A survey. Int. J. Robot. Res. 2013, 32, 1238–1274.

  • 7.

    Kaelbling, L.; Littman, M.; Moore, A. Reinforcement learning: A survey. J. Artif. Intell. Res. 1996, 4, 237–285.

  • 8.

    Puterman,M. Markov Decision Processes: Discrete Stochastic Dynamic Programming; JohnWiley & Sons: Hoboken, NJ, USA, 2014.

  • 9.

    Mnih, V.; Kavukcuoglu, K.; Silver, D.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533.

  • 10.

    Schulman, J.; Wolski, F.; Dhariwal, P.; et al. Proximal policy optimization algorithms. arXiv 2017. preprint, arXiv:1707.06347.

  • 11.

    Garcıa, J.; Fern´andez, F. A comprehensive survey on safe reinforcement learning. J. Mach. Learn. Res. 2015, 16, 1437–1480.

  • 12.

    Recht, B. A tour of reinforcement learning: The view from continuous control. Annu. Rev. Control Robot. Auton. Syst. 2019, 2, 253–279.

  • 13.

    Liu, D.; Xue, S.; Zhao, B.; et al. Adaptive dynamic programming for control: A survey and recent advances. IEEE Trans. Syst. Man Cybern. Syst. 2020, 51, 142–160.

  • 14.

    Kiumarsi, B.; Vamvoudakis, K.; Modares, H.; et al. Optimal and autonomous control using reinforcement learning: A survey. IEEE Trans. Neural Netw. Learn. Syst. 2017, 29, 2042–2062.

  • 15.

    Bellman, R.; Corporation, R.; Collection, K.M.R. Dynamic Programming; Rand Corporation research study, Princeton University Press: Princeton, NJ, USA, 1957.

  • 16.

    Powell, W. Approximate Dynamic Programming: Solving the Curses of Dimensionality; John Wiley & Sons: Hoboken, NJ, USA, 2007; Vol. 703.

  • 17.

    Werbos, P. Beyond Regression: New Tools for Prediction and Analysis in the Behavioral Sciences. PhD Thesis, Committee on Applied Mathematics, Harvard University, Cambridge, MA, USA, 1974.

  • 18.

    Prokhorov, D.; Wunsch, D. Adaptive critic designs. IEEE Trans. Neural Netw. 1997, 8, 997–1007.

  • 19.

    Sutton, R. Learning to predict by the methods of temporal differences. Mach. Learn. 1988, 3, 9–44.

  • 20.

    Bertsekas, D.; Tsitsiklis, J. Neuro-dynamic programming: An overview. In Proceedings of 1995 34th IEEE Conference on Decision and Control, New Orleans, LA, USA, 13–15 December 1995; Vol. 1, pp. 560–564.

  • 21.

    Lewis, F.; Liu, K.; Yesildirek, A. Neural net robot controller with guaranteed tracking performance. IEEE Trans. Neural Netw. 1995, 6, 703–715.

  • 22.

    Doya, K. Reinforcement learning in continuous time and space. Neural Comput. 2000, 12, 219–245.

  • 23.

    Vrabie, D.; Pastravanu, O.; AbuKhalaf, M.; et al. Adaptive optimal control for continuous-time linear systems based on policy iteration. Automatica 2009, 45, 477–484.

  • 24.

    Modares, H.; Lewis, F.; NaghibiSistani, M. Integral reinforcement learning and experience replay for adaptive optimal control of partially-unknown constrained-input continuous-time systems. Automatica 2014, 50, 193–202.

  • 25.

    Lv, Y.; Na, J.; Ren, X. Online H∞ control for completely unknown nonlinear systems via an identifier-critic-based ADP structure. Int. J. Control 2019, 92, 100–111.

  • 26.

    Tang, Z.; Rossiter, J.A.; Panoutsos, G. A reinforcement learningbased approach for optimal output tracking in uncertain nonlinear systems with mismatched disturbances. In Proceedings of the 2024 UKACC 14th International Conference on Control (CONTROL), Winchester, UK, 10–12 April 2024; pp. 169–174.

  • 27.

    Tang, Z.; Rossiter, J.A.; Dong, Y.; et al. Reinforcement learningbased output stabilization control for nonlinear systems with generalized disturbances. In Proceedings of the 2024 IEEE International Conference on Industrial Technology (ICIT), Bristol, UK, 25–27 March 2024; pp. 1–6.

  • 28.

    Qin, C.; Pang, M.; Wang, Z.; et al. Event-Triggered Zero-Sum Game for Input Saturated Nonlinear Systems with State Constraints. IEEE Trans. Consum. Electron. 2026, 72, 455–465.

  • 29.

    Wang, N.; Jia, W.; Wu, H.; et al. Event-triggered self-organizing swarm control of distributed unmanned surface vehicles. IEEE Trans. Intell. Transp. Syst. 2024, 26, 3431–3445.

  • 30.

    Dong, L.; Zhong, X.; Sun, C.; et al. Event-triggered adaptive dynamic programming for continuous-time systems with control constraints. IEEE Trans. Neural Netw. Learn. Syst. 2016, 28, 1941–1952.

  • 31.

    Bai, N.; Chan, C.P.; Yin, Q.; et al. Deep Reinforcement Learning Optimization for Uncertain Nonlinear Systems via Event-Triggered Robust Adaptive Dynamic Programming. arXiv 2025, preprint, arXiv:math/2512.15735.

  • 32.

    Qin, C.; Pang, M.; Wang, Z.; et al. Observer based fault tolerant control design for saturated nonlinear systems with full state constraints via a novel event-triggered mechanism. Eng. Appl. Artif. Intell. 2025, 161, 112221.

  • 33.

    Li, B.; Zhan, H.; Zhang, X.; et al. Single-network based adaptive dynamic programming optimal control for robotic manipulator with joint limitation under input saturation. Int. J. Syst. Sci. 2024, 55, 3355–3370.

  • 34.

    Xue, S.; Liu, Z.;Wang, L.; et al. Adaptive dynamic programming based event-triggered multi-H∞ control. Neurocomputing 2025, 638, 130157.

  • 35.

    Wang, Q.; Zhang, C.; Hu, B. Dynamic programming with metareinforcement learning: A novel approach for multi-objective optimization. Complex Intell. Syst. 2024, 10, 5743–5758.

  • 36.

    Lin, W.; Yang, P. Adaptive critic motion control design of autonomous wheeled mobile robot by dual heuristic programming. Automatica 2008, 44, 2716–2723.

  • 37.

    Szuster, M.; Hendzel, Z. Discrete Globalised Dual Heuristic Dynamic Programming in Control of the Two-Wheeled Mobile Robot. Math. Probl. Eng. 2014, 2014, 628798.

  • 38.

    Stojanovic, V.; Djordjevic, V.; Dubonjic, L.; et al. ADP-based control of a two-wheeled self-balancing mobile robot. Eng. Today 2024, 3. https://doi.org/10.5937/engtoday2400018S.

  • 39.

    Azimi, A.; Shamshiri, R.; Ghasemzadeh, A. Adaptive dynamic programming for robust path tracking in an agricultural robot using critic neural networks. Agric. Eng. 2025, 80. https://doi.org/10.15150/ae.2025.3327.

  • 40.

    Liu, X.; Sun, Y.; Han, Q.; et al. A novel adaptive dynamic optimal balance control method for wheel-legged robot. Appl. Math. Model. 2025, 137, 115737.

  • 41.

    Liang, J.; Tang, S.; Jia, B. Control of Parallel Quadruped Robots Based on Adaptive Dynamic Programming Control. Machines 2024, 12, 875.

  • 42.

    Wei, X.; Ye, J.; Xu, J.; et al. Adaptive dynamic programmingbased cross-scale control of a hydraulic-driven flexible robotic manipulator. Appl. Sci. 2023, 13, 2890.

  • 43.

    Xia, H.; Xue, D.; Wang, Y.; et al. Critic-Identifier Structure ADP Based Near-optimal Decentralized Tracking Control of Modular and Reconfigurable Robots. In Proceedings of the 2019 34rd Youth Academic Annual Conference of Chinese Association of Automation (YAC), Jinzhou, China, 6–8 June 2019; pp. 440–445.

  • 44.

    Tang, Y.; Dou, L.; Zhang, R.; et al. Adaptive dynamic programming–based fault-tolerant control of multi-UAV formation system. Trans. Inst. Meas. Control 2025, 47, 1893–1905.

  • 45.

    Luo, Z.; Zhang, P.; Ding, X.; et al. Adaptive affine formation maneuver control of second-order multi-agent systems with disturbances. In Proceedings of the 2020 16th International Conference on Control, Automation, Robotics and Vision (ICARCV), Shenzhen, China, 13–15 December 2020; pp. 1071–1076.

  • 46.

    Liu, H.; Yi, X.; Liu, D.; et al. A review of unmanned vehicle control with adaptive dynamic programming implementations. J. Intell. Robot. Syst. 2025, 111, 10.

  • 47.

    Lewis, F.; Vrabie, D. Reinforcement learning and adaptive dynamic programming for feedback control. IEEE Circuits Syst. Mag. 2009, 9, 32–50.

  • 48.

    AbuKhalaf, M.; Lewis, F. Nearly optimal control laws for nonlinear systems with saturating actuators using a neural network HJB approach. Automatica 2005, 41, 779–791.

  • 49.

    Vamvoudakis, K.; Lewis, F. Online actor–critic algorithm to solve the continuous-time infinite horizon optimal control problem. Automatica 2010, 46, 878–888.

  • 50.

    Jiang, Y.; Jiang, Z. Computational adaptive optimal control for continuous-time linear systems with completely unknown dynamics. Automatica 2012, 48, 2699–2704.

  • 51.

    Kamalapurkar, R.; Andrews, L.;Walters, P.; et al. Model-based reinforcement learning for infinite-horizon approximate optimal tracking. IEEE Trans. Neural Netw. Learn. Syst. 2016, 28, 753–758.

  • 52.

    Khalil, H.K.; Grizzle, J.W. Nonlinear Systems; Prentice Hall: Upper Saddle River, NJ, USA, 2002; Vol. 3.

  • 53.

    Bhat, S.P.; Bernstein, D.S. Finite-time stability of continuous autonomous systems. SIAM J. Control Optim. 2000, 38, 751–766.

  • 54.

    Polyakov, A. Nonlinear feedback design for fixed-time stabilization of linear control systems. IEEE Trans. Autom. Control 2011, 57, 2106–2106.

  • 55.

    Lakshmikantham, V.; Leela, S.; Martynyuk, A.A. Practical Stability of Nonlinear Systems; World Scientific: Singapore, 1990.

  • 56.

    Brunke, L.; Greeff, M.; Hall, A.W.; et al. Safe learning in robotics: From learning-based control to safe reinforcement learning. Annu. Rev. Control Robot. Auton. Syst. 2022, 5, 411–444.

  • 57.

    Ames, A.D.; Xu, X.; Grizzle, J.W.; et al. Control barrier function based quadratic programs for safety critical systems. IEEE Trans. Autom. Control 2016, 62, 3861–3876.

  • 58.

    Ames, A.D.; Coogan, S.; Egerstedt, M.; et al. Control barrier functions: Theory and applications. In Proceeding of the 2019 18th European Control Conference (ECC), Naples, Italy, 25–28 June 2019; pp. 3420–3431.

  • 59.

    Cheng, R.; Orosz, G.; Murray, R.M.; et al. End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks. In Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA, 27 January–1 February 2019; Vol. 33, pp. 3387–3395.

  • 60.

    Taylor, A.; Singletary, A.; Yue, Y.; et al. Learning for safetycritical control with control barrier functions. In Proceedings of the 2nd Conference on Learning for Dynamics and Control, Virtual Event, 10–11 June 2020; pp. 708–717.

  • 61.

    Wabersich, K.P.; Zeilinger, M.N. A predictive safety filter for learning-based control of constrained nonlinear dynamical systems. Automatica 2021, 129, 109597.

  • 62.

    Hewing, L.; Wabersich, K.P.; Menner, M.; et al. Learning-based model predictive control: Toward safe learning in control. Annu. Rev. Control Robot. Auton. Syst. 2020, 3, 269–296.

  • 63.

    Bansal, S.; Chen, M.; Herbert, S.; et al. Hamilton-jacobi reachability: A brief overview and recent advances. In Proceedings of the 2017 IEEE 56th Annual Conference on Decision and Control (CDC), Melbourne, VIC, Australia, 12–15 December; pp. 2242–2253.

  • 64.

    Fisac, J.F.; Akametalu, A.K.; Zeilinger, M.N.; et al. A general safety framework for learning-based control in uncertain robotic systems. IEEE Trans. Autom. Control 2018, 64, 2737–2752.

  • 65.

    Gierlak, P.; Szuster, M.; ˙ Zylski, W. Discrete dual–heuristic programming in 3DOF manipulator control. In Proceedings of the International Conference on Artificial Intelligence and Soft Computing, Zakopane, Poland, 13–17 June 2010; pp. 256–263.

  • 66.

    Szuster, M.; Gierlak, P. Approximate dynamic programming in tracking control of a robotic manipulator. Int. J. Adv. Robot. Syst. 2016, 13, 16.

  • 67.

    Zhang, X.; Yang, Z.; Liu, H.; et al. Optimal Sliding Mode Fault-Tolerant Control for Multiple Robotic Manipulators via Critic-Only Dynamic Programming. Sensors 2025, 25, 5410.

  • 68.

    Yang, X.; He, H. Adaptive dynamic programming for decentralized stabilization of uncertain nonlinear large-scale systems with mismatched interconnections. IEEE Trans. Syst. Man Cybern. Syst. 2018, 50, 2870–2882.

  • 69.

    Zhao, B.; Liu, D. Event-triggered decentralized tracking control of modular reconfigurable robots through adaptive dynamic programming. IEEE Trans. Ind. Electron. 2019, 67, 3054–3064.

  • 70.

    Zhou, F.; Nie, F.; An, T.; et al. Decentralized fault tolerant control of modular manipulators system based on adaptive dynamic programming. Int. J. Control Autom. Syst. 2022, 20, 3252–3253.

  • 71.

    Sveen, E.M.; Zhou, J.; Ebbesen, M.K.; et al. Decentralised adaptive learning-based control of robot manipulators with unknown parameters. J. Autom. Intell. 2025, 4, 136–144.

  • 72.

    Luy, N.T. Robust adaptive dynamic programming based online tracking control algorithm for real wheeled mobile robot with omni-directional vision system. Trans. Inst. Meas. Control 2017, 39, 832–847.

  • 73.

    Dian, S.; Fang, H.; Zhao, T.; et al. Modeling and trajectory tracking control for magnetic wheeled mobile robots based on improved dual-heuristic dynamic programming. IEEE Trans. Ind. Inform. 2020, 17, 1470–1482.

  • 74.

    Wang, C.; Zhan, H.; Guo, Q.; et al. Adaptive Dynamic Programming-Based Fixed-Time Optimal Control for Wheeled Mobile Robot. IEEE Robot. Autom. Lett. 2024, 10, 176–183.

  • 75.

    Li, S.; Ding, L.; Gao, H.; et al. ADP-based online tracking control of partially uncertain time-delayed nonlinear system and application to wheeled mobile robots. IEEE Trans. Cybern. 2019, 50, 3182–3194.

  • 76.

    Li, X.; Wang, L.; An, Y.; et al. Dynamic path planning of mobile robots using adaptive dynamic programming. Expert Syst. Appl. 2024, 235, 121112.

  • 77.

    Ghasemzadeh, A.; Amjadifard, R.; KeymasiKhalaji, A. Adaptive dynamic programming for trajectory tracking control of a tractor-trailer wheeled mobile robot. IET Control Theory Appl. 2025, 19, e12784.

  • 78.

    Luo, Y.; Li, Y.; Ding, J.; et al. Optimal moving-target circumnavigation control of multiple wheeled mobile robots based on adaptive dynamic programming. IEEE Trans. Netw. Sci. Eng. 2024, 11, 4679–4688.

  • 79.

    Zhang, Y.; Zhao, B.; Liu, D. Distributed optimal containment control of wheeled mobile robots via adaptive dynamic programming. IEEE Trans. Syst. Man Cybern. Syst. 2025, 55, 5876–5886.

  • 80.

    Lewis, F.; Vrabie, D.; Syrmos, V. Optimal Control; John Wiley & Sons: Hoboken, NJ, USA, 2012.

  • 81.

    Wang, Y.; An, T.; Dong, B.; et al. Adaptive dynamic programming-based finite-time optimal backstepping force/position control of reconfigurable robot manipulators via Pareto optimal. IEEE Trans. Autom. Sci. Eng. 2025, 22, 10660-10671.

  • 82.

    Tang, Z.; Rossiter, J.A.; Jin, X.; et al. Output Tracking for Uncertain Time-Delay Systems via Robust Reinforcement Learning Control. In Proceedings of the 2024 43rd Chinese Control Conference (CCC), Kunming, China, 28–31 July 2024; pp. 2219–2226.

  • 83.

    Nachum, O.; Gu, S.S.; Lee, H.; et al. Data-efficient hierarchical reinforcement learning. In Proceedings of the Advances in Neural Information Processing Systems 31 (NeurIPS 2018), Montreal, Canada, 3–8 December 2018; Vol. 31.

  • 84.

    Lee, E.S.; Zhou, L.; Ribeiro, A.; et al. Graph neural networks for decentralized multi-agent perimeter defense. Front. Control Eng. 2023, 4, 1104745.

  • 85.

    Vaswani, A.; Shazeer, N.; Parmar, N.; et al. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems 30 (NIPS 2017), Long Beach, CA, USA, 4–9 December 2017; Vol. 30.

  • 86.

    Qin, C.; Sun, K.; Sun, A.; et al. Dynamic event-triggered optimal safety control of interconnected nonlinear systems with discontinuous state constraints by adaptive dynamic programming. Neurocomputing 2025, 669, 132467.

  • 87.

    Qin, C.; Hou, S.; Pang, M.; et al. Reinforcement learning-based secure tracking control for nonlinear interconnected systems: An event-triggered solution approach. Eng. Appl. Artif. Intell. 2025, 161, 112243.

  • 88.

    Chen, J.; Ran, X. Deep learning with edge computing: A review. Proc. IEEE 2019, 107, 1655–1674.

  • 89.

    Mittal, S. A survey on optimized implementation of deep learning models on the nvidia jetson platform. J. Syst. Archit. 2019, 97, 428–442.

  • 90.

    Guo, K.; Zeng, S.; Yu, J.; et al. [DL] A survey of FPGA-based neural network inference accelerators. ACM Trans. Reconfig. Technol. Syst. (TRETS) 2019, 12, 1–26.

Share this article:
How to Cite
Chan, C. P.; Bai, N.; Yin, Q.; Wang, J.; Hu, B.; Yan, Y.; Tang, Z. Reinforcement Learning Based Optimal Control: A Survey of Adaptive Dynamic Programming for Manipulators and Wheeled Mobile Robots. Advanced Mechatronics 2026, 1 (1), 1.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.