Ijraset Journal For Research in Applied Science and Engineering Technology
Authors: Sachidananda M H, Prashanth Kumar R, Darshan P R
DOI Link: https://doi.org/10.22214/ijraset.2026.84647
Certificate: View Certificate
Neural-network-based reinforcement learning has driven a broad set of breakthroughs across many application areas. In the specific case of robotic manipulation, these methods raise the prospect of machines acquiring dexterous, human-like manipulation skills straight from visual input. This survey examines where reinforcement learning algorithms currently stand within this field. It lays out the core theory and terminology involved, and highlights the principal obstacles that still restrict these algorithms from being deployed on real-world robotic problems. The paper closes with the authors\' outlook on promising directions for continued research.
This text is a survey of deep reinforcement learning (RL) for robotic manipulation, focusing on research published from 2013 onward. It reviews the fundamental concepts of RL, major algorithmic approaches, their application to robotic manipulation, and the key challenges that prevent RL from being widely deployed on real robots. The survey was motivated by the lack of a comprehensive review specifically covering deep RL for robot manipulation during this period and identified 42 relevant papers through searches of IEEE Xplore, Google Scholar, and arXiv.
RL involves an agent interacting with an environment by observing states, taking actions, and receiving rewards. Its main components are:
The survey explains the Markov Decision Process (MDP) as the mathematical foundation of RL and introduces state-value and action-value functions. The Bellman equations provide the basis for recursively estimating these values and determining optimal actions.
The survey organizes RL methods primarily along two dimensions:
Model-based vs. model-free
Policy-based vs. value-based
Thus, the major trade-off is between stability and sample efficiency.
Robot manipulation presents particularly difficult conditions for RL because it involves:
The survey identifies three major challenges as especially important:
A. Sample Inefficiency
Sample inefficiency is described as perhaps the largest obstacle to applying RL to real robotic manipulation. Many algorithms require enormous amounts of interaction data, while collecting physical robot experience is slow and expensive.
The survey compares several approaches, finding that evolutionary methods tend to be least data-efficient, followed by policy-gradient approaches and replay-buffer methods such as DDPG and NAF. Model-based methods generally achieve much higher sample efficiency, with approaches such as PILCO and Guided Policy Search demonstrating particularly strong results.
Potential solutions include:
RL agents must balance exploration—trying new actions to gain information—with exploitation—using actions that are already believed to produce good results.
Different algorithms address this differently:
However, exploration remains difficult in robotics because action spaces are continuous and extremely large. Random exploration can be inefficient, and poorly designed exploration may lead to local optima or unsafe robot behavior.
The survey highlights the lack of standardized benchmarks for comparing exploration strategies and emphasizes that safety is especially important when exploration occurs on physical robots.
For RL to become practically useful, robots must be able to generalize beyond the exact conditions used during training. However, many existing methods are specialized for individual tasks and perform poorly when environments or tasks change.
Two main approaches to generalization are identified:
The survey also identifies reproducibility as a major problem. RL results can change significantly depending on implementation details, neural-network architecture, reward scaling, environment, and especially the random seed. Consequently, results based on only a few experimental runs may not reliably represent an algorithm's true performance.
This survey has outlined the current state of RL algorithms as applied to robot manipulation. Despite considerable progress in simulated domains such as games, RL\'s real-world impact on physical robots remains fairly limited. Today\'s best RL algorithms achieve strong performance in domains governed by simple, fully known rules, such as board games like Go and chess. When confronted with unfamiliar dynamics, robots can currently manage only simple manipulation tasks, and only given ample training samples. Relative to the breadth of tasks humans can learn quickly and perform well, there is still substantial ground to cover before truly intelligent robots become a reality. The research community remains split across several distinct directions, but the authors believe that progress will most likely come from combining the strengths of some or all of these approaches — or possibly from entirely new classes of algorithms not yet developed. Even so, the authors remain optimistic about RL\'s role in the future of robot manipulation, viewing algorithms in this family as a likely necessary ingredient for building genuinely intelligent robots.
[1] H. M. La, R. Lim, and W. Sheng, “Multirobot cooperative learning for predator avoidance,” IEEE Transactions on Control Systems Technology, vol. 23, no. 1, pp. 52–63, 2015. [2] M. Rahimi, S. Gibb, Y. Shen, and H. M. La, “A comparison of various approaches to reinforcement learning algorithms for multi-robot box pushing,” in Intern. Conf. on Engineering Research and Applications. Springer, 2018, pp. 16–30. [3] H. X. Pham, H. M. La, D. Feil-Seifer, and A. Nefian, “Cooperative and distributed reinforcement learning of drones for field coverage,” 2018. [4] V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wier stra, and M. Riedmiller, “Playing atari with deep reinforcement learn ing,” in NIPS Deep Learning Workshop, 2013. [5] D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton et al., “Mastering the game of go without human knowledge,” Nature, vol. 550, no. 7676, p. 354, 2017. [6] J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The Intern. J. of Robotics Research, vol. 32, no. 11, pp. 1238–1274, 2013. [7] R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction. MIT press, 2018. [8] T. Salimans, J. Ho, X. Chen, S. Sidor, and I. Sutskever, “Evolution strategies as a scalable alternative to reinforcement learning,” arXiv preprint arXiv:1703.03864, 2017. [9] V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, p. 529, 2015. [10] V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in Intern. Conf. on Machine Learning, 2016, pp. 1928–1937. [11] J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in Intern. Conf. on Machine Learning, 2015, pp. 1889–1897. [12] T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” arXiv preprint arXiv:1509.02971, 2015. [13] S. Gu, T. Lillicrap, I. Sutskever, and S. Levine, “Continuous deep q learning with model-based acceleration,” in Intern. Conf. on Machine Learning, 2016, pp. 2829–2838. [14] S. Levine and V. Koltun, “Guided policy search,” in Intern. Conf. on Machine Learning, 2013, pp. 1–9. [15] M. Deisenroth and C. Edward Rasmussen, “Pilco: A model-based and data-efficient approach to policy search.” 01 2011, pp. 465–472. [16] M. P. Deisenroth, C. E. Rasmussen, and D. Fox, “Learning to control a low-cost manipulator using data-efficient reinforcement learning,” in Robotics: Science and Systems, 2011. [17] J. Tan, T. Zhang, E. Coumans, A. Iscen, Y. Bai, D. Hafner, S. Bohez, and V. Vanhoucke, “Learning to control a low-cost manipulator using data-efficient reinforcement learning,” in Robotics: Science and Systems, 2018. [18] A. A. Rusu, M. Vecerik, T. Roth¨orl, N. Heess, R. Pascanu, and R. Hadsell, “Sim-to-real robot learning from pixels with progressive nets,” arXiv preprint arXiv:1610.04286, 2016. [19] X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Sim-to-real transfer of robotic control with dynamics randomization,” in 2018 IEEE Intern. Conf. on Robotics and Automation (ICRA). IEEE, 2018, pp. 1–8. [20] U. Viereck, A. t. Pas, K. Saenko, and R. Platt, “Learning a visuomotor controller for real world robotic grasping using simulated depth images,” arXiv preprint arXiv:1706.04652, 2017. [21] E. Tzeng, C. Devin, J. Hoffman, C. Finn, X. Peng, S. Levine, K. Saenko, and T. Darrell, “Towards adapting deep visuomotor representations from simulated to real environments,” CoRR, abs/1511.07111, 2015. [22] S. Levine, P. Pastor, A. Krizhevsky, J. Ibarz, and D. Quillen, “Learning hand-eye coordination for robotic grasping with deep learning and large scale data collection,” The Intern. J. of Robotics Research, vol. 37, no. 4-5, pp. 421–436, 2018. [23] C. J. C. H. Watkins, “Learning from delayed rewards,” Ph.D. disserta tion, King’s College, Cambridge, 1989. [24] M. Tokic, “Adaptive ?-greedy exploration in reinforcement learning based on value differences,” in Annual Conf. on Artificial Intelligence. Springer, 2010, pp. 203–210. [25] J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in Intern. Conf. on Machine Learning, 2015, pp. 1889–1897. [26] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Prox imal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017. [27] T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off policy maximum entropy deep reinforcement learning with a stochastic actor,” arXiv preprint arXiv:1801.01290, 2018. [28] S. Sukhbaatar, Z. Lin, I. Kostrikov, G. Synnaeve, A. Szlam, and R. Fergus, “Intrinsic motivation and automatic curricula via asymmetric self-play,” arXiv preprint arXiv:1703.05407, 2017. [29] M. Plappert, R. Houthooft, P. Dhariwal, S. Sidor, R. Y. Chen, X. Chen, T. Asfour, P. Abbeel, and M. Andrychowicz, “Parameter space noise for exploration,” arXiv preprint arXiv:1706.01905, 2017. [30] A. Tamar, Y. Glassner, and S. Mannor, “Optimizing the cvar via sampling.” in AAAI, 2015, pp. 2993–2999. [31] A. Rajeswaran, K. Lowrey, E. V. Todorov, and S. M. Kakade, “Towards generalization and simplicity in continuous control,” in Advances in Neural Information Processing Systems, 2017, pp. 6550–6561. [32] L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta, “Robust adversarial reinforcement learning,” arXiv preprint arXiv:1703.02702, 2017. [33] W. Yu, J. Tan, C. K. Liu, and G. Turk, “Preparing for the unknown: Learning a universal policy with online system identification,” arXiv preprint arXiv:1702.02453, 2017. [34] N. Mishra, M. Rohaninejad, X. Chen, and P. Abbeel, “Meta-learning with temporal convolutions,” arXiv preprint arXiv:1707.03141, 2017. [35] F. Sung, L. Zhang, T. Xiang, T. Hospedales, and Y. Yang, “Learning to learn: Meta-critic networks for sample efficient learning,” arXiv preprint arXiv:1706.09529, 2017. [36] P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger, “Deep reinforcement learning that matters,” arXiv preprint arXiv:1709.06560, 2017. [37] R. Islam, P. Henderson, M. Gomrokchi, and D. Precup, “Reproducibil ity of benchmarked deep reinforcement learning tasks for continuous control,” arXiv preprint arXiv:1708.04133, 2017. [38] H. Nguyen, H. M. La, and M. Deans, “Deep learning with experience ranking convolutional neural network for robot manipulator,” arXiv preprint arXiv:1809.05819, 2018. [39] A. Sehgal, H. M. La, S. J. Louis, and H. Nguyen, “Deep reinforcement learning using genetic algorithm for parameter optimization,” Submitted for Intern. Conf. on Robotic Computing (IRC), 2019 [40] K. Cobbe, O. Klimov, C. Hesse, T. Kim, and J. Schulman, “Quantifying generalization in reinforcement learning,” arXiv preprint arXiv:1812.02341, 2018. [41] H.-J. Ye, X.-R. Sheng, D.-C. Zhan, and P. He, “Distance metric facil itated transportation between heterogeneous domains.” in IJCAI, 2018, pp. 3012–3018. [42] T. Carr, M. Chli, and G. Vogiatzis, “Domain adaptation for reinforcement learning on the atari,” arXiv preprint arXiv:1812.07452, 2018. [43] J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh, “Action-conditional video prediction using deep networks in atari games,” in Advances in neural information processing systems, 2015, pp. 2863–2871. [44] S. Chiappa, S. Racaniere, D. Wierstra, and S. Mohamed, “Recurrent environment simulators,” arXiv preprint arXiv:1704.02254, 2017. [45] A. X. Lee, R. Zhang, F. Ebert, P. Abbeel, C. Finn, and S. Levine, “Stochastic adversarial video prediction,” arXiv preprint arXiv:1804.01523, 2018. [46] I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in neural information processing systems, 2014, pp. 2672 2680. [47] D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114, 2013. [48] D. Silver, H. van Hasselt, M. Hessel, T. Schaul, A. Guez, T. Harley, G. Dulac-Arnold, D. Reichert, N. Rabinowitz, A. Barreto et al., “The predictron: End-to-end learning and planning,” arXiv preprint arXiv:1612.08810, 2016. [49] A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine, “Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning,” in 2018 IEEE Intern. Conf. on Robotics and Automation (ICRA). IEEE, 2018, pp. 7559–7566. [50] E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model based control,” in Intelligent Robots and Systems (IROS), 2012 IEEE/RSJ Intern. Conf. on. IEEE, 2012, pp. 5026–5033. [51] R. Rahmatizadeh, P. Abolghasemi, L. B¨ol¨oni, and S. Levine, “Vision based multi-task manipulation for inexpensive robots using end-to-end learning from demonstration,” in 2018 IEEE Intern. Conf. on Robotics and Automation (ICRA). IEEE, 2018, pp. 3758–3765. [52] M. Wulfmeier, I. Posner, and P. Abbeel, “Mutual alignment transfer learning,” arXiv preprint arXiv:1707.07907, 2017. [53] A. Y. Ng, S. J. Russell et al., “Algorithms for inverse reinforcement learning.” in ICML, 2000, pp. 663–670.
Copyright © 2026 Sachidananda M H, Prashanth Kumar R, Darshan P R. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Paper Id : IJRASET84647
Publish Date : 2026-08-17
ISSN : 2321-9653
Publisher Name : IJRASET
DOI Link : Click Here
Submit Paper Online
