OER·harvester

← Back to the library
arXiv HTML resource

Neurosymbolic Reinforcement Learning and Planning: A Survey

The area of Neurosymbolic Artificial Intelligence (Neurosymbolic AI) is rapidly developing and has become a popular research topic, encompassing sub-fields such as Neurosymbolic Deep Learning (Neurosymbolic DL) and Neurosymbolic Reinforcement Learning (Neurosymbolic RL). Compared to traditional learning methods, Neurosymbolic AI offers significant advantages by simplifying complexity and providing transparency and e…

Licence
OPEN CC-BY-4.0
Authors
K. Acharya, W. Raza, C. M. J. M. Dourado, A. Velasquez, H. Song
Published
2023-09-02 · arXiv
Language
en
Length
14073 words
Type
narrative text

Cites 33 works

inferred
Open original ↗

II Milestone in Reinforcement Learning

Reinforcement Learning (RL) has a rich history dating back to the 1940s when B.F. Skinner introduced the concept of operant conditioning in psychology, while Walter Pitts and Warren McCulloch[23] presented a computational model based on the functioning of the human brain. Donald Hebb’s Hebbian Learning Rule[24] introduced in 1949 also formed the basis for modern neural networks. In the same year, there was use of Monte Carlo method in the nuclear reactor for predicting the behaviour of neutron during second world war[25]. In the late 1950s, two important concepts: Dynamic Programming[26] and Markov Decision Process(MDP)[27] were purposed which developed the mathematical formulation of RL, also Frank Rosenblatt developed the perceptron[28], which could learn based on associationism. Later, Temporal Difference Learning (TD Learning)[29] was introduced by Arthur Samuel, which enabled agents to learn from delayed rewards and gradually update their value estimates. Grigoryevich[30] used complex polynomial equations to statistically analyze network elements, selecting the best ones for the next layer, laying the groundwork for what would become deep learning. In the 1970s, the field of AI experienced a reduction in funding, leading only a few scientists to continue their work independently. Nonetheless, during this time, significant progress was achieved. Fukushima developed the Neocognitron neural network[31], which utilized a hierarchical multi-layer architecture to enable computers to learn visual patterns. The Neocognitron later served as a basis for the Convolutional Neural Network(CNN)[32], introduced in 1989.

The 1980s saw the emergence of the field of RL with the introduction of Actor-Critic Algorithms[33] and Q-learning[34]. Additionally, Paul Werbos introduced the backpropagation algorithm[35], which, although not widely used at the time, raised questions in cognitive psychology regarding the role of symbolic logic in human comprehension. In the following decade, the field continued to evolve with the introduction of core algorithms such as TD-Gammon[36] REINFORCE[37],Experience Replay[38] and State-Action-Reward-State-Action(SARSA)[39]. A pivotal breakthrough occurred in 1999 with the invention of the Graphics Processing Unit (GPU)[^2], which enabled RL to tackle more complex environments. This was further enhanced by the parallel computing power of NVIDIA’s Compute Unified Device Architecture (CUDA)[40] on GPUs.

The 2010s proved to be a remarkable decade for RL. In 2012, the introduction of the Arcade Learning Environment (ALE)[41] opened the gateway to the use of RL in gaming environments. DRL, which combines neural networks with RL to learn high-dimensional state-action value functions, was introduced, leading to breakthroughs in game playing and robotics, such as the Deep Q-Network (DQN)[42]. Many researchers became active in modifying existing algorithms, resulting in the development of numerous new algorithms in the RL domain in 2015, such as Trust Region Policy Optimization (TRPO)[43], Deep Deterministic Policy Gradient (DDPG)[44], Double DQN[45], Dueling DQN[46] and Prioritized Experience Replay [47]. Year 2016 saw the rise of algorithms like Asynchronous Advantage Actor-Critic (A3C)[48],Generative Adversarial Imitation Learning (GAIL)[49]. The introduction of OpenAI Gym[50], an open-source toolkit for developing and comparing reinforcement learning algorithms, opened the door for exploring RL algorithms among the RL community members. In same year of 2016, RL achieved a significant milestone in the history of AI by defeating the world champion in the game of Go[2].Year of 2017 also came up with new RL algorithms like Model-Agnostic Meta-Learning (MAML)[51], Distributional RL[52],Proximal Policy Optimization (PPO)[53],Intrinsic Curiosity Module (ICM)[54] and Rainbow DQN[55]. RL model for the first time could learnt by playing itself and was able to beat its previous version which beat the world champion in the game of Go[56]. RL Algorithms like Soft Actor Critic(SAC)[57], Quality Value Iteration Optimization (QT-Opt)[58] and Imapala[59] were introduced in year 2018. RL continued to conquer many other games against humans, such as chess and shogi[60] and Starcraft II[61]. In 3D games as well RL model gave the human level performance[62].

As we have progressed into the 2020s, RL continues to be a dynamic field of research, with researchers exploring novel algorithms and applications, such as multi-agent RL, meta-RL, and RL for safety-critical systems. In addition, Alphazero has been successful in discovering faster multiplication algorithms[63] which showed that RL can also contribute in other field also as a superhuman. Recently, in late 2022, OpenAI released ChatGPT[^3], a chatbot that utilizes RL techniques to be trained and generate diverse responses to various inquiries and concerns. Reinforcement Learning from Human Feedback (RLHF)[64] has been introduced to enhance the alignment of the model with human values such as helpfulness, honesty, and harmlessness. RLHF aims to leverage the expertise and knowledge of humans to accelerate the learning process of an AI agent, allowing it to learn more efficiently and effectively. This involves fine-tuning using feedback data collected from humans. Other major technology companies are also in the race to develop their own innovative AI solutions. A summary of significant milestones in the history of RL is presented in Fig.2.

Fig. 2: RL Milestone Timeline