OER·harvester

← Back to the library
arXiv HTML resource

Neurosymbolic Reinforcement Learning and Planning: A Survey

The area of Neurosymbolic Artificial Intelligence (Neurosymbolic AI) is rapidly developing and has become a popular research topic, encompassing sub-fields such as Neurosymbolic Deep Learning (Neurosymbolic DL) and Neurosymbolic Reinforcement Learning (Neurosymbolic RL). Compared to traditional learning methods, Neurosymbolic AI offers significant advantages by simplifying complexity and providing transparency and e…

Licence
OPEN CC-BY-4.0
Authors
K. Acharya, W. Raza, C. M. J. M. Dourado, A. Velasquez, H. Song
Published
2023-09-02 · arXiv
Language
en
Length
14073 words
Type
narrative text

Cites 33 works

inferred
Open original ↗

Neurosymbolic Reinforcement Learning (RL) is an emerging area of AI that is currently lacking in literature. This section aims to provide an overview of the implementation stages of Neurosymbolic RL, including the state of the art, current trends, and proposed research studies. Various notable works have been analyzed, detailing the neural and symbolic components used in Table III, and information about the RL algorithm, reward space, action space, and policy module in Table IV. These works have been classified into three main RL models: Learning for Reasoning, Reasoning for Learning, and Learning-Reasoning, and further sub-divided according to their areas of application. A summary of this classification can be found in Table V. The classification and analysis of these works provides insights into the use of Neurosymbolic RL models and their potential for further development in the future.

V-A Learning for Reasoning RL model

This particular Neurosymbolic RL model involves the use of a neural network as an auxiliary tool to extract crucial symbols and information, which helps to reduce the search space of the symbolic system. This results in a faster problem-solving process, making it especially useful for problems that require reasoning. This architecture has been used for the following primary goals:

V-A1 Transforming unstructured data into a symbolic representation

Symbolic systems, which rely on logical rules and representations of symbolic data, are often limited in their ability to process unstructured data such as images, videos, and natural language text. Most of the real world data are inherently present in the unstructured form so there must be a model to transform them to the symbolic form before being processed by the symbolic models. DNNs have been shown to be very effective in processing and generating such unstructured data and can be used to generate structured data that can be used as input to a symbolic system.

Deep Symbolic Reinforcement Learning (DSRL)[84] consists of two main components: a deep neural network that learns a low-level continuous representation of the state space and map it to low-dimensional symbolic space, and a symbolic model that distills the learned policy into a more interpretable form by mapping symbolic representation to action. The authors use this framework to learn policies for a range of environments, and demonstrate that their system outperforms traditional reinforcement learning algorithms in terms of interpretability, generalization, and efficiency.Symbolic Reinforcement Learning with Common Sense (SRL+CS)[85] a novel extension of DSRL where the authors create a meaningful symbolic representation of the world using sub-states before applying learning and decision-making algorithms. They have two modification in Q-value function: restricting the updates to specific sub-state and assigning importance on the basis of distance of the objects. This approach provides better generalization and explainability. Neural Symbolic Reinforcement Learning (NSRL)[86], includes a reasoning module based on neural attention networks, which performs relational reasoning on symbolic states and induces the RL policy, enabling end-to-end learning with prior symbolic knowledge. It can extract the logical rules selected by the attention modules instead of storing all the rules, saving memory budget and improving scalability.Deep Symbolic Policy[87], uses an autoregressive recurrent neural network to generate symbolic policies, which are optimized using a risk-seeking policy gradient. To scale to environments with multi-dimensional action spaces, the authors propose an ”anchoring” algorithm that distills pre-trained neural network-based policies into fully symbolic policies. The authors also introduce two novel methods to improve exploration in DRL-based combinatorial optimization which are hierarchical entropy regularizer and a soft length prior.

Detect, Understand, Act (DUA) [88], composed of three components: Detect, which consists of a traditional computer vision object detector and tracker, Understand, which provides an answer set programming (ASP) paradigm for symbolically implementing a meta-policy over options, and Act, which houses a set of options that are high-level actions enacted by pre-trained DRL policies. The paper evaluates the DUA framework on the Animal-AI (AAI) competition testbed and achieves state-of-the-art results in multiple categories. It is modular approach, allowing for straightforward generalization and transfer to other complex tasks. Another study, Symbolic Options for Reinforcement Learning (SORL)[89], proposes a method for automatically discovering and learning symbolic options, which are higher-level actions with specified preconditions and postconditions, to assist deep reinforcement learning (DRL) agents in complex environments. It was successful in mitigating the problem of sparse and delayed reward along with improving efficiency.Neurosymbolic Logic Neural Network (LNN) for RL algorithm[90], supplies fast convergence and interpretability for RL policies in text-based interaction games by extracting first-order logical facts from text observation using semantic parser(ConceptnNet) and history, then trains the symbolic rules with logical functions in the neural networks.

V-A2 Knowledge Graph Reasoning

Knowledge Graph (KG) reasoning is the task of inferring new information from a given KG, which consists of a set of entities and their relationships. KG reasoning is important for a wide range of applications, including question answering, information retrieval, and recommender systems.

DeepPath[91], uses a policy-based agent with continuous states based on knowledge graph embeddings to sample the most promising relation and extend the multi-hop relational path. The authors demonstrate that their method outperforms path-ranking based algorithms and knowledge graph embedding methods on two standard reasoning tasks on Freebase and Never-Ending Language Learning datasets. Meandering In Networks of Entities to Reach Verisimilar Answers (MINERVA)[92]outperforms DeepPath, as Deeppath cannot be applied to query answering tasks where the second entity is unknown.It uses neural reinforcement learning to learn how to navigate the knowledge graph conditioned on the input query to find predictive paths.

V-A3 Verification

The process of verification aims to determine if a model meets a particular desired property, and can play a key role in enhancing quality and safety. In the context of reinforcement learning (RL), it is important to verify a model’s convergence, correctness, and robustness to ensure it functions effectively.

Verifiability via Iterative Policy ExtRaction(VIPER)[93],is a modification of Q-DAGGER algorithm used for verifying the correctness of deep reinforcement learning policies. The approach involves extracting a small, interpretable model from a deep neural network policy, which can then be verified using existing verification techniques. The extracted model approximates the original policy well and can be used to analyze the policy’s convergence, correctness, and robustness.Another model, Reinforcement Learning with Verified Exploration (REVEL)[94], incorporates a differentiable symbolic planner that generates a set of safe exploration actions, which the RL agent executes to find optimal policies while avoiding potentially unsafe states. The proposed framework is formally verified using the Coq proof assistant, which ensures that the system is free from runtime errors and satisfies desired safety properties.

V-A4 Gaming

Neurosymbolic RL is a promising approach in the field of game playing, where the goal is to develop agents that can learn to play games at a human-like level or beyond. It involves combining neural networks and symbolic systems to develop agents that can learn the rules of the game and develop strategies to play the game effectively. By combining symbolic reasoning with deep learning, Neurosymbolic RL models can provide explanations for the decisions made by the agent, making it easier for humans to understand and evaluate the agent’s performance.AlphaGo Zero[56], achieves superhuman performance in the game of Go without using human data or domain knowledge. The system is based on a combination of deep neural networks and Monte Carlo tree search, with the neural networks trained through a reinforcement learning process using self-play. It beats the older AlphaGo version in game Go with score of 100-0.

Research NN Knowledge Base
[84] CNN First Order Logic
[85] CNN First Order Logic
[86] Transformer First Order Logic
[87] RNN Decision Tree
[88] DNN Answer Set Programming
[89] DNN Propositional Logic
[90] LNN First Order Logic
[91] DNN Knowledge Graph
[92] LSTM Knowledge Graph
[93] DNN Decision Tree
[94] DNN Symbolic Policies
[56] DNN Decision Tree
[95] CNN Finite Trace Linear Temporal Logic
[96] CNN Finite Trace Linear Temporal Logic
[97] DNN Omega Regular Language
[98] DNN Programmatic Policy
[99] DNN Programmatic Policy
[100] NN State Machines
[101] RNN Programmatic Policy
[102] DNN Deterministic Finite Automaton
[103] DNN Propositional Logic
[104] CNN First Order Logic

TABLE IV: Neural and symbolic components in related works

V-B Reasoning for Learning RL model

In this type of Neurosymbolic RL model, symbolic system acts as a helper, it provides symbolic knowledge to the neural network from where the final decision is made. This approach is particularly useful in complex applications, such as robotics, where the environment is often uncertain and dynamic, and where the use of symbolic knowledge can facilitate high-level reasoning and decision-making. It has been used in the following application area:

V-B1 Reward Shaping

In the field of RL, one of the primary challenges is dealing with sparse rewards. One solution to this problem is reward shaping, which involves incorporating domain knowledge. Rather than relying on a single, final reward, intermediate rewards are provided to the agent for exhibiting desirable behavior. This encourages the agent to take effective actions early on in the learning process, leading to faster convergence.

Monte Carlo Tree Search with Automaton-Guided Reward Shaping (MCTS-A)[95], helps to improve the learning process and final performance of the agent in domains with sparse rewards. The method introduces an automaton that guides the reward shaping process, allowing for a dynamic and flexible approach. The automaton’s states correspond to the different learning phases of the agent, and each state has its own shaping rules that change over time. Transfer learning between the two different environments with the same objective is also eased with this approach. Authors in[96], extended the work by introducing Multiagent Tree Search Algorithm with reward shaping(MATS-A) so that it can be applied to multi-agent scenario and can handle both stochastic and deterministic transition in Multi-agent Non-Markovian Reward Decision Process. They prove that sharing the same search tree and DFA objective can be used to develop competitive and cooperative behavior among the agents, within and across the team.Research work [97] first converts an omega-regular specification into a Buchi automaton. It is then used to construct an average reward objective which can then be optimized by standard RL algorithms.The authors prove that the learned policy converges to the optimal policy and demonstrate the effectiveness of the method.

V-B2 Programmatic Policy Design

A programmatic policy refers to a decision-making algorithm that governs the behavior of an agent that can make decisions. The programmatic policy takes inputs from the environment and computes a set of actions that the agent should take in response.The design of programmatic policies can vary based on the complexity of the task and available data, including decision trees, state machines, and programs.

Imitation-Projected Programmatic Reinforcement Learning (PROPEL)[98], proposes a new approach for combining imitation learning with reinforcement learning, called imitation-projected programmatic reinforcement learning (IP-PRL). The approach uses a programmatic policy to encode a priori knowledge about the task and trains an agent through a combination of imitation learning and reinforcement learning. The agent first learns to imitate an expert’s behavior through supervised learning, and then the agent’s policy is updated through reinforcement learning while being constrained to stay close to the expert’s policy. IP-PRL outperforms both pure imitation learning and pure reinforcement learning in terms of sample efficiency and final performance. Another work Programmatically Interpretable Reinforcement Learning (PIRL)[99], uses Neurally Directed Program Synthesis (NDPS) algorithm to generate interpretable neural policies which can be verified through the symbolic approach. It first learn a neural policy network using deep reinforcement learning and then performing a local search over programmatic policies that seeks to minimize the distance from this neural oracle. [100] proposes a novel approach for synthesizing policies for automated decision-making systems that can generalize to new situations. The authors use inductive programming techniques, specifically ”program synthesis by example”, to generate policies that satisfy a set of example-based specifications. Framework[101] synthesize programmatic policies that are more interpretable and generalizable than neural network policies produced by deep reinforcement learning methods. It uses a program representation and only requires minimal supervision compared to prior programmatic reinforcement learning and program synthesis works. It learns a program embedding space that parameterizes diverse behaviors in an unsupervised manner and then searches over this space to find a program that maximizes the return for a given task.

V-B3 Task Segmentation

Main goal or task is broken down in to the smaller task with their own set of rewards so that the task become more generalizable and reasonable. DeepSynth[102], uses automata synthesis to automatically segment a task. A task was broken down into smaller subgoals, each with its own reward. The proposed approach learns a model of an automaton that represents the state machine of a task and uses it to segment the task into subgoals. The learned automaton is then used to guide the agent in finding the optimal policy for the task.

V-B4 Knowledge Initialized Model

Researches has found that the model give higher convergence rate, reasoning ability if initialized with knowledge base.Before starting the learning process, the knowledge base of the agent is initialized with some prior information instead of starting from zero. Propositional Logic Nets (PROLONETS)[103],enables warm start of learning process by efficient initialization of RL agents using human-specified policies, without requiring an Imitation Learning (IL) phase. This approach helps RL agents to navigate complex environments that pose challenges to randomly initialized models, and allows for greater exploration. It outperforms baseline RL approaches such as IL and knowledge-based techniques.

V-C Learning-Reasoning RL model

This Neurosymbolic RL model uses symbolic and neural systems as an alternate process. They both complement each other by performing abstraction and regularization to give the final results.Symbolic Deep Reinforcement Learning (SDRL) [104], uses planner-controller-meta-controller architecture where planner uses prior symbolic knowledge for long term planning, controller uses DRL algorithms for intrinsic rewards and meta-controller evaluate training performance of controller based on extrinsic rewards along with proposing new intrinsic goals to the planner.

Research RL-Algorithm State Space Action Space Policy Module
[84] Q-Learning Multi-dimensional vector Multi-dimensional vector Tabular Q Learning
[85] Q-Learning Multi-dimensional vector Multi-dimensional vector Q-Table
[86] Double DQN Set of Predicates Set of Predicates Multi-layer Perceptron
[87] Policy Gradients Multi-dimensional vector Multi-dimensional vector RNN
[88] PPO Multi-dimensional vector Discrete action space NN
[89] Double Q-Learning High level dimension sapce 5-dimensional vector Option Set
[90] DQN Multi-dimensional vector Discrete set of 10 different actions LNN
[91] Policy Gradients Entities in Knowledge Graph Relations in Knowledge Graph Fully connected NN
[92] REINFORCE Entities in Knowledge Graph Relations in Knowledge Graph LSTM
[93] VIPER Multi-dimensional vector Leaf nodes in Decision Tree Decision Tree
[94] Policy Gradients Real vector space Real vector space NN
[56] Policy Gradients Muti-dimensional vector Mutli-dimensional vector DNN
[95] MCTS-A Nodes of Automata Transitions of Automata CNN
[96] MATS-A Nodes of Automata Transitions of Automata CNN
[97] Differential Q-Learning Nodes in GFM automaton Relation in GFM automaton DNN
[98] Policy Gradients Continuous Continuous Programatic and Neural
[99] NDPS Unconstrained Policy Space Continuous Programatic and Deterministic
[100] Gradient Based Optimization Continuous Continuous State Machine
[101] REINFORCE Program Embedding Space Program Execution Trace RNN
[102] DQN Multi-dimensional vector Multi-dimensional vector DNN
[103] PPO 193D and 37D 44D and 10D Decision Tree
[104] Double Q-Learning High dimensional Set of Primitive Action DNN

TABLE V: RL components of the related works

RL model Areas of Application Related Researches
Learning for Reasoning Transforming unstructured data into a symbolic representation [84][85][86][87][88][89][90]
Knowledge Graph Reasoning [91][92]
Verification [93][94]
Gaming [56]
Reasoning for Learning Reward Shaping [95][96] [97]
Programatic Policy Design [98][99][100][101]
Task Segmentation [102]
Knowledge Initialized Model [103]
Learning-Reasoning Task Segmentation [104]

TABLE VI: Summary of classification of related researches