OER·harvester

← Back to the library
arXiv HTML resource

Neurosymbolic Reinforcement Learning and Planning: A Survey

The area of Neurosymbolic Artificial Intelligence (Neurosymbolic AI) is rapidly developing and has become a popular research topic, encompassing sub-fields such as Neurosymbolic Deep Learning (Neurosymbolic DL) and Neurosymbolic Reinforcement Learning (Neurosymbolic RL). Compared to traditional learning methods, Neurosymbolic AI offers significant advantages by simplifying complexity and providing transparency and e…

Licence
OPEN CC-BY-4.0
Authors
K. Acharya, W. Raza, C. M. J. M. Dourado, A. Velasquez, H. Song
Published
2023-09-02 · arXiv
Language
en
Length
14073 words
Type
narrative text

Cites 33 works

inferred
Open original ↗

IV Neurosymbolic Reinforcement Learning

RL, a long-standing topic in the field of AI, has faced the curse of dimensionality, but the introduction of DRL solved this problem. However, DRL has several limitations. For instance, DRL can be extremely data-inefficient. In a paper by Deepmind [55], they demonstrated that the Rainbow DQN method can achieve state-of-the-art performance in terms of both performance and data efficiency. Nevertheless, it required almost 83 hours (about 3 and a half days) of playtime in addition to the training time. Conversely, many people can achieve this level of performance in just a few minutes. Another issue with DRL is that, except for rare scenarios, domain-specific algorithms work better than DRL. In the field of robotics, Boston Dynamics[^4] is a leading research institution that focuses mainly on classical robotics techniques such as time-varying Linear Quadratic Regulator(LQR), Quadratic Programming(QP) solvers, and convex optimization. Another main issue with RL is the reward system, which can be easily functionalized, but the challenge arises when trying to encourage appropriate behavior while still making it learnable. Sparse rewards are problematic because they only supply rewards in the goal state, making them difficult to shape. Shaped rewards are easier to learn because they provide positive feedback even when the whole solution has not yet been figured out. However, the problem with shaped rewards is that they are biased. The agent becomes focused on maximizing the reward instead of finding the complete solution. For example, in a study on text summarization [83], the RL model focused on increasing the ROUGE score, which it succeeded in doing. However, it failed to achieve the actual task of generating readable summaries. In contrast, summarized text generated by the model with lower ROUGE scores was found to be more readable and efficient.

The combination of Neurosymbolic systems and RL appears to be a solution to many of the issues identified in previous DRL methods. This approach not only adds reasoning and explaining capabilities to DRL but also provides a breakthrough in the field of RL. There are multiple ways in which the Neurosymbolic counterpart can be combined with RL, each with its own unique features. In this context, we will discuss three different approaches.

IV-A Learning for Reasoning RL model

The Learning for Reasoning RL model combines a neural component with a symbolic system to improve reasoning capabilities. The neural component functions as a co-actor, while the symbolic system handles the problem of reasoning. The DNNs in the model help to reduce the symbolic space, leading to faster convergence and improved performance. In cases where the data presented to the model are unstructured, DNNs can transform them into a symbolic form that the symbolic system can utilize. Furthermore, DNNs can also distill the learning policy to the symbolic system, which enhances verifiability. Serialization characterizes the neural and symbolic counterparts in this model, as shown in Fig.7.

Possible uses of the model can be:

  • Handling Unstructured Data: This type of model can handle unstructured data by using the neural part to transform it into a symbolic form that the symbolic system can work with. This is particularly useful when dealing with real-world data that may be unstructured or lack a clear symbolic representation.
  • Faster Convergence and Improved Performance: By leveraging DNNs, the model can benefit from their ability to learn complex patterns and generalize from data. This lead to faster convergence during training and improved performance in terms of accuracy and efficiency.
  • Verifiability and Interpretability: The model enhances verifiability by distilling the learned policies from the neural component to the symbolic system. The symbolic system can understand and interpret the reasoning process, making it easier to verify and understand the decisions made by the model.

Fig. 7: Learning for Reasoning RL model

IV-B Reasoning for Learning RL model

The Reasoning for Learning RL model is a different approach that utilizes symbolic models to guide the output of the neural network. By incorporating structured knowledge from the symbolic system, the performance and interpretability of the DNNs can be improved. The symbolic model can also help with reward shaping to enable faster convergence and improved performance of the DNNs agent. Additionally, the symbolic system can aid in generating the programmatic policy, making the RL model more interpretable and explainable. This type of model is characterized by parallelization, as shown in Fig.8. Application areas of the Reasoning for Learning RL model can be:

  • Improved Performance and Interpretability: By incorporating structured knowledge from the symbolic system, the model enhances the performance and interpretability of the neural network component. The symbolic system supplies guidance and constraints to the neural network, improving its decision-making and enabling more transparent and understandable outputs.
  • Efficient Reward Shaping: The symbolic system in the RL model can assist in reward shaping, which involves designing reward functions that guide the learning process. By leveraging the structured knowledge from the symbolic system, reward shaping can be more effective, leading to faster convergence and improved performance of the agent.
  • Interpretable and Explainable RL: The model’s integration of a symbolic system allows for the generation of programmatic policies. This means that the RL model’s decision-making process can be expressed in a human-readable and interpretable form, making it easier to understand and explain the reasoning behind the model’s actions.

Fig. 8: Reasoning for Learning RL model

IV-C Learning-Reasoning RL model

In the Learning-Reasoning RL model, the neural and symbolic components work bidirectionally, where the output of one can be the input of the other. This approach combines the benefits of both the Learning for Reasoning RL and Reasoning for Learning RL models, resulting in a balanced combination of interpretability and reasoning capacity. The symbolic part provides the structured knowledge to DNNs to enhance their interpretability and performance, while the neural component reduces the symbolic space, enabling the symbolic counterpart to achieve faster convergence. Moreover, the two parts work together which make the RL model more interpretable and explainable. This type of model is characterized by bidirectional communication, as depicted in Fig.9. Some of the areas where this model can contribute are:

  • Enhanced Interpretability and Reasoning Capacity: By integrating the neural and symbolic components bidirectionally, the model achieves a balanced combination of interpretability and reasoning ability. The symbolic part provides structured knowledge to the DNNs, improving their interpretability, while the neural component reduces the symbolic space and enables faster convergence of the symbolic counterpart. This combination enhances the overall interpretability and reasoning capacity of the RL model.
  • Improved Performance and Decision-Making: The bidirectional communication between the neural and symbolic components allows for an exchange of information and knowledge. Structural from the symbolic system can guide the decision-making of the neural network, leading to improved performance and more informed choices. At the same time, the neural component can transform unstructured data into a symbolic form that the symbolic system can utilize, enabling the integration of both types of information for more effective decision-making.
  • Handling Complex and Hybrid Domains: This model is particularly well-suited for complex and hybrid domains that require a combination of symbolic reasoning and neural network learning. It can handle scenarios where structured knowledge, reasoning, and interpretation are essential, while also benefiting from the learning capacity and flexibility of neural networks.The neural component can distill the learning policy to the symbolic system, enhancing its verifiability and enabling more efficient learning. The symbolic system, in turn, provides guidance and constraints to the neural network, aiding in reward shaping and reducing the search space for better convergence.

The aforementioned three RL models are currently in their initial stages, undergoing extensive research to optimize their utilization. Table III presents a comparative summary of all three RL models.

Fig. 9: Learning-Reasoning RL model

Parameters Learning for Reasoning RL Reasoning for Learning RL model Learning-Reasoning RL
Neural Component Provide abstraction to symbolic component Generate the actions Provide abstraction to symbolic component
Symbolic Component Generate action Provide Regularization to the Neural component Provide Regularization along with generating the actions
Connection Structure Serial Parallel Bi-directional
Interaction Unidirectional(Neural to symbolic) Unidirectional(Symbolic to Neural) Bi-directional
Key Focus Improved reasoning Improved learning Balance learning and reasoning
Application Areas Transforming unstructured data into a symbolic representation, Knowledge Graph Reasoning, Verification, Gaming Reward Shaping, Programatic Policy Design, Task Segmentation, Knowledge Initialized Model Task Segmentation

TABLE III: Comparison of Neurosymbolic RL models