III Background and Preliminaries
In this section, we keep focus on the background information related to the topics Neurosymbolic AI followed by RL.
III-A Neurosymbolic AI
The field of artificial intelligence has been centered around the goal of developing machines that can achieve human-like levels of intelligence. Two major approaches have been pursued in this effort. The first, symbolic AI, is a rule-based approach that was prevalent from the 1950s to the 1980s. The second approach is a data-based approach known as connectionist AI. While symbolic AI requires a large amount of information to be supplied, it can learn from this information on its own. The primary disadvantage of connectionist AI is its inability to explain the reasoning or logical processes behind the model, leading to these models being referred to as black boxes. Symbolic reasoning provides an explainable inference process and employs powerful declarations to represent knowledge, as well as offering benefits such as fast initial coding, explicit method control, and abstraction of knowledge[65]. However, this approach is limited in its ability to handle vast amounts of incomplete data and to generalize from such data. Psychologist Daniel Kahneman has distinguished between two human cognitive processes, system 1 and system 2. System 1 is fast, automatic, and unconscious, akin to deep learning, while system 2 is slow, effortful, and conscious, similar to symbolic AI[66]. In the context of AI, there have been discussions of ways to combine these two approaches, as the authors of a study[67] conclude that only a combination of both fields is likely to enable the development of human-like intelligence.
Neurosymbolic AI is a subfield of AI that combines two historically prominent approaches: connectionist AI and symbolic AI. This integration enables more efficient derivation of knowledge and general concepts from data, focusing on learning from experience and reasoning about what has been learned from uncertain environments. Hybrid Neurosymbolic systems require less training data and are capable of tracking the steps required to draw conclusions and make inferences which is the reason Neurosymbolic AI has been regarded as the $3^{rd}$ wave of AI[68]. By combining symbolic reasoning with deep learning, ideal results can be obtained with a limited number of datasets, error correction with recoveries, and enhanced explanatory capabilities that are not possible with deep learning alone[69]. Numerous applications necessitate both learning and reasoning abilities. On the neural aspect, models learn from data provided to them, while the symbolic aspect seeks to retain the innate explanatory power of these systems. The Neurosymbolic AI domain, as previously discussed, can be employed to develop various applications across different fields, such as medical diagnostic systems, recommender systems, and text mining[70]. By incorporating deep human expert knowledge into the system’s design and function, Neurosymbolic AI can be leveraged to its fullest potential in creating such applications.Fig.3 depicts the evaluation of the Neurosymbolic AI process within a model design that integrates neural network and symbolic artificial intelligence, harnessing the full strength of both fields in hybrid models.

Fig. 3: General Representation Neurosymbolic AI
Numerous researchers have provided insights into how the fields of neural and symbolic AI can be combined practically. Three noteworthy works have contributed significantly to organizing the research in Neurosymbolic systems. The first notable work was a survey paper published in 2005 by Sebastian and Pascal [71]. They identified three main axes of Neurosymbolic integration: Interrelation, Language, and Usage. Each of these axes was further divided into several sub-divisions. Fig.4 provides a simplified visualization of the eight dimensions along with their axes. Another researcher, Henry Kautz[72], proposed a way to classify Neurosymbolic systems into six different categories. He gave them distinctive names, which are detailed in Table I. In a recent survey[6], Neurosymbolic systems were analyzed based on three parameters: efficiency, generalization, and interpretability. The authors proposed a novel taxonomy consisting of three different classes: learning for reasoning, reasoning for learning, and learning-reasoning. Table II provides detailed descriptions of each of these classes.

Fig. 4: Classification by Sebastian and Pascal[71]
| Classification | Characteristic Features |
|---|---|
| Symbolic Neuro symbolic | Symbolic input is converted to feature vectors for the neural networks which give final results in the symbolic form |
| Symbolic[Neuro] | Neural pattern recognition subroutine within a symbolic problem solver |
| Neuro $\|$ Symbolic | A cascade from neural network into symbolic reasoner |
| Neuro: Symbolic → Neuro | Symbolic rules are input which are compiled so that their knowledge end up in the neural network |
| Neuro_{Symbolic} | Uses direct encodings of logical statements into neural structures |
| Neuro[Symbolic] | Embed symbolic reasoning inside neural engine to enable both superhuman and super combinatorial reasoning |
TABLE I: Classification by Henry Kautz[72]
| Classification | Characteristic Features |
|---|---|
| Learning for reasoning | Neural network play the role of the helper, it extracts the important symbols and information so that the search space of the symbolic system narrowed down |
| Reasoning for learning | Symbolic system act as a helper, it provides symbolic knowledge to the neural network from where the final decision is made |
| Learning-reasoning | Uses symbolic and neural systems as an alternate process. They both complement each other to give the final results |
TABLE II: Classification by D. Yu and et al. [6]
III-B Reinforcement Learning

Fig. 5: RL Methods and Techniques
The fast-learning algorithms and wide-ranging applications of RL have made it increasingly popular in academia and industry, thanks to significant technological advancements [73, 74]. In earlier literature, RL was described as a class of problems that an agent encounters in a dynamic, unpredictable environment and solves through trial and error. Nowadays, RL is viewed as a machine learning paradigm that trains an agent to make decisions based on its immediate surroundings to optimize rewards. The training process involves a loop of interaction with the environment, including observing, receiving rewards, making decisions, and obtaining feedback signals [75]. RL has proven its ability to solve complex real-world problems, such as natural language processing, image classification, speech recognition, and decision-making, which has improved planning and perception in various applications [76].RL is an essential component of autonomous driving cars and robots, which can perform tasks such as food preparation without human intervention or specific programming. RL-based strategies could play a crucial role in enabling fully autonomous systems in the future [77]. RL employs algorithms and methods to enable an agent to obtain optimal control in an environment, and the agents in RL can range from a game player to a stock trading bot. Another field which is similar to RL is Intrinsic Motivation(IM) but it lack feedback mechanism. Many research papers in RL have utilized IM to address complex problems in sparse reward platforms [78, 79]. The interactions that occur between an agent and its environment are typically modeled as a Markov Decision Problem (MDP)[80], or a Partially Observable Markov Decision Process (POMDP). An MDP is a framework used for sequential decision-making in Markovian dynamical systems, which extends the Multi-Armed Bandits (MAB) framework by allowing the system state to stochastically change based on the actions taken and their resulting outcomes. On the other hand, a POMDP is a newer version of MDP, where the system state is not directly observable. In certain cases, MDPs can be solved analytically, while in many cases, they can be solved iteratively through the use of dynamic or linear programming. When no model is present, RL methods can be employed to obtain sample trajectories and directly interact with the system[81]. As the number of computing devices continues to increase rapidly, it is expected that the number of devices capable of handling complex and dynamic systems with minimal programming will grow exponentially, potentially reaching billions.
Looking at the bigger picture, RL can be classified as a type of sample-based approach for solving MDP problems. The RL technique uses sample trajectories and the agent’s interaction with the system, which can be obtained from a simulation. This approach is quite common in practical applications, where a simulation is available, and a clear transition-probability model is not required. In such scenarios, dynamic or linear programming may not be suitable, making the RL method a more practical option[8].Main Components of Reinforcement Learning are:
- Policy: It refers to the way an agent behaves at a given time, which is generally a mapping from perceived states to the action that needs to be taken when in those states of the environment. The goal of the policy is to maximize the expected cumulative reward received by the agent over time. There are various types of policies, including deterministic policies and stochastic policies, which define the agent’s behavior in different ways.
- Reward Signal: It refers to the objective or goals of the problem and is a number delivered to the agent by the environment at each time step. The reward signal is used to train the agent to learn a behavior that maximizes the cumulative reward over time. It is a crucial component as it guides the agent to take actions that lead to achieving the desired goals.
- Value Function: It represents the expected long-term cumulative reward that an agent can obtain by following a specific policy. It estimates the value of each state or state-action pair, which allows the agent to choose the best action in each state. The value function can be expressed mathematically as the expected sum of discounted future rewards starting from a given state or state-action pair. The estimation of the value function can be done through various methods, such as Monte Carlo methods, Temporal Difference learning, or Bellman equations.
- Model of the Environment: It refers to the representation of how the environment behaves in response to the actions taken by the agent. It allows the agent to predict the next state and reward given the current state and action. The model can be either known or unknown, and the goal is to use it to optimize the agent’s behavior. In cases where the model is known, dynamic programming techniques can be used to find the optimal policy. The model can be represented in different forms, including transition probabilities, state-transition diagrams, or function approximators.
Reinforcement Learning (RL) can be categorized into various types based on different parameters, such as the environment, policy, model, and others. These categories provide a framework to understand and classify different RL approaches. Fig.5 summarizes the various types of RL based on these parameters, including environment type, horizon type, optimization type, and more.
The success of RL largely depends on the quality of its algorithm. Numerous RL algorithms have been developed, tailored to specific contexts, and based on various parameters such as environment type, action space type, and model type. These algorithms are continuously modified to improve their performance and expand their scope of applications[82]. Fig.6 provides a brief overview of the various types of RL algorithms that have been used to date.

Fig. 6: Overview of RL Algorithms