VI Opportunities
Real-world applications strive to minimize errors that can arise from risky exploration and exploitation, whereas Neurosymbolic RL methods employ a trial-and-error mechanism. However, to reconcile this contradiction between Neurosymbolic RL and real-world applications, a viable approach is to create an authentic simulator using real data and domain knowledge of the model dynamics. Subsequently, objectives can be designed for the agent, and the policy network can be trained in the simulator. Finally, the trained policy can be deployed in the real world with further enhancements. Though Neurosymbolic RL is in its early stage but it has started contributing to other RL areas as well. Causal Reinforcement Learning[105] has been able to produce the significant result since its collaboration with Neurosymbolic RL model. In this section, we examine the opportunities of Neurosymbolic RL methods in various fields.
VI-A Robotics and Control
Building autonomous embodied robotic systems requires designing suitable policies that ensure the system operates within reasonable mechanical constraints while maintaining safety and data efficiency. RL has been growing in robotics from very old time[20, 106]. The symbolic method has been introduced in robotic motion planning and control in 2007[107] to address these concerns. It is clear that decision-making is a crucial aspect of robotics control, and there have been various approaches to address this challenge. One notable technique is the Neurosymbolic Program Search (NSPS)[108], which produces interpretable and robust Neurosymbolic programs for autonomous driving design. Another approach is the decomposition of decision-making into two levels: what to do and how to do it[109]. This method utilizes Neurosymbolic skills and has been shown to be effective in various robotics tasks. There have also been efforts to construct robotic platforms for building manipulation environments, such as the open-source platform CausalWorld[110]. Overall, these methods and platforms have been shown to improve policy learning and performance in robotics tasks.
VI-B Gaming RL
Games are considered as suitable benchmarks with clear-cut rules and boundaries in the RL community. Over the past few years, gaming AI has exhibited extraordinary decision-making abilities, surpassing human-level performance in various decision-making games, including card games[111], board games[2], and video games[3]. Neurosymbolic RL has primarily been utilized in board games and video games, yielding state-of-the-art outcomes. It is expected that these models will also outperform others in card games. However, open-ended games such as Minecraft and XLand remain unexplored.
VI-C Intelligent Question Answering
Neurosymbolic RL has emerged as a powerful tool in natural language processing for intelligent question answering, which involves deducing the answer to a given question based on the surrounding context, often comprising both text and images. Although various studies in the context of Neurosymbolic RL have focused on knowledge graph reasoning[91, 92] for this task, combining both text and images remains a relatively unexplored research area. As such, there is considerable potential for future research to investigate and advance this topic.
VI-D Safe Reinforcement Learning
Reinforcement learning (RL) has become popular due to its ability to learn from experience and make decisions in complex environments. However, in safety-critical settings such as autonomous driving, robotics, or medical diagnosis, the failure of the system can result in severe consequences, including loss of property or human lives. Ensuring the safety of RL agents is crucial in such settings. Neurosymbolic RL has been used for the verification of the RL[93, 94] and it has given some significant result. But, it is still in its infancy period, and there are many opportunities for safely exploring the RL.
VI-E Optimizing Parameters of RL
The RL framework consists of several components, including the environment, the agent, the policy, and the reward function. Neurosymbolic RL has been applied to different components of the RL framework, combining symbolic reasoning with neural networks to solve complex RL problems. It has been successful in addressing the issue of sparse rewards by formulating reward functions that provide more informative feedback to the agent [95, 96, 97]. It has also been used to learn programmatic policies that are more generalizable and flexible to different environments [98, 99, 100, 101]. Additionally, Neurosymbolic RL has been effective in reducing the symbolic space, resulting in more efficient representations of the policy, and improving the agent’s performance [84, 85, 86, 87, 88, 89, 90].