OER·harvester

← Back to the library
arXiv HTML resource

The Duality of Generative AI and Reinforcement Learning in Robotics: A Review

Recently, generative AI and reinforcement learning (RL) have been redefining what is possible for AI agents that take information flows as input and produce intelligent behavior. As a result, we are seeing similar advancements in embodied AI and robotics for control policy generation. Our review paper examines the integration of generative AI models with RL to advance robotics. Our primary focus is on the duality be…

Licence
OPEN CC-BY-4.0
Authors
Angelo Moroncelli, Vishal Soni, Marco Forgione, Dario Piga, Blerina S…
Published
2024-10-21 · arXiv
Language
en
Length
18858 words
Type
narrative text

Cites 98 works

inferred
Open original ↗

5 Task

RL tries to perform optimal decision-making considering the interaction of an agent (e.g., a robot) with its environment, with the goal of maximizing rewards [140, 11]. In the RL framework, several core components play essential roles in guiding the agent’s behavior. First, we have states, which represent all possible situations the agent might encounter within its environment. At any given moment, the agent will find itself in a particular state, prompting it to consider the best action (e.g., the action that maximizes the reward). These actions are the set of choices available to the agent, allowing it to interact with and influence the environment. The policy represents the probability of taking action given a state (e.g., it’s the agent’s strategy). As the agent takes actions, it receives feedback from the environment through the reward function, which assigns immediate rewards based on the outcome of each action. This reward signal helps the agent understand how to act optimally to achieve long-term goals. Two key functions support the agent in making more informed decisions over time. The value function estimates the expected long-term return of being in a particular state, giving the agent an idea of how promising that state is under the current policy. The Q-function, on the other hand, is slightly more detailed. It assesses the expected return not just for states, but also for specific actions taken within those states, enabling the agent to evaluate the quality of particular actions in given situations. One of the biggest challenges in RL is designing an effective reward function [139, 101]. A poorly designed reward system can mislead the agent, resulting in suboptimal or unintended behaviors. Crafting rewards that properly guide the agent toward good outcomes is crucial. Other two significant challenges are representing the state space and exploring it. If states are not represented accurately or comprehensively, the agent may struggle to learn the true dynamics of the environment, which can impede its ability to make optimal decisions [111].

This section reviews key works that exemplify how generative models have been employed to address critical aspects of the RL training loop in robotics, specifically focusing on: (i) Reward Signal generation in the presence of sparse rewards and under goal specification constraints, (ii) State Representation learning to improve sample efficiency, and (iii) Planning & Exploration mechanisms that support generalization to new tasks. Based on the modality and purpose of the generative models employed in relation to the RL Task, the works are categorized and examined. A more thorough classification of the basic models under examination is provided in Table 8.

Network Reward Signal State Representation Planning & Exploration
LLM Textual feedback Textual state encoding High-level guidance
VLM Multi-modal feedback Multi-modal state encoding Multi-modal evaluation
Diffusion Model Reward generation State generation Action sampling
WM and VPM Reward prediction State prediction Model-based planning

Table 8: Generative AI tools for RL integration. Classification of generative AI models as tools based on integration into RL, as anticipated in Section 3 and Section 4.

5.1 Reward Signal

An increasing amount of research explores how generative AI models, particularly foundation models, can help automate or enhance reward design [94, 163, 135, 39, 44, 143, 24, 72, 101, 36, 13, 84, 3, 40, 86, 104, 128, 10, 27, 49]. These models offer new ways to interpret user input, perceive environments, and produce reward signals that guide effective policy learning.

5.1.1 Reward design with LLMs

Recently, LLMs such as GPT-4 [2] or Llama 3 [47], have been used to address the challenges of reward function design and direct reward specification. While LLMs struggle to directly control robots due to their ineffectiveness to produce control commands, they are highly valuable in evaluating the performance of RL agents, performing in-context reasoning on natural language problems. Traditionally, reward functions in RL are meticulously hand-crafted by researchers, often requiring significant domain expertise and limiting the agent’s ability to adapt to new situations. However, foundation models offer a compelling alternative by providing rich, generalist knowledge representations that can be leveraged to automatically generate or inform reward functions. Imagine a robot tasked with “cleaning the kitchen”. An LLM could identify sub-tasks like “wiping the table” and “putting dishes in the dishwasher” and assign rewards based on their successful completion. This structured approach breaks down complex tasks and provides clear goals for the RL agent.

Some approaches use LLMs for zero-shot reward generation [101, 163, 135] to convert natural language descriptions of desired behaviors into mathematical reward functions. TEXT2REWARD [163] and Self-Refined LLM [135] use LLMs to generate reward functions as executable code, while EUREKA [101] employs LLMs for evolutionary optimization of reward code. These methods leverage the LLM’s ability to understand task semantics and translate them into reward functions. Others leverage the in-context learning capabilities of LLMs to learn from demonstrations or feedback provided in natural language, shaping the reward function iteratively [39, 44, 143, 36, 171]. For example, Lafite-RL [36] uses LLM feedback to iteratively refine the reward function, employing few-shot learning to guide the LLM’s generation process; while Language to Rewards [171] uses LLM feedback to define reward parameters that are optimized for specific tasks. While TEXT2REWARD and Lafite-RL focus on generating dense reward functions that provide feedback at each timestep, other methods, like Language to Rewards, may generate sparse rewards that are only provided at the end of an episode. Some methods, such as EUREKA and Language to Rewards, incorporate interactive feedback from humans to refine the reward function. In FoMo Rewards [96], LLMs are employed to evaluate the likelihood of an instruction accurately describing a task given a trajectory of observations. This likelihood then serves as a reward signal. While the use of LLMs for reward function generation is still an emerging field, it shows significant promise for improving the development and deployment of intelligent agents. This approach can streamline the reward function design process, making it more general and task agnostic. Lastly, Nair et al. [113] showcase the use of LLMs to learn language-conditioned rewards from offline data and crowd-sourced annotations. This approach avoids the need for extensive human demonstrations, offering a scalable solution for learning complex behaviors. These different approaches highlight the versatility of LLMs in reward generation for RL in robotics, while the choice of approach depends on the specific requirements of the task and the available resources. However, one main drawback we found in existing work leveraging LLMs to craft rewards is that, since they work with text-only input, they are limited in acquiring perception from the real world [103]. This especially limits their capability to simulation environments where full state information is readily available.

5.1.2 Reward design with VLMs

The integration of vision-language models [124, 125] in RL task evaluation represents a significant advancement over methods that rely solely on textual or numerical feedback [104, 128, 3, 10, 136, 151, 3, 40, 86, 167]. By leveraging the power of visual understanding, VLMs can be used for task evaluation [3, 40] based on prompts and image frames as input.

Many robotic tasks involve interacting with a complex environment filled with diverse sensory information. Vision foundation models can leverage their ability to process different data modalities (text, images) to construct richer reward functions. Imagine a robot tasked with sorting laundry. A VLM could analyze an image of the clothing to identify the fabric type and color, and combine this information with text instructions to create a reward function that promotes sorting based on pre-defined categories, performing multi-modal reward learning. In Mahmoudieh et al. [104], Rocamonde et al. [128] and Adeniji et al. [3] VLMs are employed as zero-shot reward models. These approaches leverage the in-context learning ability of VLMs to evaluate if a visual scene coming from a camera sensor effectively match the language task description. This approach eliminates the need for manually designing reward functions or gathering extensive human feedback; moreover, it does not need environment state information. Differently, in RoboCLIP [136] rewards are generated by evaluating how closely the robot’s actions match the provided example. Interestingly, Code as Reward [146] utilizes VLMs to generate reward functions through code, thereby reducing the computational overhead of directly querying the VLM every time. Wang et al. [151] takes a slightly different route, querying VLMs to express preferences over pairs of image observations based on task descriptions, and subsequently learning a reward function from these preferences.

5.1.3 Reward design with diffusion models

Similar to other types of generative AI tools, diffusion models can be leveraged to infer rewards in RL. Huang et al. [71] and Nuti et al. [117] present methods for deriving reward functions by conditional diffusion models that are trained on expert visual demonstrations to generate rewards, similarly to the usage of diffusion models for conditional image generation [129]. On a different approach, Mazoure et al. [108] introduces the Diffused Value Function (DVF) algorithm, which leverages diffusion models to estimate value functions from states and actions, enhancing efficiency without explicitly learning rewards. They show improved long-term decisions. To conclude, diffusion models are a relatively new approach to reward design, but they show strong promise; particularly due to adaptability, dense reward generation, and fast inference. We do not yet see immediate applications for text-conditioned reward generation.

5.1.4 Reward design with video prediction

The approaches for reward design presented so far in our taxonomy often lack accuracy in predicting changes in the robot’s visual state within the environment. Recent methods, such as those in Chen et al. [27] and Escontrela et al. [49], utilize the ability to predict future video frames based on current observations to shape rewards more effectively, leveraging environmental models (e.g., video prediction models) to enhance reward specification in RL systems. Specifically, Chen et al. [27] propose DVD, a discriminator that learns multi-task reward functions by classifying whether two videos perform the same task. This approach allows generalization to unseen environments and tasks by learning from a small amount of robot data and a large dataset of human videos. Escontrela et al. [49] introduce VIPER, which uses pre-trained VPMs to provide reward signals for RL based on frame observations rather than the actions the robot takes. VIPER achieves expert-level robot control across various tasks in simulations without predefined rewards, demonstrating the potential of VPMs for online reward specification during robotic task execution. At the same time, other researchers propose HOLD [7] and VIP [100]. This approaches generalize to unseen robot embodiments and environments, effectively accelerating RL training on various manipulation tasks zero-shot; the second one is trained on large-scale human videos.

5.2 State Representation

Foundation models can offer a promising avenue to address the challenge of more realistic state representations for data augmentation and training in RL [28, 116, 94, 116, 149, 169, 133, 150, 109, 158, 20]. However, the successful integration of foundation models as world representations for meaningfully training robots on augmented data and through model-based RL techniques hinges on addressing a fundamental issue mentioned many times in our discussion: grounding. While these models excel at processing abstract representations, their ability to effectively model realistic scenarios depends on establishing a meaningful connection between these representations and the real-world objects or concepts they abstract [159]. This section surveys research papers that utilize learned representations to predict and interact with environments in RL.

5.2.1 Learning representations from videos

Integrating VPMs with RL enhances the training of robots by allowing them to predict future frames from a sequence of previous images. VPMs forecast how a state will evolve, enabling the RL agent to plan actions and make decisions with more data-efficient learning and safer exploration, by anticipating future scenarios while reducing reliance on real-world trials [45, 170, 13, 105]. However, the success of this approach depends on the accuracy of the predictions, as well as managing the increased computational complexity, that often limits the possibility to use a single pre-trained model for a wide variety of robotics tasks. As a result, a common trend in research is to train or fine-tune different VPMs for specific applications, but future work could focus on scaling VPMs to improve their generalization.

Du et al. [45] present UniPi, a unique approach that casts sequential decision-making as a text-conditioned video generation problem. UniPi leverages the knowledge embedded in language and videos to generalize to novel goals and tasks across diverse environments and to transfer effectively to downstream RL tasks. This approach enables multi-task learning and action planning, showcasing the potential of video generation for policy learning. Foundation RL [170] takes this objective a step further, they trained an RL model that leverages foundation priors from large-scale pre-training for embodied agents and they tested it on simulated robotic tasks.

5.2.2 Foundation world models for model-based RL

Many RL algorithms rely on accurate models of the environment’s dynamics to plan and make decisions [139]. World models [57], particularly those trained on data that includes physical interactions, can be used to learn these models of the world. The learned model can help the RL agent to predict the consequences of its actions (e.g., its future states) and make informed decisions based on predictions [159].

World models capture the action space of a robot. They can be constructed in a variety of ways, utilizing diverse data sources and employing different learning architectures. One recent approach to building a world model for abstract representation is through the use of large language models [116, 172, 149, 132, 133]. For example, Zala et al. [172] propose EnvGen, a framework where an LLM generates and adapts training environments for small RL agents. By creating diverse environments tailored to specific skills, EnvGen enables agents to learn more efficiently in parallel. The LLM receives feedback on agent performance and iteratively refines the environments, focusing on weaker skills. On a similar approach Wang et al. [149] introduces GenSim, which uses GPT-4 [2] to automate the generation of diverse simulation tasks. Differently, Yang et al. [169] further advance the state-of-the-art in learning interactive real-world simulators for robotics as foundation models. They leverage diverse datasets, each rich in different aspects of real-world experience, to simulate the visual outcomes of both high-level instructions and low-level controls. The resulting simulator is used to train policies that can be deployed in real-world scenarios, showcasing the potential of bridging the sim-to-real gap in embodied learning.

World models can also be trained at a very large scale from scratch, leading to the development of multi-modal foundation world models capable of generalizing greatly in vertical domains. Building on this for efficient robot training, Wang et al. [150] present RoboGen, a generative robotic agent that automatically learns diverse skills through environment simulation. Similarly, Mazzaglia et al. [109] introduce GenRL, a framework for model-based RL training of generalist agents, that learns a multi-modal foundation world model. Moreover, Wu et al. [158] presents iVideoGPT, a scalable autoregressive transformer framework for interactive world models. It is pre-trained on millions of human and robotic manipulation trajectories, demonstrating its versatility in various downstream tasks. While Mao et al. [106] focuses on zero-shot safety prediction for autonomous robots using foundation world models. They propose a world model that combines foundation models with interpretable embeddings, addressing the distribution shift issue in standard world models. Their approach demonstrates superior state prediction and excels in safety predictions, highlighting the potential of foundation models for safety-critical applications. Lastly, Bruce et al. [20] introduces Genie, the first generative interactive environment trained in an unsupervised manner from unlabeled Internet videos.

While significant progress has been made in model-based RL and world models could unlock unlimited data availability for training, Wolczyk et al. [156] discuss the issue of catastrophic forgetting in post-training RL models, where pre-trained knowledge can be lost as new tasks are learned. This problem is particularly evident in compositional tasks, where different parts of the environment are introduced at different stages of training.

5.3 Planning & Exploration

Policy learning refers to the process of determining a set of actions that can be executed to reach the desired target state for a given task [140, 54]. In this section, we briefly highlight interesting findings in the literature where generative AI models are used to learn effective policies for RL. In particular, two important concepts are exploration, which enables the agent to discover new states and maximize rewards, and planning, which involves combining skills to develop more comprehensive policies [94, 110, 122]. Various approaches leveraging LLMs, VLMs, and diffusion models have been explored to enhance both exploration and planning in RL settings [38, 65, 175, 33, 98, 40, 91, 182, 17].

5.3.1 LLMs for planning and exploration

LLMs can be used for planning in RL, where the LLM acts as a strategic semantic planner, guiding the application of learned RL skills to new tasks based on task-specific prompts. Ahn et al. [5] and Huang et al. [72] introduce two methods that combines LLMs with pre-trained skills and affordance functions extracted from the RL training to ground language in robotic actions. The LLM is used to propose high-level actions, while the affordance functions, often learned through RL, determine the feasibility of these actions in the current context. This approach enables robots to execute complex tasks based on natural language instructions by ensuring that the proposed actions are both semantically relevant and physically feasible. Similarly, Plan-Seq-Learn (PSL) [41], is a modular approach that uses motion planning to connect abstract language from LLMs with learned low-level control for solving long-horizon robotics tasks. PSL breaks down tasks into sub-sequences, uses vision and motion planning to translate these sub-sequences into actionable steps, and then employs RL to learn the necessary low-level control strategies. This approach enables robots to efficiently learn and execute complex tasks by leveraging the strengths of both LLMs and RL.

Regarding exploration, Colas et al. [38] were the first to work on it by specifically incorporating language-driven imagination into RL. According to their IMAGINE architecture, “imagination” is the process of applying a learned goal recognizer to relabel prior experiences with other, language-based objectives. By predicting which natural language descriptions could realistically correspond to an episode’s outcome, this methodology enables the agent to associate new hypothetical objectives with previous ones. The agent may efficiently learn from these relabeled goals by combining a modular policy with a language-conditioned reward function. This improves generalization and exploration in a variety of language-specific tasks. The research underscores the critical role of language in enhancing creative, goal-driven exploration in RL. LLMs have shown promise in directly generating RL policies, instead of generating reward functions: GLAM [24], InstructRL [65], and BOSS [175], explore this, each with distinct approaches and contributions. GLAM focuses on grounding LLMs in interactive environments through online RL. It utilizes an LLM as the policy for an agent operating in a textual environment, refining the LLM’s understanding of the environment through continuous interaction and feedback. This approach aims to address the challenge of aligning the LLM’s knowledge with the actual environment dynamics, improving its ability to make decisions and achieve goals. The key innovation of GLAM lies in its online learning approach, where the LLM is not just pre-trained on existing data but actively learns and adapts as it interacts with the environment. This allows for a more dynamic and context-aware policy generation process. InstructRL, on the other hand, introduces a framework where humans provide high-level natural language instructions to guide the agent’s behavior. These instructions are used to generate a prior policy using LLMs, which then regularizes the RL objective. This approach aims to align the agent’s actions with human preferences and expectations, making it more suitable for collaborative tasks. InstructRL bridges the gap between human intentions and agent actions. The authors acknowledge that their work is limited by the challenges of abstracting certain actions, like continuous robot joint angles, into language, but they are optimistic about future advancements in multi-modal models expanding their applicability. Addressing the problem of generating realistic robotic policies and providing a more robust framework for translating abstract concepts into precise commands for robots, BOSS tackles the challenge of learning long-horizon tasks with minimal supervision. It starts with a set of primitive skills and progressively expands its skill repertoire through a bootstrapping phase. During this phase, the agent practices chaining skills together, guided by LLMs that suggest meaningful combinations. This approach enables the agent to learn complex behaviors autonomously, reducing the need for extensive human demonstrations or reward engineering. BOSS’s innovation lies in its ability to leverage the knowledge embedded in LLMs to guide the exploration and learning of new skills, making it a promising approach for developing generalist agents capable of performing a wide range of tasks.

5.3.2 VLMs for exploration

Another major theme in the research is the use of FMs that work with images together with text to guide exploration in RL agents. VLMs offer deeper grounding in robotics applications compared to LLMs because they directly associate visual perception with the input task in text form and do not require a separate vision module for perception, hence enhancing efficient exploration based on goal evaluation [98, 102, 40]. By processing real-time sensory data from the robot (e.g., camera images), the model can identify unexpected situations or deviations from the expected plan. This information can be used to modify the reward function online, penalizing actions that lead to undesirable outcomes and encouraging exploration of alternative strategies. For example, if a robot attempting to pick up a cup encounters an obstacle, the foundation model could adjust the reward function to prioritize navigating around the obstacle before resuming the grasping attempt. As demonstrated by Cui et al. [40], VLMs can enable zero-shot task specification in robotic manipulation by supporting more general and user-friendly goal representations—such as internet images or hand-drawn sketches—which promote exploration more closely aligned with the intended task. Other studies, however, assert that general-purpose VLMs might struggle with exploration and result in agents that act too roughly, especially in online RL [33].

5.3.3 Diffusion models for planning and exploration

Conventional RL planning techniques frequently use deterministic algorithms, which might perform poorly in complicated or unpredictable contexts [77, 48, 179]. Diffusion models, which were first created for generative tasks, have been modified to provide flexibility and stochasticity to the planning process in order to add robustness [91, 182, 17, 88, 61]. These models allow for flexible decision-making by iteratively converting noise into planned action sequences. One important strategy, Diffuser [76], combines RL with diffusion models to create reward-conditioned plans based on prior knowledge. It is excellent at long-term planning. To improve sample efficiency and generalization, extensions such as EDGI [17] treat planning as conditional sampling and add domain constraints. Other studies focus on safe planning, guaranteeing constraint satisfaction, and refining plans that are not feasible [160, 85]. Diffusion models also improve exploration by stochastically generating a wide range of possible states and objectives [153, 88, 174, 81, 29].

Lastly, diffusion models for planning and exploration works well in offline RL, where we can train diffusion models to generate high-quality actions from large datasets or benchmarks [53, 95, 137, 79, 65, 59, 122, 75, 185, 82]. Some methods [137, 95], use gradient-based planning or energy functions to steer action generation and improve reward outcomes. Other works [79], focus on optimizing the sampling process to make diffusion models more practical and faster. Researchers also developed techniques like Implicit Diffusion Q-Learning [59] and Q-Score Matching [122], incorporating Q-learning ideas to better connect action choices with predicted rewards training offline [65, 63, 145, 129, 75].