OER·harvester

← Back to the library
arXiv HTML resource

The Duality of Generative AI and Reinforcement Learning in Robotics: A Review

Recently, generative AI and reinforcement learning (RL) have been redefining what is possible for AI agents that take information flows as input and produce intelligent behavior. As a result, we are seeing similar advancements in embodied AI and robotics for control policy generation. Our review paper examines the integration of generative AI models with RL to advance robotics. Our primary focus is on the duality be…

Licence
OPEN CC-BY-4.0
Authors
Angelo Moroncelli, Vishal Soni, Marco Forgione, Dario Piga, Blerina S…
Published
2024-10-21 · arXiv
Language
en
Length
18858 words
Type
narrative text

Cites 98 works

inferred
Open original ↗

6 RL Pre-Training

Until now, our survey has examined Generative AI as a tool within the RL pipeline (Generative Tools for RL). We now turn the coin over and explore the converse perspective: using RL itself to pre-train, fine-tune, and distill generative policy models. In RL Pre-Training we survey methods that pre-train transformer- or diffusion-based policies with RL objectives.

6.1 Transformer Policy

Autoregressive transformer models were originally restricted to the domain of natural language, but they now form an integral part of RL, especially when the task can be framed as a sequential decision-making problem [68, 21, 87, 157, 126]. In particular, specialized adaptations of the Transformer architecture, such as Decision Transformers (DTs) [32], are becoming more popular. Like language models, DTs generate tokens (i.e., actions) one at a time, conditioned on the previous context, and use the Transformer to predict the next action in a sequence based on previous return, state, and action tuples. Recent transformer-based RL studies focus on scalability, adaptivity, and transferability [67, 84, 154, 165]. Results from Wen et al. [154] and Xu et al. [165] show how large models may generalize across various RL tasks. These methods show how transformer backbones can learn from little supervision and generalize well. HarmoDT [68], LATTE [21], and PACT [15] would be examples of such architectures that have trained transformers at scale for RL-based robotic control on real robots. Tackling a different problem, Q-Transformer [26] leverages offline data, human demonstrations, and trajectory rollouts to learn Q-functions with transformers so that robots can perform tasks appropriately. AnyMorph [142] pushes the generalization frontier by adapting to different robot morphologies. We describe here the small number of DT variants applicable to robotics, but we envision that many RL problems may be addressed by foundation models utilizing DTs backbones [155]. Lastly, a major trend is merging language and reasoning into policies: Mezghani et al. [110] merge language generation with action prediction, allowing agents to reason in natural language while planning actions. This has not been validated directly in real-world robotics but gives an interesting approach for tasks which require long-term planning. Recently, GATO [126] emerged as a generalist agent due to its ability to execute tasks across modalities, spanning from gaming to real-world robotics, using a single transformer backbone that performs input fusion into a RL policy.

Aspect Transformer Autoregression Diffusion Model
Generation Process Sequential, one step at a time. Iterative, gradual refinement from noise.
Output Type Typically discrete (e.g., tokens, discrete actions). Typically continuous (e.g., trajectories, continuous actions).
Training Objective Maximize likelihood of observed data. Predict noise added during forward diffusion.
Inference Speed Faster (if same size). Slower due to iterative refinement.
Error Propagation Errors can compound over time. Less prone to compounding errors.
Expressiveness Highly expressive for sequential data. Highly expressive for continuous data.
Applications Discrete action RL policies. Continuous action RL policies.
RL Fine-tuning Ease Straightforward with standard RL algorithms (e.g., PPO, Q-learning). More complex, may require custom integration with RL.
Sample Efficiency Moderate; improves with pretraining. High; performs well in low-data regimes.
Action Modeling Limited multi-modality; can struggle in complex spaces. Strong multi-modal capabilities; excels in complex control.
Best For Long-horizon, complex reasoning tasks; fast inference scenarios. Smooth, continuous control; planning; multi-modal outputs.

Table 9: Comparison of generative policies. Transformer autoregressive policies versus diffusion non-autoregressive policies in RL.

6.2 Diffusion Policy

Using iterative refinement instead of step-by-step prediction, diffusion models have recently acquired popularity as alternative to autoregressive techniques in robotic policy generation [35, 76] (see Table 9 for a comparison). They are ideal for dynamic robotic environments due to their ability to manage uncertainty and provide a variety of behaviors [74, 184, 152]. Diffusion models can model complicated or high-dimensional dynamics by refining probability distributions to construct action sequences, in contrast to classic RL methods that yield deterministic outputs [139, 61]. They can incorporate constraints such as safety or energy efficiency and allow expressive, context-sensitive behaviors by directly modeling action distributions [115, 90]. Hegde et al. [62] present a method to condense a large archive of policies, trained using RL, into a single generative model. A diffusion model is then trained on these compressed representations to generate new policies conditioned on specific behaviors, either through quantitative measures or language descriptions. Other studies use diffusion models as policy in offline RL, improving performance and scalability [31, 43]. Moreover, Diffusion-QL [152] uses conditional diffusion to match IL with Q-learning while preserving demonstration data proximity and optimizing rewards. Lastly, frameworks such as Decision Diffuser [6] and DIPO [168], formulate the sequential decision-making problem as a conditional generative modeling problem using RL to train a diffuser.