6 RL Pre-Training
Until now, our survey has examined Generative AI as a tool within the RL pipeline (Generative Tools for RL). We now turn the coin over and explore the converse perspective: using RL itself to pre-train, fine-tune, and distill generative policy models. In RL Pre-Training we survey methods that pre-train transformer- or diffusion-based policies with RL objectives.
6.1 Transformer Policy
Autoregressive transformer models were originally restricted to the domain of natural language, but they now form an integral part of RL, especially when the task can be framed as a sequential decision-making problem [68, 21, 87, 157, 126]. In particular, specialized adaptations of the Transformer architecture, such as Decision Transformers (DTs) [32], are becoming more popular. Like language models, DTs generate tokens (i.e., actions) one at a time, conditioned on the previous context, and use the Transformer to predict the next action in a sequence based on previous return, state, and action tuples. Recent transformer-based RL studies focus on scalability, adaptivity, and transferability [67, 84, 154, 165]. Results from Wen et al. [154] and Xu et al. [165] show how large models may generalize across various RL tasks. These methods show how transformer backbones can learn from little supervision and generalize well. HarmoDT [68], LATTE [21], and PACT [15] would be examples of such architectures that have trained transformers at scale for RL-based robotic control on real robots. Tackling a different problem, Q-Transformer [26] leverages offline data, human demonstrations, and trajectory rollouts to learn Q-functions with transformers so that robots can perform tasks appropriately. AnyMorph [142] pushes the generalization frontier by adapting to different robot morphologies. We describe here the small number of DT variants applicable to robotics, but we envision that many RL problems may be addressed by foundation models utilizing DTs backbones [155]. Lastly, a major trend is merging language and reasoning into policies: Mezghani et al. [110] merge language generation with action prediction, allowing agents to reason in natural language while planning actions. This has not been validated directly in real-world robotics but gives an interesting approach for tasks which require long-term planning. Recently, GATO [126] emerged as a generalist agent due to its ability to execute tasks across modalities, spanning from gaming to real-world robotics, using a single transformer backbone that performs input fusion into a RL policy.
| Aspect | Transformer Autoregression | Diffusion Model |
|---|---|---|
| Generation Process | Sequential, one step at a time. | Iterative, gradual refinement from noise. |
| Output Type | Typically discrete (e.g., tokens, discrete actions). | Typically continuous (e.g., trajectories, continuous actions). |
| Training Objective | Maximize likelihood of observed data. | Predict noise added during forward diffusion. |
| Inference Speed | Faster (if same size). | Slower due to iterative refinement. |
| Error Propagation | Errors can compound over time. | Less prone to compounding errors. |
| Expressiveness | Highly expressive for sequential data. | Highly expressive for continuous data. |
| Applications | Discrete action RL policies. | Continuous action RL policies. |
| RL Fine-tuning Ease | Straightforward with standard RL algorithms (e.g., PPO, Q-learning). | More complex, may require custom integration with RL. |
| Sample Efficiency | Moderate; improves with pretraining. | High; performs well in low-data regimes. |
| Action Modeling | Limited multi-modality; can struggle in complex spaces. | Strong multi-modal capabilities; excels in complex control. |
| Best For | Long-horizon, complex reasoning tasks; fast inference scenarios. | Smooth, continuous control; planning; multi-modal outputs. |
Table 9: Comparison of generative policies. Transformer autoregressive policies versus diffusion non-autoregressive policies in RL.
6.2 Diffusion Policy
Using iterative refinement instead of step-by-step prediction, diffusion models have recently acquired popularity as alternative to autoregressive techniques in robotic policy generation [35, 76] (see Table 9 for a comparison). They are ideal for dynamic robotic environments due to their ability to manage uncertainty and provide a variety of behaviors [74, 184, 152]. Diffusion models can model complicated or high-dimensional dynamics by refining probability distributions to construct action sequences, in contrast to classic RL methods that yield deterministic outputs [139, 61]. They can incorporate constraints such as safety or energy efficiency and allow expressive, context-sensitive behaviors by directly modeling action distributions [115, 90]. Hegde et al. [62] present a method to condense a large archive of policies, trained using RL, into a single generative model. A diffusion model is then trained on these compressed representations to generate new policies conditioned on specific behaviors, either through quantitative measures or language descriptions. Other studies use diffusion models as policy in offline RL, improving performance and scalability [31, 43]. Moreover, Diffusion-QL [152] uses conditional diffusion to match IL with Q-learning while preserving demonstration data proximity and optimizing rewards. Lastly, frameworks such as Decision Diffuser [6] and DIPO [168], formulate the sequential decision-making problem as a conditional generative modeling problem using RL to train a diffuser.