2 Taxonomy
In our review, we analyze how prominent generative AI tools and RL converge to enhance control policies for robots, specifically, action reference signals in the robot’s Cartesian operational space. To achieve this, we propose a new taxonomy that categorizes the 169 papers we examined, tracing the evolution of generative AI and RL integration for multi-modal data fusion in robotics policy generation. Our taxonomy (see Figure 2 for an abstract representation) primarily explores the duality between Generative AI tools for RL and RL for Generative Policies. We chose to analyze only papers focusing on specific generative models—particularly modern architectures widely adopted in recent works for policy generation, primarily based on transformers or diffusion models. We further classify the approaches along six principal dimensions based on orthogonal features of the works. For Generative Tools for RL, we categorize papers based on their underlying model architecture, which we refer to as the Base Model; the input and output modalities, referred to as Modality; and finally, the aim of the RL process, which we call the Task. For RL for Generative Policies, we analyze research along: pre-training strategies for different policy architectures, which we refer to as RL pre-training; methods for adapting policies through interaction, referred to as RL fine-tuning; and the transfer or compression of policies, which we call Policy Distillation. We chose these categories to reflect two clear and growing trends in recent literature, that we depict in Figure 3: the increasing use of generative models as tools within RL pipelines (blue outline), and the rise of RL techniques for training and refining generative policies (red outline), especially in robotic manipulation. Model architecture, modality, and task alignment are now central to foundation model choice [55], while RL techniques like fine-tuning and policy distillation are increasingly used to adapt generative policies [107]. Our taxonomy captures this duality by focusing on the most recurrent and emphasized dimensions in recent works.

Figure 2: Taxonomy. A taxonomy of the generative AI tools and RL integration for robotics.
2.1 Generative Tools for RL
Generative Tools for RL explores how selected generative AI architectures can be integrated into the RL training loop (see Figure 3(a)). We analyze prior work on leveraging generative and foundation models to enhance robotics, focusing on Transformer and Diffusion backbones. “As tools” highlights that pre-trained foundation models (like LLMs) are not being retrained end-to-end with the RL agent, but are instead leveraged in a modular way—as plug-and-play components that provide capabilities (such as understanding or generating specific modalities) that the RL agent can use during training or decision making. Their in-context learning ability allows them to flexibly adapt to tasks without needing full retraining, making them useful tools for enhancing RL agents. Our analysis is structured along three key secondary dimensions: (i) Base Model, (ii) Modality, and (iii) Task.
Base Model
Base Model classifies the models based on their underlying architecture, which directly influences their capabilities. Different architectures process and generate various types of data, shaping how they interact with RL. We analyze how deeply each base model can be incorporated into the RL framework based on its structural characteristics. Furthermore, for each model, approaches are analyzed along with (i) the framework names, (ii) code availability, (iii) the data used for training, and (iv) whether the pipeline is in simulation or the real world. Under this dimension, we consider the following models: LLMs, VLMs, Diffusion Models, World Models, and Video Prediction Models.
Modality
Under the Modality dimension, we analyze approaches based on their Input and Output modalities, such as text, images, trajectories, or low-level sensory signals. These modalities determine how generative AI integrates with RL by tackling input data interpretation and representation. Understanding these modalities helps assess the suitability of different generative AI tools for various RL applications.
Task
Generative AI models often serve as powerful priors within the RL training loop, as they are typically pre-trained modules that enhance specific aspects of learning. In this section, we delve deeper into how different works have employed generative models to address key Tasks in the RL training loop such as: Reward Signal generation, State Representation and Planning & Exploration. Figure 3(a) illustrates scenarios where various generative AI tools enhance RL tasks. As an example, LLMs support symbolic reasoning for rewards or plan generation based on task descriptions, while VLMs contribute to reward augmentation or plan generation through scene understanding.
2.2 RL for Generative Policies
The second primary dimension of our taxonomy examines RL methods used to train generative models, offering a complementary perspective to Generative AI Tools for RL. Here, we analyze works that employ RL-based approaches to pre-train, fine-tune, or distill generative policies—where RL is used directly to optimize models for action generation. We organize our discussion along three secondary dimensions, which we refer to as: (i) RL-Based Pre-Training, (ii) RL-Based Fine-Tuning, and (iii) Policy Distillation (see Figure 3(b)).
RL-Based Pre-Training
We survey various RL methods used to pre-train Transformer- and Diffusion-based policy backbones, enabling generalist generative policies that can process complex instructions; we divide them in Transformer Policy and Diffusion Policy respectively. Alternatively, RL can be used to train any set of simpler, task-specific policies for low-level control. These policies serve as a library of primitives (Primitive Generator), which can then be orchestrated by a foundation model-based planner for more complex behaviors.
RL-Based Fine-Tuning
Fine-tuning pre-trained policies is a standard approach to adapting models to new tasks or datasets. We classify fine-tuning methods for generative policies based on their architecture and policy size—we call them Policy Specific Methods. Additionally, we discuss emerging policy-agnostic fine-tuning techniques that aim to improve generalization across different generative architectures—called Policy Agnostic Methods.
Policy Distillation
Beyond policy fine-tuning, RL includes the concept of Policy Distillation, which refers to transferring knowledge and learned skills from a “teacher” policy to a “student” policy [131]. Recently, researchers have begun investigating how to distill knowledge from large pre-trained Vision-Language-Action (VLA) models into smaller and efficient RL-based expert policies (From Generalist to Expert), as well as how to use RL pre-trained single-task policies to inject more knowledge into VLA models (From Experts to Generalist).

(a) This figure illustrates generative AI models (e.g., LLMs, VLMs, diffusion models, world/video prediction models) as tools for RL-based decision making.

(b) This figure illustrates RL-based pre-training, fine-tuning and policy distillation, highlighting the complementary role of RL in enhancing generative policies.
(a) This figure illustrates generative AI models (e.g., LLMs, VLMs, diffusion models, world/video prediction models) as tools for RL-based decision making.