3 Base Model
Generative models as tools for RL can be primarily classified by their Base Model—i.e., the backbone architecture. In our review, we identify five main architectures that are typically used in RL: LLMs, VLMs, diffusion models, world models, and video prediction models. Figure 4(a) visually represents these models, along with reference papers [101, 3, 71, 169] for each category and their associated input/output modalities.
The Base Model section classifies the papers according to their architecture, briefly describes their features, and summarizes key aspects in tables. These aspects are important when selecting a tool for the RL tasks defined in Section 2. The table columns represent the following:
- Paper: Lists the corresponding research papers.
- Framework: Specifies the generative AI framework used.
- Base Model/ Architecture: Specifies the generative AI architecture used.
- Network: Indicates the actual neural network model (e.g., GPT-3.5, GPT-4 for LLMs).
- Code Available: Highlights whether the implementation is accessible.
- Training Data: Lists the environments or datasets employed.
- Task: Describes the key task or feature handled by the framework (e.g., reward function design, exploration).
- Simulation or Real Exp: Distinguishes between simulated environments and real-world experiments.
Table 2 offers details on LLM-based works; Table 3 covers VLMs; Tables 4 and 5 summarize diffusion models; and Table 7 presents information on world models and video prediction models.
3.1 LLM
LLMs are changing how agents learn through RL. These models enable agents to acquire skills more effectively and flexibly by generating high-level plans, interpreting instructions, and providing structured feedback [101, 103, 98, 12, 177, 138]. LLMs are incorporated into RL systems mostly because of their capacity to leverage vast amounts of existing information [171, 23]. Because of this integrated knowledge, agents are less penalized by learning everything from scratch, which facilitates exploration and convergence on efficient behaviors. For example, GPT-4 is utilized in the EUREKA framework [101] to automatically generate code for RL tasks. Similarly, in open-ended textual settings with undefined learning objectives, LMA3 [39] uses GPT-3.5. Adaptability is another benefit. LLMs can take zero-shot decisions through prompts or updated instructions, in contrast to classic RL techniques that frequently demand for big datasets and retraining when situations change[5]. But flexibility is not without its drawbacks. Because pre-trained models’ reactions are limited by the data they were trained on, they may perform poorly in contexts that are novel or that change quickly. If not handled appropriately, this can result in biased behavior and false assumptions [51]. However, adding LLMs to RL also presents a number of difficulties. Interpretability is a major problem. It is challenging to comprehend or follow the logic behind the outputs of LLMs because they function mostly as “black boxes.” In safety-critical applications such as robotics, this lack of transparency becomes a significant concern [36, 44, 173]. Furthermore, there are real-world limitations due to the computing requirements of LLMs [47, 148, 183] and few works utilize open-source models.
Recent work combining LLMs with RL is summarized in Table 2, which covers the main frameworks, models, code availability, training data, and use cases in the RL training loop. The BOSS framework [175], for instance, makes use of GPT-3.5 to provide skill chaining across long-horizon activities, assisting agents in learning complicated behaviors through the extension and reuse of acquired abilities. FoMo [96] increases agents’ sensitivity to task context by using LLMs to dynamically scale reward signals based on visual inputs. The functions that the LLMs perform, including reward design, plan creation, or skill discovery, are represented by the Task in the table. These task modules are integrated into larger RL pipelines. More details are given in Sections 4 and 5.
Paper Framework Base Model/ Architecture Network Code Training Data Task Simulation or Real-World Exp. Chu et al. [36] Lafite-RL LLM, World Model GPT 3.5 and GPT 4 N/A Simulated robotic environment Reward Signal Simulation Colas et al. [39] Language Model Augmented Autotelic Agent (LMA3) LLM ChatGPT (gpt-3.5-turbo-0301) N/A Text-based environment (CookingWorld) Planning & Exploration Simulation Zhang et al. [175] Bootstrapping your Own Skills (BOSS) LLM GPT 3.5 Yes Alfred Dataset, Demonstration data Planning & Exploration Simulation + Real-World Ahn et al. [5] SayCan LLM GPT-3, FLAN and LAMBDA Yes Various real-world robotic tasks Planning & Exploration Real-World Ma et al. [101] EUREKA LLM GPT 4 Yes Variety of robotic tasks Reward Signal Simulation (NVIDIA Isaac Gym) Lubana et al. [96] Foundation Models as Reward Functions (FoMo Rewards) LLM N/A N/A Visual observations from agent’s trajectory Reward Signal Simulation Huang et al. [72] Grounded Decoding (GD) LLM InstructGPT and PALM N/A Various embodied tasks Planning & Exploration Simulation Carta et al. [24] Grounded Language Models (GLAM) LLM N/A Yes Textual environment (BabyAI-Text) State Representation Simulation Du et al. [44] Exploring with LLMs (ELLM) LLM GPT-3, GPT-2 and InstructGPT Yes Various environments (Crafter, Housekeep) State Representation Simulation Colas et al. [38] Intrinsic Motivations And Goal Invention for Exploration (IMAGINE) LLM N/A Yes Procedurally-generated scenes Planning & Exploration Simulation Hu and Sadigh [64] LLM Prior Policy with Regularized RL (instructRL) LLM GPT 3.5, GPT J Yes Toy game and Hanabi benchmark State Representation Simulation Yu et al. [171] LLM-based Reward Translator LLM, VLM GPT-4 Yes Simulated quadruped robot, dexterous manipulator robot tasks Reward Signal Simulation Dalal et al. [41] Plan-Seq-Learn (PSL) LLM GPT-4 Yes Challenging robotics tasks benchmarks Planning & Exploration Simulation Song et al. [135] Self-Refined Large Language Model LLM GPT-4 Available Soon Various continuous robotic control tasks Reward Signal Simulation (NVIDIA Isaac Gym) Xie et al. [163] Text2Reward Framework LLM GPT-4 Yes ManiSkill2, MetaWorld, MuJoCo Reward Signal Simulation Triantafyllidis et al. [143] Intrinsically Guided Exploration from Large Language Models (IGE-LLMs) LLM GPT-4 N/A Robotic manipulation tasks Planning & Exploration Simulation
Table 2: Summary of large language models as tools for RL.
3.2 VLM
The integration of multi-modal VLMs into RL is a significant step forward with respect to single modality LLMs, enabling more useful and informative text and visual input fusion [40, 146, 3, 102, 151, 136, 167, 42, 128, 10, 33, 104, 98]. This is particularly useful in complex robotic tasks where goal states are better represented through images rather than numbers or text. One of the key advantages of VLMs is their ability to generalize across tasks without requiring environment-specific fine-tuning, similar to LLMs [40]. Moreover, VLMs can work directly with real-world perception and don’t need to rely on accurate environment information translated into text, compared to LLM frameworks such as EUREKA [101]. The computational demand of VLMs remains significant, if not greater than that of LLMs. Few methods [146, 151] aim to mitigate this by solving RL tasks without excessive computational costs or calling the model inference at lower frequency. Moreover, fine-tuning VLMs to optimize their responses for RL specific tasks remains sometimes necessary [70]. The effectiveness of VLMs in RL also depends on the choice of architecture—e.g., several works use specialized modules such as CLIP (Contrastive Language-Image Pre-Training), a model that predicts the most relevant text sequence given an image [124, 134, 136], while others use general-purpose foundation models like GPT-4—and experimental setup, inheriting the issues of LLMs. Some studies focus on simulation environments with synthetic data [151], while others explore real-world robotic applications [98], leading to variations in prompting strategies, RL methodologies, and evaluation metrics. Notably, in Table 3 some models are used primarily for reward generation—similar to LLMs but based on visual input (e.g., Reward Generation in Venuto et al. [146])—while others use VLMs to define the goal task or condition the learning to it (classified as State Representation), as seen in Cui et al. [40]. See Section 5 for more details on column Task.

(a) Base models modularity.
(b) Trade-off between abstraction and grounding, when using generative AI tools for RL.
(a) Base models modularity.
Paper Framework Base Model/ Architecture Network Code Training Data Task Simulation or Real-World Exp. Cui et al. [40] Zero-Shot Task Specification (ZeST) VLM ImageNet-supervised ResNet50, ImageNet-trained MoCo, CLIP N/A Simulated robot manipulation tasks and real-world datasets State Representation Simulation Venuto et al. [146] VLM-CaR (Code as Reward) VLM GPT-4 N/A Discrete and continuous environments Reward Signal Simulation Adeniji et al. [3] Language Reward Modulated Pretraining (LAMP) VLM, World Model CLIP Yes Ego4D, video Dataset, automatically Generated language Dataset Reward Signal Simulation Ma et al. [102] Language-Image Value Learning (LIV) VLM CLIP with ResNet 50 backbone Yes EpicKitchen, simulated and real-world robot environments Reward Signal Simulation + Real-World Wang et al. [151] RL-VLM-F VLM Gemini and GPT-4 Vision Yes Various domains including classic control and manipulation tasks Reward Signal Simulation Sontakke et al. [136] RoboCLIP VLM S3D Yes Howto100M Reward Signal Simulation Yang et al. [167] ROBOFUME VLM, LLM MiniGPT4 Available soon Diverse datasets, real robot experiments Reward Signal Simulation + Real-World Di Palo et al. [42] Unified Agent with Foundation Models VLM, LLM CLIP, FLAN-T5 N/A Image and Text datasets State Representation Simulation Rocamonde et al. [128] Zero-Shot Vision-Language Models (VLM-RMs) VLM CLIP Yes Human motion task dataset Reward Signal Simulation Baumli et al. [10] Vision-Language Models (VLMs) like CLIP VLM CLIP N/A Playhouse Reward Signal Simulation Chen et al. [33] PR2L (Promptable Representations for RL) VLM, LLM Vicuna-7B version of the InstructBLIP, a Llama2-7B Prismatic VLM Yes Minecraft, Habitat environments State Representation Simulation Mahmoudieh et al. [104] Zero-Shot Reward Model (ZSRM) VLM ResNet-50 N/A Large dataset of captioned images Reward Signal Simulation Ma et al. [98] ExploRLLM VLM, LLM GPT-4 N/A Image-based robot manipulation tasks dataset Planning & Exploration Simulation + Real-World
Table 3: Summary of vision-language models as tools for RL.
3.3 Diffusion Model
Diffusion models, known for success in image generation [129, 50], are emerging in RL [95, 79, 137, 59, 82, 145, 65, 122, 75, 30]. Unlike models that use only textual or visual inputs, diffusion models can also operate directly in the trajectory space, generating continuous actions [185, 115]. These models learn to reverse a diffusion process that adds noise to states, rewards or policy parameters, generating new quantities from random noise. This approach has several advantages. First, it promotes behavior diversity, encouraging exploration and the discovery of novel solutions by generating a wide range of control policies. Second, it is well-suited for tasks with continuous action spaces, thereby avoiding the suboptimal performance that can arise from discretization. Finally, it improves sample efficiency by learning from limited demonstrations, reducing the need for extensive data collection [61, 184, 152]. The ability of diffusion models to be conditioned on language instructions or goal prompts enhances adaptability, making zero-shot policy generation possible. While still early in development, diffusion models offer new perspectives compared to more widely explored LLMs and VLMs [161, 91, 115]. Their iterative process allows for fine adjustments, making them ideal for complex tasks like grasping and balancing [85]. Additionally, their capacity to model complex dynamics and manage noise and uncertainty makes them robust in real-world robotic environments. However, their effectiveness in high-dimensional tasks is limited [61, 182, 92] and issues such as sensitivity to training data persist. We investigate further these aspects in Section 5 and we classify diffusion-based papers in Tables 4 and 5.
Paper Framework Base Model/ Architecture Code Training Data Task Simulation or Real-World Exp. Lu et al. [95] Contrastive Energy Prediction (CEP) Diffusion model Yes D4RL benchmarks State Representation Simulation (D4RL benchmarks) Kang et al. [79] EDP Diffusion model Yes D4RL benchmark datasets State Representation Simulation (D4RL benchmark) Suh et al. [137] Score-Guided Planning (SGP) Diffusion model Yes Cart-pole system, D4RL benchmark (MuJoCo tasks), pixel-based single integrator environment Planning & Exploration Simulation (Various simulations) Hansen-Estruch et al. [59] Implicit Diffusion Q-Learning (IDQL) Diffusion model Yes D4RL benchmark (halfcheetah, hopper, walker2d, antmaze), Maze2D State Representation Simulation (D4RL benchmark, Maze2D) Kim et al. [82] DuSkill Diffusion Model N/A Rule-based expert policies, multi-stage Meta-World tasks (slide puck, close drawer) State Representation Simulation (Multi-stage Meta-World) Venkatraman et al. [145] LDCQ Diffusion Model, World Model N/A MuJoCo benchmarks (Hopper, Walker, Ant), Maze2D, AntMaze, FrankaKitchen, CARLA State Representation Simulation (Various environments) Hu et al. [65] Temporally-Composable Diffuser (TCD) Diffusion Model N/A Gym-MuJoCo environments (HalfCheetah, Hopper, Walker2D), Maze2D, Hand Manipulation tasks State Representation Simulation (Gym-MuJoCo, Maze2D, Hand Manipulation) Psenka et al. [122] Q-Score Matching (QSM) Diffusion Model Yes DeepMind Control Suite (Cartpole Balance, Cartpole Swingup, Cheetah Run, Hopper Hop, Walker Walk, Walker Run, Quadruped Walk, Humanoid Walk) State Representation Simulation (DeepMind Control Suite) Jain and Ravanbakhsh [75] Merlin Diffusion Model N/A PointReach, PointRooms, Reacher, SawyerReach, SawyerDoor, FetchReach, FetchPush, FetchPick, FetchSlide, HandReach State Representation Simulation (Various simulated environments) Zhu et al. [185] MADIFF Diffusion Model Yes Multi-agent particle environments, MA Mujoco, StarCraft Multi-Agent Challenge (SMAC), NBA dataset State Representation Simulation (Multi-agent environments) Ni et al. [115] MetaDiffuser Diffusion Model Yes MuJoCo benchmarks (Hopper-Param, Walker-Param), Point-Robot 2D navigation State Representation Simulation (MuJoCo, Point-Robot) Chen et al. [30] SfBC Diffusion Model N/A D4RL benchmarks, AntMaze tasks, Maze2d, FrankaKitchen, Bidirectional-Car tasks State Representation Simulation (Various benchmarks)
Table 4: Summary of diffusion models as tools for RL.
Paper Framework Base Model/ Architecture Code Training Data Task Simulation or Real-World Exp. Liang et al. [91] AdaptDiffuser Diffusion Model Yes Synthetic expert data Planning & Exploration Simulation (Maze2D, MuJoCo) He et al. [61] Multi-Task Diffusion Model (MTDIFF) Diffusion model N/A Meta-World and Maze2D environments Planning & Exploration Simulation (Meta-World, Maze2D) Brehmer et al. [17] EDGI Diffusion Model, World Model N/A Offline trajectory datasets (3D navigation, Kuka robotic arm) Planning & Exploration Simulation Li et al. [88] Hierarchical Diffusion (HDMI) Diffusion model, World Model N/A Maze2D, AntMaze, D4RL environments, NeoRL benchmark Planning & Exploration Simulation (Various environments) Xiao et al. [160] SafeDiffuser Diffusion model Yes Maze2D environments, MuJoCo environments (Walker2D, Hopper), Pybullet environments Planning & Exploration Simulation (Multiple environments) Chen et al. [29] Hierarchical Diffuser Diffusion Model, World Model N/A Maze2D (U-Maze, Medium, Large), Multi2D, AntMaze, Gym-MuJoCo, FrankaKitchen Planning & Exploration Simulation (Various benchmarks) Kim et al. [81] Sub-trajectory Stitching with Diffusion (SSD) Diffusion Model Yes Maze2D environments, Fetch environments Planning & Exploration Simulation (Maze2D, Fetch) Lee et al. [85] Restoration Gap Guidance (RGG) Diffusion model Yes Maze2D environments, Gym-MuJoCo locomotion tasks, block stacking tasks with Kuka iiwa robotic arm Planning & Exploration Simulation (Multiple environments) Zhang et al. [174] Language Control Diffusion (LCD) Diffusion Model Yes CALVIN language robotics benchmark, CLEVR-Robot benchmark Planning & Exploration Simulation (CALVIN, CLEVR-Robot) Janner et al. [76] Diffuser Diffusion Model Yes Maze2D environments, block stacking tasks, D4RL locomotion suite Planning & Exploration Simulation (Multiple environments) Nuti et al. [117] Relative Reward Function Diffusion model Yes Maze2D, D4RL locomotion tasks, I2P dataset Reward Signal Simulation (Maze2D, D4RL) Mazoure et al. [108] Diffused Value Function (DVF) Diffusion Model, World Model N/A Maze2D environments, PyBullet environments, D4RL offline suite Reward Signal Simulation (Multiple environments)
Table 5: Summary of diffusion models as tools for RL (continued).
3.4 World Model and Video Prediction Model
World models in RL introduce new methods for learning representations and dynamic models that serve as internal simulators for planning and prediction in RL [159, 109, 60, 172]. Typically, a world model consists of an encoder that incorporates environmental observations, a dynamics model that is learned during pre-training and allows for internal simulation, and an optional decoder that reconstructs information from the latent space. A reward model, which forecasts rewards based on the learned representation, might also be included. Although they are not commonly employed as primary dynamic predictors, LLMs and VLMs have recently being integrated into world models. LLMs are mainly used for high-level reasoning and task specification [149, 116], while VLMs [109] support state representation by providing rich visual-linguistic embeddings. The core transition dynamics are typically modeled using architectures like RNNs, transformers, or diffusion models.
Table 7 demonstrates that State Representation is the primary use of world models, with almost all works concentrating on this function [57, 109, 158, 150]. Few studies focus on Reward Signal generation [116, 27], or Planning & Exploration [169, 170] as the main tasks. With little application to real-world data, the majority of the studied techniques are trained and assessed in simulated contexts. Though, few works [149, 133, 170, 27], include real-world experiments.
Video prediction models, special models that learn the temporal dynamics of a visual environment, have also proved good results in visual RL, with frameworks like VIPER [49]. However, the main limitation is the reliance on specific environments for training [27, 7].
See Table 6 for a comparison between features of VPMs and WMs for RL, where action-conditioned means that the model’s prediction can be influenced also by the action taken by the RL agent, and Sections 5 for more details.
| Feature | Video Prediction Models | World Models |
|---|---|---|
| Output | Future frames (pixels) | Latent states, frames, rewards |
| Use Case | Future frame generation | Future state/control generation |
| Reward Modeling | Rarely included | Often included |
| Action-Conditioned | No | Usually |
Table 6: Comparison between video prediction models and world models for RL.
Paper Framework Model Class Code Training Data Task Simulation or Real-World Exp. Ha and Schmidhuber [57] MDN-RNN World Model Yes CarRacing-v0, DoomTakeCover-v0 State Representation Simulation Wang et al. [149] GENSIM (GPT-4, GPT-3.5, Code Llama) World Model, LLM Yes Generated 100 tasks from a task library State Representation Simulation + Real-World Mazzaglia et al. [109] MFWM (GRU-based, InternVideo2) World Model, LLM, VLM Yes Walker, Cheetah, Quadruped, Stickman, Kitchen State Representation Simulation Wu et al. [158] iVideoGPT World Model Yes OXE Dataset, SSv2, RoboNet State Representation Simulation Mao et al. [106] Segment Anything Model (SAM), LLMs (GPT-3.5) World Model, LLM N/A Cart Pole, Lunar Lander State Representation Simulation Bruce et al. [20] Genie (ST-Transformer) World Model N/A Platformers dataset, robotics dataset State Representation Simulation Yang et al. [169] Observation Prediction Model (Video Diffusion Model) World Model, VLM, Diffusion Model N/A Simulated environments, real robot data, human activity videos, panorama scans Planning & Exploration Simulation Seo et al. [132] Masked World Models (MWM) World Model Yes Meta-world, RLBench State Representation Simulation Seo et al. [133] Multi-View Masked World Models (MV-MWM) World Model Yes RLBench State Representation Simulation + Real-World Ye et al. [170] Foundation Actor-Critic (FAC) World Model, Diffusion Model N/A Internet-scale robotics datasets, Meta-World Planning & Exploration Simulation + Real-World Chen et al. [27] Domain-agnostic Video Discriminator (DVD) World Model N/A Something-Something-V2, robot videos in various environments Reward Signal Simulation + Real-World Nottingham et al. [116] DECKARD World Model, LLM Yes Minecraft environment Reward Signal Simulation Zala et al. [172] EnvGen World Model, LLM Yes Generated and original environments State Representation Simulation Wang et al. [150] RoboGen Generative Simulation World Model, LLM Yes Generated tasks, scenes State Representation Simulation
Table 7: Summary of world models as tools for RL.