4 Modality
This section focuses on the classification of five types of generative AI models used in RL, as introduced in Section 3, with an emphasis on how their input/output modalities shape their role within RL frameworks.
Bridging the gap between symbolic concepts and real-world experience is a concept referred to as grounding, introduced in Section 1. Our classification highlights a fundamental trade-off between abstraction and grounding: while LLMs [47, 2] and VLMs [8, 176, 16, 183, 33] are adept at abstract reasoning and symbolic processing, their input/output operations are less directly tied to physical control [24, 12, 37, 5, 101, 163]. On the other hand, diffusion models, by generating low-level, continuous control actions and operating directly in the action space, provide a more sample-efficient approach to policy learning [63, 129, 152, 35, 184, 76], though they may lack the higher-level abstraction capabilities inherent to larger language models. Models that can effectively fuse abstraction and grounding for low-level action generation in robotics remain limited, although world models [159, 132, 169] and video prediction models [49, 45] represent promising approaches by integrating both high-level abstraction and grounded sensory information.
The diagram in Figure 4(b) shows that, based on our analysis, each generative AI tool used in RL contributes distinct input/output modalities. LLMs, working with text as both input and output excel at symbolic reasoning and can synthesize or interpret textual inputs such as task goals, environment documentation, or learned RL primitives, enabling them to produce reward signals or task refinements aligned with high-level objectives [101, 103]. VLMs, in contrast, can process visual inputs and generate reasoning over visual scenes. This makes them valuable for visual feedback mechanisms [151]. Their ability to bridge visual and textual modalities allows flexible integration into tasks where visual context is critical. Diffusion models can be highly specialized for policy learning and state generation, as they can model complex distributions over low-level states and actions. Operating directly in continuous action spaces, they can produce precise control signals that are easily integrated with RL algorithms [35, 71]. This modality specificity makes them ideal for applications in robotics and control where fine-grained action generation is required. Lastly, world and video prediction models, with their broad support for multi-modal input fusion and internal representation learning capacity, provide an even more flexible foundation. They can incorporate and generate rich multi-modal state representations (including visual, textual, proprioceptive, and other sensory data), making them well-suited for learning predictive models of environment dynamics and for supporting planning and model-based RL [159, 132, 169].
The level of abstraction and the size of the models (in terms of parameters) strongly influence how easily they can be integrated into the RL training loop. Larger, more powerful, and generalist models (e.g., current LLMs, which are text-based) can be quickly incorporated in a zero-shot fashion in several frameworks [101, 103]. However, they are often less tailored and less adaptable to the specific input/output requirements of RL tasks. Moreover, the choice of model—whether a diffusion model, an LLM, or a VLM—directly impacts the integration strategy within RL frameworks. Different models impose distinct computational demands and offer varying levels of flexibility in adapting to real-world dynamics [61, 180, 71]. The diagram in Figure 4(b) visually summarizes these trade-offs across models, mapping representative tools based on their degree of symbolic reasoning and grounding in physical control. This complements our analysis of modality diversity and integration strategies by illustrating how models such as UniSim, EUREKA, RL-VLM-F, Reward Diffusion, and KCGG occupy different positions in this space.