9 Challenges, Recommendations and Future Research Directions
We organize challenges, recommendations and few important insights for the future into three macro-categories based on our taxonomy—Section 9.1 is related to challenges in Generative Tools for RL, Section 9.2 presents challenges in RL for Generative Policies and Section 9.3 is on Future Research Directions. Each category addresses specific limitations and opportunities that must be tackled to enable scalable, safe, and adaptive robot learning that integrates generative AI and RL.
9.1 Challenges in Generative Tools for RL
Every generative model we discussed has unique advantages for RL; consequently, different generative AI models present distinct challenges, and we highlight those we identified as most urgent to be addressed.
LLMs for reward signals
Despite their strength in symbolic reasoning, LLMs lack grounding in real-world tasks [37, 12]. Main challenges are: (i) limited contextual understanding [5, 72, 44, 39, 143, 41], (ii) real-world reduced adaptability [101, 103, 36, 96, 163] because most works are based on the assumption that the LLM has access to the environmental states from the simulation, and (iii) symbolic-physical misalignment [12, 37, 12, 24].
Recommendation: Similarly to Ma et al. [103], we believe that LLMs can improve reward generation in real-world. Develop architectures that pair LLMs with RL controllers capable of translating high-level goals into more controlled rewards. Employ LLM fine-tuning gradually increasing real-world sensor data exposure to avoid simulation environments dependency.
VLMs for reward signals and state representation
VLMs excel at perception and language integration and they are more effective in real-world tasks than LLMs, but they struggle with translating visual-semantic representations into actionable RL policies. Challenges are: (i) detailed scene understanding [40, 146, 102, 151, 167, 33, 104, 98] and (ii) inferring physical affordances or dynamics [70, 3, 136, 42, 128, 10].
Recommendation: Associate VLMs with with more meaningful information to improve precision: extracting pose of relevant objects, keypoints, or region of interest in the frames. Use positive and negative visual examples to generate semantic rewards, since it’s proved to be more effective than only positive ones by Huang et al. [70]. Use auxiliary tasks like affordance prediction to narrow the abstraction gap, and combine VLMs outputs with filtering mechanisms, as in Ahn et al. [5].
Diffusion models for planning and exploration
Diffusion models offer strong distribution modeling and short-term planning, but suffer from: (i) in-context understanding [184, 59, 122, 75, 30, 91, 182, 153, 17], (ii) multi-modality extension [137, 82, 145, 115, 61, 76], and (iii) poor reasoning [95, 79, 65, 185, 88, 160, 29].
Recommendation: Embed hierarchical control into diffusion models to manage long-horizon planning; integrate LLMs for reasoning levels with low-level diffusion-based modules. Combine diffusion models with RL through receding-horizon approaches for stable planning, as shown in recent approaches [184]. Guide diffusion models generation by conditioning them on state information or meaningful variables [6].
World models and video prediction for environment modeling
World models and video prediction models give agents the ability to “imagine” future without actually interacting with the environment, which is particularly helpful in situations with limited data. However, they have drawbacks: (i) restricted generalization [25, 57, 158, 106, 169], (ii) compounding errors over extended rollouts [144, 132, 133, 116, 172], and (iii) misalignment between actionable state representations and visual accuracy [149, 109, 20, 170, 150]. Similar to this, video prediction algorithms try to predict changes at the pixel level, but they frequently produce (i) misaligned forecasts [27] and have (ii) poor physically significant dynamics [49] over time.
Recommendation: Use foundation world models [109] rather than video predictors for more robust representation learning. Employ stochasticity and uncertainty modeling to mitigate compounding errors in prediction. Augment world models with high-fidelity sensory data [20, 149]. Hybrid frameworks combining analytical simulators with learned models (e.g., [9]) can improve realism and transferability. Fine-tune video prediction models on task-specific environments and prefer them only for data augmentation or internal reward shaping.
Scalability and resource demand of foundation models in RL training
Recent research draw attention to the difficulties of incorporating foundation models straight into the RL training loop, where RL training have (i) memory limitations [156, 157, 144] and foundation models have (ii) high inference costs [69, 162, 23] that reduce efficiency on complex tasks.
Recommendation: Give priority to modular architectures, such as Ma et al. [101]. Deploy distilled RL policies for low-latency real-world control and use foundation models mainly for offline learning and reasoning in the RL training loop. Create hybrid pipelines in which smaller controllers make decisions in real time, while LLMs or VLMs offer symbolic reasoning.
9.2 Challenges in RL for Generative Policies
RL pre-training of generative policies is widely studied and shares challenges with classic RL training as described in prior work [111, 139, 112], we first focus on RL fine-tuning, which presents new and specific open challenges unique to generative policies. While, unpredictability, safety verification, and robustness under uncertainty are inherent challenges that are common to both pre-training and fine-tuning. We highlight key open problems in these directions.
Policy agnostic RL fine-tuning
Fine-tuning of generative policies remains a critical challenge for deploying policies in real-world or dynamic environments. The main challenges are: (i) obtaining new fine-tuning RL methods and fine-tuning methods independent from the backbone model [157, 186, 34, 127], and (ii) achieving good fine-tuning for any size of the policy [66].
Recommendation: Investigate new specific methods for efficient RL fine-tuning of VLA models [66, 107], or smaller diffusion models [127]. Develop methods that are agnostic from the policy architecture, such as in Mark et al. [107].
Online RL fine-tuning of VLA models
Four major obstacles stand in the way of fine-tuning VLA models with RL: (i) data gathering is expensive in real-world, and VLA models are usually trained with real-world data through IL, limiting GPU parallelization in simulation [80, 119]. It is (ii) computationally demanding to conduct online training [66], and it is (iii) difficult to assess VLA model performance [89]. They also suffer from (iv) catastrophic forgetting [156, 157, 144].
Recommendation: Leverage realistic simulations for massive training and few-shot real-world fine-tuning to close the gap between simulation and reality. Integrate offline RL pipelines to minimize the need for online training. Use distillation to condense huge VLA models into lightweight, task-specific policies that allow for effective RL deployment [78]. Test in simulations the VLA models using frameworks like [89]. Combine replay-based continual RL with regularization to mitigate forgetting [144, 25].
Safety and failure modes of generative policies
Generative policies present serious safety and reliability issues in robotics, due to their black-box nature and absence of clear constraints. They may behave in an unpredictable manner. Challenges are: (i) closed-loop verification and (ii) real-time monitoring [68, 21, 87, 157, 126, 35, 76].
Recommendation: Incorporate safety-aware control modules that monitor and oversee the outputs of generative policies, to enforce physical constraints, and direct policy correction in the face of uncertainty [4, 130, 54]. Create closed-loop systems that use environment feedback to continually test and improve policy behavior, and implement hybrid architectures where generative policies are constrained by explicit safety layers or fallbacks [120]. Apply verification techniques such as control barrier functions [114].
9.3 Future Research Directions
We highlight three promising research directions based on our findings.
9.3.1 RL from human feedback
Techniques such as RL from Human Feedback (RLHF) and related approaches (e.g., Preference Based RL [1]) have been used for aligning large language models and other generative systems to produce outputs that are both useful and aligned with human input[186, 97]. Future research should explore RLHF to fine-tune generalist generative policies—such as VLA models. However, evaluating robotic actions remains harder than assessing tasks like chatbot responses.
9.3.2 Actor-critic foundation models
A new research direction could be to develop specialized foundation models that act as critics in RL or generate in-context rewards zero-shot in real-world settings for policy training. Traditional RL depends on manually designed rewards, which can be limiting. Instead, specific foundation models can dynamically assess actions using multi-modal inputs, improving learning efficiency and adaptability with respect to generic LLMs [101].
9.3.3 Constraint-aware generative models with optimal control
Based on recent advances in model-based training of IL policies [120, 161] and extending our discussion on safety in generative policies (Section 9), we propose that integrating constraint satisfaction methods from optimal control theory, such as control barrier functions, into RL presents a promising direction for safer and more grounded diffusion policies (as suggested in Figure 5(b)). Moreover, adaptive learning-based techniques that dynamically respond to changing environments can significantly enhance the deployment of diffusion policies across various control tasks, much like model predictive safety filters [147]. Our argument is further supported by recent works leveraging diffusion models within a receding horizon framework for control action execution [147, 178, 58, 166, 180, 123, 35], which integrate model predictive control frameworks obtaining stronger stability.