5 Top Reinforcement Learning Techniques to Train VLA Models
A robot sees a mug, hears the words "put it on the shelf," and then it moves. The loop sounds simple, but training a model to run it reliably is one of the hardest problems in robotics today.
Most vision-language-action (VLA) models learn by copying human demonstrations. Copying works until the scene changes. A new object or a shifted table is enough to break the learned pattern, and the policy starts to drift while small errors compound with every step.
Reinforcement learning (RL) offers a way out. The robot tries a task, earns a reward, and learns from its own mistakes. This guide covers five RL techniques for VLA models, and it explains how each one trains, what results it reached, and where it fits best. Every number comes from a published paper.
How a VLA Model Works
VLA architecture
Before the algorithms, you need the architecture. A VLA model has three connected parts, and every reinforcement learning method below plugs into one of them.
The vision encoder turns camera frames into visual tokens, and the language backbone reads those tokens together with the task instruction to decide what to do next. The action head then turns that decision into robot motion. Some models output discrete action tokens. Others use a diffusion or flow matching expert to produce smooth, continuous actions.
Training runs in two stages. First, supervised fine-tuning (SFT) teaches a base skill from demonstrations. Second, RL improves that skill through a continuous feedback loop in which the model acts in an environment, the environment returns a reward signal, and an optimization algorithm updates the model weights.
training pipeline
SFT has a hard limit because it only learns from data that humans collected, so it cannot learn from its own failures. RL can, and the five techniques below differ in how they close that loop, and that choice shapes the cost, stability, and speed of training.
1. PPO: The Stable Baseline
ppo pipeline
Proximal Policy Optimization (PPO) is the classic actor-critic method. The VLA acts as the actor, and a small value head acts as the critic. The critic estimates how valuable the current state is, and the gap between the observed reward and that estimate gives the advantage. PPO then updates the policy in small, clipped steps so that no single update can wreck the model.
A Tsinghua team tested this approach on OpenVLA by comparing PPO, GRPO, and DPO, and PPO won across the board. The authors point to two causes: robot tasks have shifting dynamics that hurt GRPO, and sparse rewards hurt DPO.
The team also built a lean training recipe in which a shared backbone for the actor and critic saved 45% of VRAM and trained 53% faster, a short warm-up cut environment steps by about 50%, and fewer PPO epochs saved wall-clock time without hurting sample efficiency.
RL beat SFT on semantic understanding and execution robustness, while visual robustness stayed on par, and that result matters because generalization, not peak accuracy, is what fails first in the real world.
Supporting Research Paper - Liu et al., "What Can RL Bring to VLA Generalization? An Empirical Study." arXiv:2505.19789 (NeurIPS 2025).
Best for: teams that own a simulator and a solid GPU budget.
2. GRPO: Critic-Free Training at Scale
GRPO Pipeline
Group Relative Policy Optimization (GRPO) drops the critic entirely. The model runs a group of rollouts on the same task, and a simple reward marks each one as a success or a failure. Each rollout is then scored against the rest of its group. Removing the critic frees a whole network from memory, which makes large models easier to train.
SimpleVLA-RL brought this idea to VLA models. It added dynamic sampling, a higher rollout temperature, and a wider clip range. The paper credits these exploration tricks with gains of 10 to 15%.
The numbers are strong. On OpenVLA-OFT, the LIBERO average rose from 91.0% to 99.1%, and with just one demonstration per task, LIBERO-Long jumped from 17.3% to 91.7%, which beat full-data SFT.
Note that the two studies disagree on PPO versus GRPO, so setup clearly matters and you should test both on your own task.
Supporting Research Paper - Li et al., "SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning." arXiv:2509.09674 (ICLR 2026).
Best for: settings with very few demonstrations.
3. RECAP: Advantage-Conditioned RL
RECAP Pipeline
Physical Intelligence built RECAP for real robots. It skips policy gradients, which are hard to run on billion-parameter models, and trains a value function first. That function scores how much each action helps the robot finish the task.
The score becomes a text tag, either "Advantage: positive" or "Advantage: negative," and the VLA trains with that tag as extra input. At run time, you simply ask for positive advantage, and the policy acts better than its average training data.
RECAP learns from three data types: demonstrations, autonomous rollouts, and human corrections. Then the cycle repeats as the team collects data, retrains the value function, and retrains the policy.
The results on the π*0.6 model are practical. Throughput more than doubled on some of the hardest tasks, and failure rates fell by 2x or more. The robot made espresso drinks for 13 hours straight and folded new laundry in a new home for over two hours, and in the paper it beat a policy gradient baseline.
Supporting Research Paper - Physical Intelligence, "π*0.6: a VLA That Learns From Experience." arXiv:2511.14759.
Best for: real-world fleets that keep collecting data.
4. ConRFT: Offline-to-Online Q-Learning
ConRFT Pipeline
ConRFT uses value-based RL in two stages. The offline stage trains on demonstrations with Cal-QL, a calibrated Q-learning method. The online stage keeps the same objective and adds human help. An operator steps in when the robot is about to fail, and those fixes become high-value training data. A light consistency policy serves as the action head.
The results came fast, and across eight real-world tasks the average success rate reached 96.3% after only 45 to 90 minutes of online tuning. That is a 144% gain over supervised baselines, with episodes 1.9x shorter, while HG-DAgger and PA-RL reached 65% and 71.3%.
The authors list sensitivity to reward design as a limit.
Supporting Research Paper - Chen et al., "ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy." arXiv:2502.05450 (RSS 2025).
Best for: contact-rich tasks that need quick real-world gains.
5. DSRL: Steer the Model, Do Not Retrain It
DSRL Pipeline
Diffusion Steering via Reinforcement Learning (DSRL) leaves the VLA frozen. A diffusion or flow action head starts from random noise and cleans it into an action, so DSRL trains a small policy to pick that starting noise instead. RL now works in noise space, and the frozen model turns each noise choice into an action.
The weights stay untouched, which protects the broad generalization that the model already has. The method also works with black-box access, so you can steer a model you cannot open. In the RLinf setup for π0, the trainable agent has about 500K parameters. The authors show gains on simulated tasks, real robots, and pretrained generalist policies.
Supporting Research Paper - Wagenmaker et al., "Steering Your Diffusion Policy with Latent Space Reinforcement Learning." arXiv:2506.15799 (CoRL 2025).
Best for: small compute budgets and closed models.
How to Choose the Right Technique
Start with your constraints, because they decide most of the choice. If you have a simulator and strong GPUs, try PPO or GRPO. If you have very few demonstrations, lean on GRPO. If you work on a real robot with human operators nearby, pick RECAP or ConRFT, and if you need to keep the model frozen, use DSRL.
Conclusion
VLA models no longer have to stop at imitation, because RL lets them practice, fail, and improve. PPO gives stability, GRPO gives data efficiency, RECAP scales to real deployment, ConRFT delivers fast real-world gains, and DSRL adds gains without touching the weights.
The field moves quickly, so treat each result as a starting point and match the method to your data, your hardware, and your robot. Then measure the gains yourself.
Every method above depends on clean demonstrations and reliable success labels, and Labellerr helps teams label the video and sensor data behind those pipelines.
FAQs
What are the main RL techniques used for VLA models?
The main techniques covered are PPO, GRPO, RECAP, ConRFT, and DSRL. They differ in how they handle policy updates, value estimation, demonstrations, human feedback, and whether the VLA weights are updated.
How does reinforcement learning improve VLA model generalization?
RL lets a VLA learn from its own successes and failures instead of relying only on human demonstrations. This can improve robustness when objects, environments, or task conditions differ from the training data.
Which RL technique is suitable for a frozen VLA model?
DSRL is designed to steer a frozen VLA without updating its weights. It trains a small RL policy to select the starting noise for a diffusion or flow-based action head, allowing the underlying VLA to remain unchanged.