RL for VLAs
Published:
Today in Sergey’s VLA seminar, Seohong talked about RL for VLAs. Here are brief notes from the two lectures:
- LLMs have cheap rollouts, thus predominantly use off-policy policy gradient algorithms such as PPO and GRPO.
- VLA RL occupies a different regime, preferring sample-efficient value-based methods like DDPG, TD3, SAC; robots have slow rollouts. However, PPO and GRPO is still often used in simulation.
- The problem with learning from fixed data is that the Q function is never trained on actions outside the dataset, so the policy is low quality. Nearly every offline method tries to maximizing return while staying close to the data.
- Flow matching action heads (e.g. \(\pi_0\)) make RL awkward, since likelihoods are intractable and maximizing Q means backpropagating through the ODE. Most Gaussian RL tricks still carry over though, e.g. BC regularization, advantage weighting, rejection sampling. Need to review this section.
- A cheap alternative is to freeze the VLA and train a small Gaussian policy with SAC on top, either steering its noise or adding residual edits to its actions.
- On-policy RL is expensive but good (Go, Dota, reasoning LLMs). Off-policy RL is sample efficient but hard, and the bottleneck is value learning rather than policy extraction.
- Seohong proposed some endgames: on-policy RL with a huge robot fleet, on-policy RL in a learned world model, or literally solving value learning.

