RL for VLAs

1 minute read

Published:

Today in Sergey’s VLA seminar, Seohong talked about RL for VLAs. Here are brief notes from the two lectures:

  • LLMs have cheap rollouts, thus predominantly use off-policy policy gradient algorithms such as PPO and GRPO.
  • VLA RL occupies a different regime, preferring sample-efficient value-based methods like DDPG, TD3, SAC; robots have slow rollouts. However, PPO and GRPO is still often used in simulation.
  • The problem with learning from fixed data is that the Q function is never trained on actions outside the dataset, so the policy is low quality. Nearly every offline method tries to maximizing return while staying close to the data.
  • Flow matching action heads (e.g. \(\pi_0\)) make RL awkward, since likelihoods are intractable and maximizing Q means backpropagating through the ODE. Most Gaussian RL tricks still carry over though, e.g. BC regularization, advantage weighting, rejection sampling. Need to review this section.
  • A cheap alternative is to freeze the VLA and train a small Gaussian policy with SAC on top, either steering its noise or adding residual edits to its actions.
  • On-policy RL is expensive but good (Go, Dota, reasoning LLMs). Off-policy RL is sample efficient but hard, and the bottleneck is value learning rather than policy extraction.
  • Seohong proposed some endgames: on-policy RL with a huge robot fleet, on-policy RL in a learned world model, or literally solving value learning.