Meet Q-Learning with World Models (QWM)
World models have emerged as one of the biggest directions in physical AI.
At the same time, RL fine-tuning is unlocking capabilities in frontier models beyond what pretraining can achieve on its own.
Can we get the best of both worlds?
Meet Q-Learning with World Models (QWM).
QWM isn’t tied to a specific RL algorithm it’s a general framework for test-time scaling in RL.
On top of RLPD, it also shows consistent gains.
QWM scales to frontier video models.
QWM with frontier video models also consistently improves over the base RL method
World models have emerged as one of the biggest directions in physical AI.
At the same time, RL fine-tuning is unlocking capabilities in frontier models beyond what pretraining can achieve on its own.
Can we get the best of both worlds?
Meet Q-Learning with World Models (QWM).
QWM isn’t tied to a specific RL algorithm it’s a general framework for test-time scaling in RL.
On top of RLPD, it also shows consistent gains.
QWM scales to frontier video models.
QWM with frontier video models also consistently improves over the base RL method