Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Hongpeng Guo from Bytedance Seed shares how the team is advancing reinforcement learning for large language models through verl, an open-source framework designed for building scalable, end-to-end RL pipelines with LLMs.
He explains why scaling RL with billion-parameter models remains difficult—existing frameworks often lack the abstractions needed to orchestrate complex dataflows or cannot efficiently manage resources across large GPU clusters. verl addresses these challenges with a Ray-based hybrid-controller architecture that provides high-level abstractions for dataflow orchestration and resource management. The entire RL workflow runs as a single controller process on the Ray driver, delegating computation to WorkerGroup and ResourcePool components that distribute work across clusters with high throughput and strong extensibility.
Hongpeng highlights how verl has gained traction in both academia and industry, integrating with major training backends (FSDP, FSDP2, Megatron-LM), inference engines (vLLM, SGLang), and supporting RL algorithms such as PPO, GRPO, and DAPO with effortless scaling.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
đź”— Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
At Ray Summit 2025, Hongpeng Guo from Bytedance Seed shares how the team is advancing reinforcement learning for large language models through verl, an open-source framework designed for building scalable, end-to-end RL pipelines with LLMs.
He explains why scaling RL with billion-parameter models remains difficult—existing frameworks often lack the abstractions needed to orchestrate complex dataflows or cannot efficiently manage resources across large GPU clusters. verl addresses these challenges with a Ray-based hybrid-controller architecture that provides high-level abstractions for dataflow orchestration and resource management. The entire RL workflow runs as a single controller process on the Ray driver, delegating computation to WorkerGroup and ResourcePool components that distribute work across clusters with high throughput and strong extensibility.
Hongpeng highlights how verl has gained traction in both academia and industry, integrating with major training backends (FSDP, FSDP2, Megatron-LM), inference engines (vLLM, SGLang), and supporting RL algorithms such as PPO, GRPO, and DAPO with effortless scaling.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
đź”— Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com










