Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Will Brown and Johannes Hagemann from Prime Intellect share how they designed and scaled the core infrastructure that powers large-scale distributed reinforcement learning across their platform.
They begin by introducing the major components of the Prime Intellect RL stack, including prime-rl, their async-first RL trainer built for massive distributed runs spanning multiple clusters, fault-tolerant execution, and heterogeneous inference pools that can leverage spot compute for rollout workers. prime-rl supports complex multi-turn environments powered by verifiers, Prime Intellect’s library for building agentic protocols around an OpenAI-compatible API—enabling direct offline evaluation using any model endpoint.
The speakers then discuss how large RL training runs—such as those for their upcoming INTELLECT-3 model—draw environment implementations from the Environments Hub, a community-driven platform for sharing train-ready RL environments packaged as importable Python modules. This hub enables modularity, rapid experimentation, and reuse across complex training pipelines.
Finally, they highlight the Prime Compute platform, Prime Intellect’s multi-cloud compute marketplace that underpins everything from large-scale training clusters and inference deployments to the secure sandboxes required for sophisticated agentic environments.
Attendees will gain an inside look at how Prime Intellect architectures distributed RL at scale, designed tooling for multi-turn agentic workflows, and built the compute substrate that supports next-generation RL-driven AI systems.
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
At Ray Summit 2025, Will Brown and Johannes Hagemann from Prime Intellect share how they designed and scaled the core infrastructure that powers large-scale distributed reinforcement learning across their platform.
They begin by introducing the major components of the Prime Intellect RL stack, including prime-rl, their async-first RL trainer built for massive distributed runs spanning multiple clusters, fault-tolerant execution, and heterogeneous inference pools that can leverage spot compute for rollout workers. prime-rl supports complex multi-turn environments powered by verifiers, Prime Intellect’s library for building agentic protocols around an OpenAI-compatible API—enabling direct offline evaluation using any model endpoint.
The speakers then discuss how large RL training runs—such as those for their upcoming INTELLECT-3 model—draw environment implementations from the Environments Hub, a community-driven platform for sharing train-ready RL environments packaged as importable Python modules. This hub enables modularity, rapid experimentation, and reuse across complex training pipelines.
Finally, they highlight the Prime Compute platform, Prime Intellect’s multi-cloud compute marketplace that underpins everything from large-scale training clusters and inference deployments to the secure sandboxes required for sophisticated agentic environments.
Attendees will gain an inside look at how Prime Intellect architectures distributed RL at scale, designed tooling for multi-turn agentic workflows, and built the compute substrate that supports next-generation RL-driven AI systems.
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com










