Uploaded November 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, David Hall from Stanford shares how the Marin project brings true open-source principles to foundation model development—closing the gap between “open-weight” releases and fully reproducible ML research.
He explains how most open-weight model releases lack key components such as training code, data recipes, and logs, making them impossible to reproduce or audit. Marin tackles this problem by ensuring that every training run begins as a GitHub pull request, where hypotheses are stated explicitly and configurations are fully pinned.
In this talk, David walks through how Ray orchestrates each job across preemptible Google Cloud TPUs, streaming metrics in real time and storing artifacts tightly linked to the exact commit that launched the run. Successes, failures—even restarts—are all publicly visible, preserving the iterative scientific process instead of hiding it.
Attendees will learn how Marin leverages Ray to make large-scale training transparent, inspectable, and reproducible—bringing foundation model development closer to the standards of open-source software.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com
At Ray Summit 2025, David Hall from Stanford shares how the Marin project brings true open-source principles to foundation model development—closing the gap between “open-weight” releases and fully reproducible ML research.
He explains how most open-weight model releases lack key components such as training code, data recipes, and logs, making them impossible to reproduce or audit. Marin tackles this problem by ensuring that every training run begins as a GitHub pull request, where hypotheses are stated explicitly and configurations are fully pinned.
In this talk, David walks through how Ray orchestrates each job across preemptible Google Cloud TPUs, streaming metrics in real time and storing artifacts tightly linked to the exact commit that launched the run. Successes, failures—even restarts—are all publicly visible, preserving the iterative scientific process instead of hiding it.
Attendees will learn how Marin leverages Ray to make large-scale training transparent, inspectable, and reproducible—bringing foundation model development closer to the standards of open-source software.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
X: https://x.com/anyscalecompute
Website: anyscale.com










