Uploaded July 2026 | Updated September 2026, 3 weeks ago
Alex Shaw and Ryan Marten present a rollout-centered view of evaluating and improving AI agents. Drawing on their work on Harbor, Terminal-Bench, and OpenThoughts-Agent, they connect sandboxed environments, agent evaluations, and optimization workflows into a practical framework for generating and learning from rollouts.
Speakers:
Alex Shaw — Member of Technical Staff, Laude Institute
Alex is the creator of Harbor, a framework for evaluating and optimizing agents and language models in sandboxed environments.
linkedin.com/in/alexgshaw
Ryan Marten — Member of Technical Staff, Laude Institute
Ryan builds Harbor and works on research-to-production efforts including Terminal-Bench and OpenThoughts-Agent.
linkedin.com/in/ryan-marten
Harbor: harborframework.com
GitHub: github.com/harbor-framework/harbor
Alex Shaw and Ryan Marten present a rollout-centered view of evaluating and improving AI agents. Drawing on their work on Harbor, Terminal-Bench, and OpenThoughts-Agent, they connect sandboxed environments, agent evaluations, and optimization workflows into a practical framework for generating and learning from rollouts.
Speakers:
Alex Shaw — Member of Technical Staff, Laude Institute
Alex is the creator of Harbor, a framework for evaluating and optimizing agents and language models in sandboxed environments.
linkedin.com/in/alexgshaw
Ryan Marten — Member of Technical Staff, Laude Institute
Ryan builds Harbor and works on research-to-production efforts including Terminal-Bench and OpenThoughts-Agent.
linkedin.com/in/ryan-marten
Harbor: harborframework.com
GitHub: github.com/harbor-framework/harbor






![[Full Workshop] Building Metrics that actually work — David Karam, Pi Labs (fmr Google Search)
[Full Workshop] Building Metrics that actually work — David Karam, Pi Labs (fmr Google Search) [Full Workshop] Building Metrics that actually work — David Karam, Pi Labs (fmr Google Search)](https://i.ytimg.com/vi/jxrGodnopHo/mqdefault.jpg)



