Uploaded July 2025 | Updated September 2026, 3 weeks ago
Building an evaluation from the ground up requires iteration and testing. In this video, we walk through how to use Arize Phoenix to create a benchmark dataset with annotations, then develop a custom LLM evaluator. We refine the evaluator against the golden dataset to ensure it meets quality standards, highlighting practical techniques for improving evaluator accuracy over time.
Notebook: arize.com/docs/phoenix/cookbook/evaluation/creating-a-custom-llm-evaluator-with-a-benchmark-dataset
Arize Community Slack: arize.com/community
Make a free Phoenix account: app.phoenix.arize.com
Arize Phoenix docs: arize.com/docs/phoenix
Custom Annotations Example: youtu.be/JK2JQUqpcqM?si=HHh4qPuVJMstPnp8
More about LLM as a Judge: arize.com/llm-as-a-judge
Building an evaluation from the ground up requires iteration and testing. In this video, we walk through how to use Arize Phoenix to create a benchmark dataset with annotations, then develop a custom LLM evaluator. We refine the evaluator against the golden dataset to ensure it meets quality standards, highlighting practical techniques for improving evaluator accuracy over time.
Notebook: arize.com/docs/phoenix/cookbook/evaluation/creating-a-custom-llm-evaluator-with-a-benchmark-dataset
Arize Community Slack: arize.com/community
Make a free Phoenix account: app.phoenix.arize.com
Arize Phoenix docs: arize.com/docs/phoenix
Custom Annotations Example: youtu.be/JK2JQUqpcqM?si=HHh4qPuVJMstPnp8
More about LLM as a Judge: arize.com/llm-as-a-judge










