Uploaded December 2025 | Updated September 2026, 2 weeks ago
At Ray Summit 2025, Jason Lopatecki from Arize AI shares a new paradigm for iterative model improvement—Prompt Learning (PL)—a technique inspired by reinforcement learning but performed entirely through prompts, without weight updates, large datasets, or gradient-based training.
He begins by framing the central question: What if we could achieve RL-like gains in model performance simply by refining prompts using natural-language feedback? Building on ideas from NVIDIA’s Voyager and broader community discussions, Arize explores how structured English feedback—evaluations, explanations, annotations—can act as an “error term” that guides models toward better behavior.
Jason walks through the conceptual foundation of Prompt Learning and why encoding corrections directly in language can dramatically reduce data requirements—sometimes down to only a handful of examples.
Key highlights include:
The Prompt Learning framework: How PL adapts RL-style improvement loops to the prompt layer
Why natural-language feedback is powerful: Turning textual “error signals” into prompt refinements without extra training
PL vs RL efficiency: A head-to-head comparison showing where PL outperforms traditional RL-driven refinement
Real-world experiment: Guiding JSON generation with latent constraints using only a few examples
TextPRO system: Arize’s PL engine that performs end-to-end optimization through single LLM calls—at a fraction of the supervision cost
Jason shares experimental results showing that Prompt Learning can achieve up to 70% test accuracy using just five feedback rules, outperforming standard prompting techniques and demonstrating that lightweight, language-based optimization can yield surprisingly strong results.
Attendees will leave with a new lens on iterative LLM improvement—one that opens the door to powerful, low-cost prompt optimization without full RL pipelines or large training sets.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
At Ray Summit 2025, Jason Lopatecki from Arize AI shares a new paradigm for iterative model improvement—Prompt Learning (PL)—a technique inspired by reinforcement learning but performed entirely through prompts, without weight updates, large datasets, or gradient-based training.
He begins by framing the central question: What if we could achieve RL-like gains in model performance simply by refining prompts using natural-language feedback? Building on ideas from NVIDIA’s Voyager and broader community discussions, Arize explores how structured English feedback—evaluations, explanations, annotations—can act as an “error term” that guides models toward better behavior.
Jason walks through the conceptual foundation of Prompt Learning and why encoding corrections directly in language can dramatically reduce data requirements—sometimes down to only a handful of examples.
Key highlights include:
The Prompt Learning framework: How PL adapts RL-style improvement loops to the prompt layer
Why natural-language feedback is powerful: Turning textual “error signals” into prompt refinements without extra training
PL vs RL efficiency: A head-to-head comparison showing where PL outperforms traditional RL-driven refinement
Real-world experiment: Guiding JSON generation with latent constraints using only a few examples
TextPRO system: Arize’s PL engine that performs end-to-end optimization through single LLM calls—at a fraction of the supervision cost
Jason shares experimental results showing that Prompt Learning can achieve up to 70% test accuracy using just five feedback rules, outperforming standard prompting techniques and demonstrating that lightweight, language-based optimization can yield surprisingly strong results.
Attendees will leave with a new lens on iterative LLM improvement—one that opens the door to powerful, low-cost prompt optimization without full RL pipelines or large training sets.
Liked this video? Check out other Ray Summit breakout session recordings youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh
Subscribe to our YouTube channel to stay up-to-date on the future of AI! youtube.com/c/anyscale
🔗 Connect with us:
LinkedIn: linkedin.com/company/joinanyscale
