AI EngineerHave you ever launched an awesome agentic demo, only to realize no amount of prompting will make it reliable enough to deploy in production? Agent reliability is a famously difficult problem to solve!
In this talk we’ll learn how to use GRPO to help your agent learn from its successes and failures and improve over time. We’ve seen dramatic results with this technique, such as an email assistant agent that whose success rate jumped from 74% to 94% after replacing o4-mini with an open source model optimized using GRPO.
We’ll share case studies as well as practical lessons learned around the types of problems this works well for and the unexpected pitfalls to avoid.
About Kyle Corbitt Kyle Corbitt is the co-founder and CEO of OpenPipe, the RL post-training company. OpenPipe has trained thousands of customer models for both enterprises and tech-forward startups.
Before founding OpenPipe, Kyle led the Startup School team at Y Combinator, which was responsible for the product and content that YC produces for early-stage companies. Prior to that he worked as an engineer at Google and studied ML at school.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletter
Timestamps:
[00:00] - Introduction to building reliable agents with RL.
[00:49] - Case Study: ART-E, an AI email assistant.
[02:19] - The importance of starting with prompted models before moving to RL.
[03:17] - Performance improvements of RL over prompted models.
[05:18] - Cost and latency benefits of the RL approach.
[08:02] - The two hardest problems in modern RL: realistic environments and reward functions.
[13:13] - Optimizing agent behavior with "extra rewards."
[15:25] - The problem of "reward hacking" and how to address it.
How to Train Your Agent: Building Reliable Agents with RL — Kyle Corbitt, OpenPipeAI Engineer2025-07-19 | Have you ever launched an awesome agentic demo, only to realize no amount of prompting will make it reliable enough to deploy in production? Agent reliability is a famously difficult problem to solve!
In this talk we’ll learn how to use GRPO to help your agent learn from its successes and failures and improve over time. We’ve seen dramatic results with this technique, such as an email assistant agent that whose success rate jumped from 74% to 94% after replacing o4-mini with an open source model optimized using GRPO.
We’ll share case studies as well as practical lessons learned around the types of problems this works well for and the unexpected pitfalls to avoid.
About Kyle Corbitt Kyle Corbitt is the co-founder and CEO of OpenPipe, the RL post-training company. OpenPipe has trained thousands of customer models for both enterprises and tech-forward startups.
Before founding OpenPipe, Kyle led the Startup School team at Y Combinator, which was responsible for the product and content that YC produces for early-stage companies. Prior to that he worked as an engineer at Google and studied ML at school.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletter
Timestamps:
[00:00] - Introduction to building reliable agents with RL.
[00:49] - Case Study: ART-E, an AI email assistant.
[02:19] - The importance of starting with prompted models before moving to RL.
[03:17] - Performance improvements of RL over prompted models.
[05:18] - Cost and latency benefits of the RL approach.
[08:02] - The two hardest problems in modern RL: realistic environments and reward functions.
[13:13] - Optimizing agent behavior with "extra rewards."
[15:25] - The problem of "reward hacking" and how to address it.
[18:37] - The solution to reward hacking:Latent Space Paper Club: AIEWF Special Edition (Test of Time, DeepSeek R1/V3) — VIbhu SapraAI Engineer2025-07-25 | Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletter
Timestamps:
00:00:00 Paper Club Year in Review & Future Plans
00:08:00 DeepSeek Paper Discussion
00:09:10 DeepSeek R1 (May 28th Update)
00:12:40 DeepSeek Distillation
00:16:51 Original DeepSeek Model Overview (DeepSeek V3 and R1)
00:21:15 Development of reasoning capabilities through a pure RL process
00:24:46 DeepSeek R10
00:39:05 DeepSeek R1 four-stage training pipeline
00:35:01 Emergence of "reflection moments" and "aha moments"
00:44:15 Distillation Strategy
00:52:34 Community and Call to ActionHuman seeded Evals — Samuel Colvin, PydanticAI Engineer2025-07-25 | In this talk I'll introduce the concept of Human-seeded Evals, explain the principle and demo them with Pydantic Logfire.
---related links---
https://x.com/samuel_colvin linkedin.com/in/samuel-colvin github.com/samuelcolvin https://pydantic.dev/Building AI Products That Actually Work — Ben Hylak (Raindrop), Sid Bendre (Oleve)AI Engineer2025-07-24 | You've made the demo. How do you make the product? A lot of AI products don't actually work. Even worse, a lot of the techniques being advertised for making AI products better don't work either. We'll cover the challenges + techniques we've seen actually work in the real world.
About Ben Hylak Ben Hylak is co-founder at Raindrop, building Sentry for AI products. He was previously a designer at Apple for 4 years, building the Apple Vision Pro.
About Sid Bendre Sid Bendre is the co-founder of Oleve, a company building a portfolio of iconic consumer software across multiple verticals. With a lean team, Oleve has already launched two virally successful consumer AI products that have amassed over 250 million views across social media platforms. One of their products reached #4 on the App Store's Education charts in 2024 and #5 in 2025, competing alongside giants like Photomath (Google) and Duolingo. Backed by Neo, Cal Henderson (co-founder of Slack), Russell Kaplan (President of Cognition), and Maria Zhang (ex-CTO of Tinder), Oleve is building the AI infrastructure to run a $1B portfolio of consumer software over the next decade. At Oleve, Sid leads technical and AI efforts, running the “Platform” team responsible for the underlying AI infrastructure that powers their lean scaling approach. Before Oleve, Sid led AI experimentation efforts at a startup hedge fund and worked at Slack, Zendesk, and Microsoft.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterRise of the AI Architect — Clay Bavor, Cofounder, Sierra w/ Alessio FanelliAI Engineer2025-07-24 | As the amount of consumer facing AI products grows, the most forward leaning enterprises have created a new role: the AI Architect. These leaders are responsible for helping define, manage, and evolve their company's AI agent experiences over time.
In this session, Clay Bavor (Cofounder of Sierra) will join Alessio Fanelli (co-host of Latent Space) in a fireside chat to share what it means to be an AI Architect, success stories from the market, and the future of the role.
---related links---
https://x.com/fanahova linkedin.com/in/fanahova alessiofanelli.comAI That Pays: Lessons from Revenue Cycle — Nathan Wan, Ensemble HealthAI Engineer2025-07-24 | While much of the AI innovation in healthcare has centered on clinical and patient-facing applications, Revenue Cycle Management (RCM) remains an underexplored yet critical domain. Given the growing financial pressures facing providers, rethinking how healthcare gets paid is essential to ensuring access and sustainability. The combination of which makes RCM an opportune area for AI disruption.
This session explores how the combination of vast structured and unstructured data, often rule-based workflows, and direct financial opportunity to drive meaningful outcomes. We’ll also share practical lessons from our journey evolving a traditional machine learning mindset to incorporate the latest advances in Generative AI, and how that shift is reshaping what's possible in healthcare operations.
---related links---
linkedin.com/in/nwan1Structuring a modern AI team — Denys Linkov, WisedocsAI Engineer2025-07-24 | You've been given an AI mandate but don't have additional headcount, what next? Re-skilling, up-skilling and team augmentation become essential to delivering on a new mandate. In this talk we'll cover strategies to structure cross functional AI teams with domain experts, software engineers and ML engineers. We'll cover key skills and milestones that each traditional role can contribute to in unique ways.
---related links---
linkedin.com/in/denyslinkovThe Rise of Open Models in the Enterprise — Amir Haghighat, BasetenAI Engineer2025-07-24 | This year kicked off with the DeepSeek-R1 news cycle breaking out of our AI Engineering bubble into the mainstream tech and business world. Leaders at the highest levels of the largest enterprises started asking how open source models could enhance and accelerate their AI strategy.
Open source models promise increased ownership of AI systems: control over performance and price, improved uptime and reliability, better compliance, and flexible hosting options. How are these promises playing out after months of implementation? In this talk, I’ll draw on hundreds of conversations with AI leaders at enterprise companies to discuss what has — and hasn’t — changed about enterprise AI strategy in a world where open-source models compete on the frontier of intelligence.
---related links---
http://twitter.com/amiruci linkedin.com/in/amirhaghighat baseten.co/blog baseten.coMentoring the Machine — Eric Hou, Augment CodeAI Engineer2025-07-24 | You’d never let a swarm of fresh interns ship to prod on day one—same deal with AI agents. Mentoring the Machine dives into how acting like a tech lead (not just a user) turns those bots into real leverage. In this talk, Eric will deliver practical advice for working with AI agents in the SDLC. He'll also preview how effective use of AI agents changes the calculus of software engineering at both a micro and macro level.
---related links---Building Applications with AI Agents — Michael Albada, MicrosoftAI Engineer2025-07-24 | Generative AI has dramatically shortened the distance between ideas and implementation, enabling faster prototyping and deployment than ever before. But while language models can streamline individual tasks, true transformation comes from combining these capabilities into intelligent, autonomous systems—AI agents.
This talk explores how to build and deploy foundation model-enabled agent systems that go beyond simple prompt chaining or chatbots. Drawing from real-world implementations and the latest research, it offers a clear and practical path to designing both single-agent and multi-agent systems capable of handling complex workflows with minimal oversight.
Attendees will gain a deeper understanding of the core design principles behind agentic systems, the architectural trade-offs involved in orchestrating multiple agents, and the strategies required to develop tailored solutions that enhance efficiency and innovation. Whether just beginning or scaling up, participants will leave with actionable insights to navigate the rapidly evolving world of AI autonomy.
09:43 - Common Pitfall 1: Insufficient Evaluation.
11:25 - Overview of specific Evaluation Tools.
12:57 - Common Pitfall 2: Lack of Observability.
13:50 - Other Common Pitfalls (Tool issues, complexity).
14:45 - The Critical Importance of Security and Safety.
15:15 - Conclusion and Future Outlook.AX is the only Experience that Matters - Ivan Burazin, DaytonaAI Engineer2025-07-24 | If you’re building devtools for humans, you’re building for the past.
Already a quarter of Y Combinator’s latest batch used AI to write 95% or more of their code. AI agents are scaling at an exponential rate and soon, they’ll outnumber human developers by orders of magnitude.
The real bottleneck isn’t intelligence. It’s tooling. Terminals, local machines, and dashboards weren’t built for agents. They make do… until they can’t.
In this talk, I’ll share how we killed the CLI at Daytona, rebuilt our infrastructure from first principles, and what it takes to build devtools that agents can actually use. Because in an agent-native future, if agents can’t use your tool, no one will.
About Ivan Burazin Ivan Burazin co-founded Codeanywhere, the very first cloud IDE, back in 2009 where he and the team had to create everything from scratch, from the IDE itself, to the entire orchestration. Concurrently, he established Shift, the premier developer conference in Europe, which was later acquired by Infobip - a global communications cloud giant in 2021. Following the acquisition, Ivan served on the executive board of this 4,000-person company and as the Chief Developer Experience Officer, where he oversaw global developer-oriented operations.
In 2023, Ivan co-founded Daytona, a fast-growing open-source platform addressing the limitations of AI coding agents by enabling them to programmatically and securely interact with runtime environments.
Backed by $7M in funding, Daytona empowers developers, from startups to Fortune 500 companies to enable AI agents to achieve their full potential.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterHow to build Enterprise Aware Agents - Chau Tran, GleanAI Engineer2025-07-24 | While LLMs demonstrated impressive reasoning capabilities, their out-of-the-box reasoning is akin to hiring a brilliant but brand-new employee who doesn’t have the enterprise context of “how things are done at this company”. In this talk, I'll introduce “Workflow Search” as a paradigm to build enterprise-aware agents that can balance predictability on common tasks, and flexibility on unforeseen tasks.
About Chau Tran Chau Tran is a Software Engineer at Glean, currently leading the technical work on Glean Assistant and semantic search. They have been with Glean for over 3 years and have a history of impactful contributions in engineering teams. Previously, Chau worked as a Research Engineer at FAIR within Meta and held technical roles at Quora. They graduated from Brown University with a Bachelor's degree in Computer Science.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterMonetizing AI — Alvaro Morales, OrbAI Engineer2025-07-23 | As AI continues to transform industries, companies are faced with the critical challenge of effectively monetizing AI-driven products in a way that captures value, ensures customer adoption, and scales revenue sustainably. Unlike traditional SaaS models, AI-powered products have unique complexities - such as fluctuating usage patterns, variable compute costs, and evolving customer demands, making conventional pricing strategies unhelpful to the growth of an AI product-led startup.
In this session, Alvaro Morales, CEO and co-founder of Orb, will explore why the often overlooked monetization aspect of AI is critical for businesses. He’ll share real-world examples and data to demonstrate how adaptive pricing models can drive cost savings, enhance customer experience, and reduce operational bottlenecks.
Alvaro will lead a live demo, showcasing how engineers can simulate AI pricing strategies and subsequently integrate them with a simple plug-and-play solution. He’ll also share how real-world revenue simulations enable companies to test and refine pricing before implementing — reducing risk, boosting adoption, and unlocking new revenue streams. As a quick example, cloud software development platform Replit was looking to adopt a usage-based pricing model for a new product, but their existing billing system couldn't support the new model, and building a new billing system would delay the launch timeline. In order to get things done, they turned to Orb, which enabled them to make pricing changes up to the last minute. After the launch, Orb became the single source of truth for both Replit and its customers - providing usage alerts to notify Replit when users hit cost thresholds and provide insights into user spend and payment methods.
Key takeaways: The challenge of AI monetization – Why traditional subscription-based SaaS pricing models don’t work for AI-powered products. Precision pricing – Exploring how usage-based, tiered, and hybrid pricing models can maximize revenue potential. Revenue simulation for AI pricing – Leveraging real-time data to test, adjust and optimize pricing strategies. Avoiding common pricing pitfalls – Identifying mistakes that can lead to revenue leakage and customer churn.
This session is designed for AI executives, product leaders, and engineering teams looking for actionable strategies to build adaptive, scalable pricing models that drive long-term growth and profitability.
---related links---
https://x.com/alvaromorales linkedin.com/in/alvaro-morales withorb.comDoes AI Actually Boost Developer Productivity? (100k Devs Study) - Yegor Denisov-Blanch, StanfordAI Engineer2025-07-23 | Forget vendor hype: Is AI actually boosting developer productivity, or just shifting bottlenecks? Stop guessing.
Our study at Stanford cuts through the noise, analyzing real-world productivity data from nearly 100,000 developers across hundreds of companies. We reveal the hard numbers: while the average productivity boost is significant (~20%), the reality is complex – some teams even see productivity decrease with AI adoption.
The crucial insights lie in why this variance occurs. Discover which company types, industries, and tech stacks achieve dramatic gains versus minimal impact (or worse). Leave with the objective, data-driven evidence needed to build a winning AI strategy tailored to your context, not just follow the trend.
About Yegor Denisov-Blanch Researcher at Stanford University researching all things developer productivity
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletter
----
The main thesis of the video is that while AI does increase developer productivity, it is not a one-size-fits-all solution. The speaker, Yegor Denisov-Blanch from Stanford, presents findings from a large-scale study on software engineering productivity to support this claim, arguing that the effectiveness of AI in software development is highly dependent on a variety of factors including task complexity, codebase maturity, language popularity, and codebase size.
timestamps:
- 00:00 Introduction and the context of AI in software development, including Mark Zuckerberg's bold claims. - 04:37 Limitations of existing studies on AI's impact on developer productivity. - 07:19 The methodology used by the Stanford research group to measure productivity. - 09:50 The overall impact of AI on developer productivity, including the concept of "rework." - 11:42 How productivity gains vary by task complexity and project maturity (Greenfield vs. Brownfield). - 14:21 The impact of programming language popularity on AI's effectiveness. - 15:42 How codebase size affects AI-driven productivity gains. - 17:22 The final conclusions of the study.How agents will unlock the $500B promise of AI - Donald Hruska, RetoolAI Engineer2025-07-23 | AI agents are on the cusp of revolutionizing work as we know it. The number of use cases software can tackle is set to explode as AI handles tasks requiring real judgment. But to cross the gap between an interesting AI prototype and an essential business tool, you need agents built by developers with real guardrails and security.
This means blending AI assistance with traditional coding in a multimodal approach that maximizes efficiency and control. The future isn't about dropping in an LLM — it requires integrating any model, any data, any system to deliver results.
Companies utilizing this approach can finally turn their slice of the $500B+ of total AI investment into real business results.
About Donald Hruska Donald is the engineering lead for Retool's new Agents product.
In his three years at Retool, Donald has led teams across AI, Mobile, and Retool's core app building product. Prior to his time at Retool, Donald co-founded and spent 5+ years growing Draftbit, a Y Combinator-backed company in the low code app building space.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterHow Intuit uses LLMs to explain taxes to millions of taxpayers - Jaspreet Singh, IntuitAI Engineer2025-07-23 | I will talk about how Intuit uses LLMs to explain tax situations to Turbotax users.
Users want explanations of their tax situations - this drives confidence in the product. Over the course of last two tax years, Intuit has built out explanations using Anthropic and openAI’s models to develop genAI powered explanations. This includes design a complex system with prompt engineered solutions and both LLM & human powered evaluations to ensure high quality bar that our users expect when filing taxes with us.
During the course of my talk, I will talk across GenAI development lifecycle at scale - including development , evaluations and scaling. And security evaluations. We also developed a fine-tuned version of Claude Haiku & shall be covering that in the presentation.
We also expanded into tax question and answering powered by RAG, including graphRAG and I would be covering those developments too.
About Jaspreet Singh I’m Jaspreet Singh, a Senior Staff Software Engineer with 12 years of experience in the tech industry. I am the tech lead for the Smart Turbotax AI team at Intuit - focusing on development of new GenAI powered experiences in Intuit Turbotax. I have worked extensively on Personalization and Recommendations problems in the past and I’m very passionate about bringing the latest in AI to help drive Taxes are done experiences for our users. I recently became a father for the first time, and enjoy spending time with my little one. As a speaker at the AI Engineer World’s Fair, I’m excited to share our journey of transforming our user’s tax filing journeys with the power of Gen AI..
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletter3 ingredients for building reliable enterprise agents - Harrison Chase, LangChain/LangGraphAI Engineer2025-07-23 | It's easy to build a prototype of an agent, but hard to put an agent in production - especially in an enterprise setting. In this section, will talk about three ingredients for building reliable agents in the enterprise.
About Harrison Chase Harrison Chase is the CEO and co-founder of LangChain, a company formed around the popular open source Python/Typescript packages. The goal of LangChain is to make it as easy as possible to use LLMs to develop context-aware reasoning applications. Prior to starting LangChain, he led the ML team at Robust Intelligence (an MLOps company focused on testing and validation of machine learning models), led the entity linking team at Kensho (a fintech startup), and studied stats and CS at Harvard.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterFrom Hype to Habit: How We’re Building an AI-First SaaS Company—While Still Shipping the RoadmapAI Engineer2025-07-23 | What does it really take to move a modern SaaS company from AI experimentation to becoming truly AI-first?
At Sprout Social, we’re in the midst of that transformation—rearchitecting strategy, systems, teams, and incentives to put AI at the heart of how we think, build, and deliver value. This is a story in motion: a behind-the-scenes look at how we’re evolving from isolated AI feature experiments to an AI-native operating model.
I’ll share what we’re learning as we navigate the innovation dilemma—integrating disruptive AI capabilities without breaking what already works or our roadmap. That includes rethinking how we define success, how we hire, reward, grow talent, and how we handle legal and ethical complexity without slowing down. We’ll explore the real-world tensions between rapid innovation, value delivery, making progress on Responsible AI, all while elevating internal AI fluency, and engaging with the broader AI ecosystem to stay at the edge.
This isn’t a playbook from the finish line—it’s a candid reflection from deep inside the journey.
About Rossella Blatt Vital Rossella Blatt Vital is a passionate AI leader with nearly 20 years of experience turning data into business value—from hands-on research and model-building to strategic executive leadership. She began her journey with a PhD in machine learning, focusing on brain-computer interfaces and cancer detection, and spent years writing code, building models, and shipping AI-powered products before stepping into leadership roles across startups, academia, and Fortune 100 companies.
As VP of AI, Data, and Data Science at Sprout Social, Rossella leads the company’s AI transformation—driving strategy across engineering, applied science, and analytics. Her team is building AI-first capabilities across product experiences, platform infrastructure, and foundational data systems.
She’s passionate about building meaningful technology—and the teams that power it—with the belief that AI, when led with vision and integrity, can help shape a more thoughtful and human-centered future.
About Deepsha Menghani Deepsha Menghani is a passionate AI leader with over a decade of experience translating data and machine learning into meaningful business impact—from predictive modeling and customer analytics to large language model applications. Her career spans hands-on data science, applied AI, and strategic leadership across global tech organizations. As Director of Engineering – AI at Sprout Social, her team is responsible for embedding AI into core product experiences and delivering insights that accelerate growth, improve customer understanding, and inform business strategy across the company.
She’s especially passionate about building AI that is not only technically robust, but also responsible, human-centered, and aligned with real-world decision-making.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterMachines of Buying and Selling Grace - Adam Behrens, New GenerationAI Engineer2025-07-23 | How to go beyond browser automation to truly agentic commerce, where AI can buy, sell and negotiate on behalf of users and merchants.
About Adam Behrens Adam Behrens is the co-founder and CEO of New Gen, a company that partners with global brands and merchants to unlock AI native commerce opportunities. New Gen builds infrastructure for brands to host their own conversational AI experiences and to connect their data into 3rd party chat clients like ChatGPT and Claude. Adam previously worked on trading infrastructure at Bridgewater and Banking-as-a-Service at Stripe. Outside of training AI models he is busy training his 8 month old Vizsla puppy.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterHow to Build Planning Agents without losing control - Yogendra Miraje, FactsetAI Engineer2025-07-23 | LLMs are getting smarter—but Agents are still unpredictable, unreliable, and hard to control.
In this talk, I’ll share practical lessons from building real-world plan-and-execute agents —covering how to steer autonomous agents using agentic workflows, blueprints, and evals.
If you’re struggling to make your agents behave (without giving up flexibility), this one’s for you.
About Yogendra Miraje I'm a backend engineer turned ML engineer turned AI engineer, with 16 years of experience building intelligent systems. I hold a Master’s degree in Computer Science from Northeastern University in Boston, and I currently work as an AI Engineer in FactSet.
I'm also the host of AI Blindspot, a podcast where we explore the frontiers of artificial intelligence—and the blind spots we often overlook in its development and deployment.
With a strong foundation in Machine Learning and software Engineering and a product-minded approach, I focus on aligning autonomous agents with real-world user goals, emphasizing safety, control, and robust evaluation techniques.
I'm passionate about building AI that’s not just powerful, but grounded, aligned, and truly useful in practice.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterBuilding Agents (the hard parts!) - Rita Kozlov, CloudflareAI Engineer2025-07-23 | AI workloads are rapidly shifting from AI being used for augmentation (co-pilots), to AI becoming responsible for full, end-to-end automation (agents). But building effective agents, and even more importantly, agent experiences that boost productivity requires many pieces. In this talk, we'll be covering the building blocks of agents, how to put them together, and what we've learned from top companies building agents along the way.
About Rita Kozlov From the very beginning, Rita has been a key figure in the development of Cloudflare's developer platform. Their initial experience as a solutions engineer, helping early enterprise customers adopt the service, gave them firsthand insight into user needs. It was the power of Cloudflare Workers that truly resonated, inspiring a vision for a serverless future. For the past eight years, they have built out the platform which now spans products including storage, compute and AI, and is used by everyone from indie millions of indie developers to Fortune 500 companies.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterPOC to PROD: Hard Lessons from 200+ Enterprise GenAI Deployments - Randall Hunt, CaylentAI Engineer2025-07-23 | The transition from experimental GenAI demonstrations to robust, production-grade systems involves significant technical and organizational complexities. Humans provide a ceiling on the true ROI of automations. This session synthesizes key patterns and practical strategies gathered from more than 200 GenAI implementations across multiple industries and business sizes.
Beyond the general lessons that apply to most products leveraging GenAI, we'll cover detailed observations within three application areas: multimodal understanding and search, enterprise knowledge retrieval, and AI agent architectures. We will share real-world comparative performance data and metrics on embedding models, vector index implementations, and explore various implementation methodologies that balance performance and cost.
Additionally, the session addresses organizational insights critical to successful AI deployments, such as the importance of clearly defined evaluation processes and understanding real-world user interaction challenges, highlighted by examples from healthcare environments. Attendees will gain an understanding of decision-making criteria, including the appropriate complexity of prompt engineering versus more elaborate orchestration methods, token/cost management strategies in multilingual settings, and the challenges in driving behavioral change with new UX and application interaction capabilities.
Participants will leave equipped with practical, data-supported insights for effectively navigating their own GenAI projects, including benchmarks and criteria for informed technology selection, and techniques to streamline the transition from initial concept to sustainable operational deployment. Please note, we all know this field evolves rapidly and we will mark which lessons we believe are immutable.
About Randall Hunt Randall Hunt is a technology leader, investor, and hands-on keyboard coder based in Los Angeles, CA. Previously, Randall led software and developer relations teams at Facebook, SpaceX, AWS, MongoDB, and NASA. Randall spends most of his time listening to customers, building demos, writing blog posts, and mentoring junior engineers. Python and C++ are his favorite programming languages, but he begrudgingly admits that Javascript rules the world. Outside of work, Randall loves to read science fiction, advise startups, travel, and ski.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterFrom Copilot to Colleague: Trustworthy Agents for High-Stakes - Joel Hron, CTO Thomson ReutersAI Engineer2025-07-23 | This keynote will explore what it takes to move from basic generative assistants to fully agentic AI—systems that don’t just suggest but plan, act, and adapt—all within the structured, high-trust environments where professionals actually work.
About Joel Hron Joel Hron is a passionate innovator driving the future of product technology and AI at Thomson Reuters. As Chief Technology Officer, he leads Product Engineering and AI Research & Development, pushing the boundaries of what’s possible in Legal, Tax, Audit, Trade, Compliance, and Risk solutions.
Joel joined Thomson Reuters in 2022 through the acquisition of ThoughtTrace, where he served as CTO. Previously, he led AI and TR Labs, launching seven groundbreaking GenAI products in just 18 months, transforming legal research, tax analysis, and contract drafting.
His approach is centered on rethinking processes through technology, building teams rooted in trust, transparency, and customer-centric innovation. Joel envisions AI not as a replacement for human expertise, but as a force that enhances professional decision-making, making expert information more accessible and impactful.
A New Orleans native, Joel’s global career spans work in London and Africa, and he now calls Zug, Switzerland home. He holds a Master’s in Mechanical Engineering from the University of Texas at Austin and a Bachelor’s in Engineering from Texas Christian University.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterHow to Hire AI Engineers when EVERYONE is cheating with AI — Beth Glenfield, DevDayAI Engineer2025-07-22 | AI broke recruitment - how to think about hiring for AI-enabled engineers in the era of AI cheating agents and AI customised resumes.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterStateful environments for vertical agents — Josh Purtell, Synth LabsAI Engineer2025-07-22 | Hey All - gave a talk on building stateful environments for vertical agents at AI tinkerers and ppl really liked it, happy to do again. Here's the repo - general code that endows environments like Pokemon Red, Minecraft, Swe-Bench, and others with the same interface for development and agent training. github.com/synth-laboratories/Environments
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterBooks reimagined: AI to create new experiences for things you know — Lukasz Gandecki, TheBrain.proAI Engineer2025-07-22 | [last round of Attendee-Led 10min lightning talks] I will showcase how I got tired of waiting for an AI assisted/no spoiler book reading experience and built my own. Check 30s video at youtu.be/JjwnYqy668M or go to demo book at https//bookgenius.net Open Sourcing!
contact: https://x.com/lgandecki
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterAI powered entomology: Lessons from millions of AI code reviews — Tomas Reimers, GraphiteAI Engineer2025-07-22 | This talk will explore insights from millions of automated code reviews, revealing trends in bugs, vulnerabilities, and code health that Graphite’s AI code review agent have uncovered. This talk will also provide meta commentary into the types of bugs AI code review agents are great at spotting, and how far the field of AI code review has come in the last year alone.
---related links---Do You Trust Your AI’s Inferences? — Sahil Yadav, Hariharan Ganesan, TelemetrakAI Engineer2025-07-22 | Enterprise AI adoption is accelerating, but with it comes a hard question: Do we trust the model’s decisions? In this 18-minute talk, I’ll explore the invisible risks behind automated decision-making in safety-critical and revenue-sensitive environments. Drawing on case studies across manufacturing, telecom, and industrial IoT, I’ll highlight how explainability, traceability, and robust guardrails drive adoption and protect enterprise value. Attendees will walk away with: • A 3-step framework for operationalizing AI trust • Real-world lessons from building guardrails in on-prem and hybrid systems • Tools and techniques for debugging and explaining inferences at scale • A blueprint for building trust between models, engineers, and executive stakeholders
---related links---
linkedin.com/in/yadavsahil telemetrak.comHow to run Evals at Scale: Thinking beyond Accuracy or Similarity — Muktesh Mishra, AdobeAI Engineer2025-07-22 | linkedin.com/in/mukteshkrmishraContinuous Profiling for GPUs — Matthias Loibl, Polar SignalsAI Engineer2025-07-22 | Continuous Profiling for GPUs extends our industry-leading continuous profiling platform to provide deep, always-on visibility into your GPU workloads.
Now you can see exactly how your GPUs are being utilized millisecond by millisecond. Our solution helps you move from guesswork to data-driven optimization.
---related links---
twitter.com/metalmatze linkedin.com/in/metalmatze matthiasloibl.com polarsignals.comTop Ten Challenges to Reach AGI — Stephen Chin, Andreas KolleggerAI Engineer2025-07-22 | an opener to the GraphRAG track!Practical GraphRAG: Making LLMs smarter with Knowledge Graphs — Michael, Jesus, and Stephen, Neo4jAI Engineer2025-07-22 | RAG has become one standard architecture component for GenAI applications to address hallucinations and integrate factual knowledge. While vector search over text is common, knowledge graphs represent a proven advancement by leveraging advanced RAG patterns to access and integrate interconnected factual information, complementing the language skills of LLMs. This talk explores GraphRAG challenges, implementation patterns, and real-world agentic examples with Google's ADK, demonstrating how this approach delivers more trustworthy and explainable GenAI solutions with enhanced reasoning capabilities.
About Michael Hunger Michael Hunger has been passionate about software development for more than 30 years.
For the last 15 years, he has been working on the open source Neo4j graph database filling many roles, most recently heading product innovation and GenAI.
As a developer Michael enjoys many aspects of software development and architecture, learning new things every day, participating in exciting and ambitious open source projects and contributing and writing software related books and articles. Michael spoke at numerous conferences and helped organize others.
Michael helps kids to learn to program by running weekly girls-only coding classes at local schools.
About Jesús Barrasa Dr. Jesús Barrasa is the AI Field CTO at Neo4j, where he works with organisations combining the power of GenAI with Knowledge Graphs. He co-authored "Building Knowledge Graphs" (O'Reilly 2023) and is cohost of the monthly Going Meta live webcast (https://goingmeta.live/) since 2022. Jesús holds a Ph.D. in Artificial Intelligence/Knowledge Representation and is an active thought leader in the KG and AI space.
About Stephen Chin Stephen Chin is VP of Developer Relations at Neo4j, conference chair of the LF AI & Data Foundation, and author of numerous titles including the upcoming GraphRAG: The Definitive Guide for O'Reilly. He has given keynotes and main stage talks at numerous conferences around the world including AI Engineer Summit, AI DevSummit, Devoxx, DevNexus, JNation, JavaOne, Shift, Joker, swampUP, and GIDS. Stephen is an avid motorcyclist who has done evangelism tours in Europe, Japan, and Brazil, interviewing developers in their natural habitat. When he is not traveling, he enjoys teaching kids how to do AI, embedded, and robot programming together with his daughters.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterKnowledge Graphs in Litigation Agents — Tom Smoker, WhyHowAI Engineer2025-07-22 | Structured Representations are pretty important in the law, where the relationships between clauses, documents, entities, and multiple parties matter. Structured Representation means Structured Context Injection. Better Context, Less Hallucinations. We walk through a couple of case studies of systems that we’ve built in production for legal use-cases - from recursive contractual clause retrieval, to HITL legal reasoning news agents.
You'll gain insights into how structured representations significantly improve the effectiveness and reliability of legal agents.
---related links---
linkedin.com/in/thomassmoker whyhow.aiWhen Vectors Break Down: Graph-Based RAG for Dense Enterprise Knowledge - Sam Julien, WriterAI Engineer2025-07-22 | Enterprise knowledge bases are filled with "dense mapping," thousands of documents where similar terms appear repeatedly, causing traditional vector retrieval to return the wrong version or irrelevant information. When our customers kept hitting this wall with their RAG systems, we knew we needed a fundamentally different approach.
In this talk, I'll share Writer's journey developing a graph-based RAG architecture that achieved 86.31% accuracy on the RobustQA benchmark while maintaining sub-second response times, significantly outperforming vector approaches.
I'll survey the key techniques behind this performance leap and why graph-based approaches excel with complex enterprise information structures like product documentation, financial documents, and technical specifications that challenge traditional RAG systems. You'll learn about using specialized LLMs to build semantic relationships, how compression techniques efficiently handle concentrated enterprise data patterns, and how infusing key data points in the memory layer of the LLM lowers hallucination.
The presentation will provide practical insights into identifying when graph-based approaches make sense for your organization's specific data challenges, helping you make informed architectural decisions for your next enterprise RAG system.
About Sam Julien Sam Julien is the Director of Developer Relations at Writer and is passionate about helping engineers improve their effectiveness and advance their careers. He loves spending time outside with his family in the Pacific Northwest. You can find more of Sam's work at samjulien.com.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterHybridRAG: A Fusion of Graph and Vector Retrieval - Mitesh Patel, NVIDIAAI Engineer2025-07-22 | Interpreting complex information from unstructured text data poses significant challenges to Large Language Models (LLM), with difficulties often arising from specialized terminology and the multifaceted relationships between entities in document architectures. Conventional Retrieval Augmented Generation (RAG) methods face limitations in capturing these nuanced interactions, leading to suboptimal performance. In our talk, we introduce a novel approach integrating Knowledge Graph-based RAG (GraphRAG) with VectorRAG, designed to refine question-answering (Q&A) systems for more effective information extraction from complex texts. Our approach employs a dual retrieval strategy that harnesses both knowledge graphs and vector databases, enabling the generation of precise and contextually appropriate answers, thereby setting a new standard for LLMs in processing sophisticated data.
About Mitesh Patel Mitesh Patel is a developer advocate manager at NVIDIA. His team is responsible for creating workflows to showcase how developers can harness GPU acceleration in their workflows using tools and frameworks popular in the developer community. Before NVIDIA, he was a senior research scientist at Fuji Xerox Palo Alto Laboratory Inc. (a research subsidiary of Fuji Xerox), where he worked on developing indoor localization technologies for applications such as asset tracking in hospitals and delivery cart tracking in manufacturing facilities. Mitesh received his Ph.D. in Robotics from the Center of Autonomous Systems (CAS) at the University of Technology Sydney, Australia in 2014.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newslettertldraw.computer - Steve Ruiz, tldrawAI Engineer2025-07-21 | Learn about tldraw's latest experiments with AI on an infinite canvas. In 2024, we created tldraw computer, a loose visual programming environment where arrows and LLMs powered every step of a graph on tldraw's canvas.
About Steve Ruiz Steve Ruiz is founder and CEO of tldraw, a London-based startup building an infinite canvas component for the web.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterExcalidraw: AI and Human Whiteboarding Partnership - Christopher ChedeauAI Engineer2025-07-21 | Covid sent everybody home and created the space of virtual whiteboards. At first the experience reused the physical constraints but soon it became better than a physical whiteboard thanks to using virtual native concepts like copy-paste and using keyboard input. The next step in this evolution is to integrate AI into the workflow. We've tried a lot of things with Excalidraw and ended up landing on turning prompt into diagram. Come to the talk to understand how it fits into the workflow and how we implemented it.
About Christopher Chedeau Co-creator of React Native and Prettier. Creator of Excalidraw, "CSS-in-JS", Yoga and React Conf.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterThe Bitter Layout or: How I Learned to Love the Model Picker — Maximillian Piras, YutoriAI Engineer2025-07-21 | Are conversational interfaces the future or, as many designers have suggested, a lazy solution that is bottlenecking AI-HCI? Despite well-documented usability issues, the design of many AI applications defaults to an input field, turn-by-turn flow, and an endless model picker — I call this “The Bitter Layout”.
In this talk, we’ll explore how Clay Christensen’s theory of commoditization from the early PC industry can explain why scaling laws require AI interfaces to remain modular until models fully commoditize. The killer feature of conversational interfaces may not be that they’re natural, but that they’re conformable. Learn how to evolve interfaces as inference scales, spot shifts in the basis of competition, and stop worrying about the next model update steamrolling your design decisions.
Foothill G 1&2: Design Engineering
About Maximillian Piras Currently the Founding Designer at Yutori working on AI web agents. Previously, Head of Design at Headliner and Sr. Designer at 8tracks. Led cross-platform UIUX design for multiple early-stage consumer startups shipping to millions of users. Contributing writer to Smashing Magazine covering AI-first design. Before that, developed graphics and animations for clients including Giphy, MIT, and Ryuichi Sakamoto.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterUX Design Principles for Semi Autonomous Multi Agent Systems — Victor Dibia, MicrosoftAI Engineer2025-07-21 | Autonomous or semi-autonomous multi-agent systems (MAS) involve exponentially complex configurations (system config, agent configs, task management and delegation, etc.). These present unique interface design challenges for both developer tooling and end-user experiences. In this session, I'll explore UX design principles for multi-agent systems, addressing critical questions: What is the true configuration space for autonomous MAS? How can users arrive at the correct mental model of an MAS's capabilities, if at all? How can we improve trust and safety through techniques like cost-aware action delegation? What makes agent actions observable? How do we enable seamless interruptibility? Attendees will gain actionable insights to create more transparent, trustworthy, and user-centered multi-agent applications, illustrated through real-world implementations in AutoGen Studio - a low code developer tool built on AutoGen (44k stars on GitHub, MIT license) and similar tools.
---related links---
https://x.com/vykthur linkedin.com/in/dibiavictor newsletter.victordibia.com victordibia.comAgentic GraphRAG: AI’s Logical Edge — Stephen Chin, Neo4jAI Engineer2025-07-21 | AI models are getting tasked to do increasingly complex and industry specific tasks where different retrieval approaches provide distinct advantages in accuracy, explainability, and cost to execute. GraphRAG retrieval models have become a powerful tool to solve domain specific problems where answers require logical reasoning and correlation that can be aided by graph relationships and proximity algorithms. We will demonstrate how an agent architecture combining RAG and GraphRAG retrieval patterns can bridge the gap in data analysis, strategic planning, and retrieval to solve complex domain specific problems.
---related links---
twitter.com/steveonjava linkedin.com/in/steveonjava http://steveonjava.com neo4j.comCIAM for AI: Authn/Authz for Agents — Michael Grinich, CEO of WorkOSAI Engineer2025-07-21 | AI agents are changing the way modern SaaS products operate. Whether automating workflows, integrating with APIs, or acting on behalf of users, AI-driven assistants and autonomous systems are becoming core product features. But securing these agents presents a fundamental challenge: How do you authenticate AI agents? How do you control what they can access? How do you ensure they act within the right permissions? This talk will explore these concepts and more while highlighting current research and best practices.
---related links---
https://x.com/grinich/ linkedin.com/in/grinich workos.com/guides workos.comGood design hasn’t changed with AI — John Pham, SF ComputeAI Engineer2025-07-21 | Bad designs are still bad. AI doesn’t make it good. The novelty of AI makes the bad things tolerable, for a short time. Building great designs and experiences with AI have the same first principles pre-AI. When people use software, they want it to feel responsive, safe, accessible and delightful. We’ll go over the big and small details that goes into software that people want to use, not forced to use.
About John Pham I'm John Pham, an engineer and a self-taught designer. I seek the dopamine hits of building delightful experiences for others. I've worked at Vercel, Microsoft and NASA doing just that.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletterBuilding Effective Voice Agents — Toki Sherbakov + Anoop Kotha, OpenAIAI Engineer2025-07-20 | How to build production voice applications and learnings from working with customers along the way!
https://x.com/tokisherbakov linkedin.com/in/akotha7What every AI engineer needs to know about GPUs — Charles Frye, ModalAI Engineer2025-07-20 | Every programmer needs to know a few things about hardware, like processors, memory, and disks. Due to AI systems' extreme demand for mathematical processing power, AI engineers need to know a few things about GPUs -- the world's most popular high-throughput mathematical co-processor.
In this talk, I will explain the fundamental engineering constraints and design decisions that shape GPUs and trace those up to some counter-intuitive facts about the performance characteristics of AI systems, with actionable insights for their deployers and consumers.
---related links---Brian Balfour: How Granola Beat Giants Like Zoom & Otter in the AI Note-Taking WarAI Engineer2025-07-20 | ...Robots as professional Chefs - Nikhil Abraham, CloudChefAI Engineer2025-07-20 | How we converted a bimanual robot into a professional chef that works in novel kitchens and learn new recipes from a single demonstration
About Nikhil Abraham Nikhil is the CEO of CloudChef - reimagining cooking using embodied AI. CloudChef builds robots that enable commercial kitchens to cook high quality meals while solving for availability of skilled chefs. Our robots are already doing full-time work in several leading commercial kitchens. Nikhil is an alum of IIT Bombay and was the cofounder of Rephrase AI (acquired by Adobe)
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletter[Full Workshop] Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel HanAI Engineer2025-07-19 | Why is Reinforcement Learning (RL) suddenly everywhere, and is it truly effective? Have LLMs hit a plateau in terms of intelligence and capabilities, or is RL the breakthrough they need?
In this workshop, we'll dive into the fundamentals of RL, what makes a good reward function, and how RL can help create agents.
We'll also talk about kernels, are they still worth your time and what you should focus on. And finally, we’ll explore how LLMs like DeepSeek-R1 can be quantized down to 1.58-bits and still perform well, along with techniques to maintain accuracy.
About Daniel Han I'm building Unsloth and we're an open-source startup trying to make AI more accessible and accurate for everyone! We have 40K GitHub stars, 10M monthly downloads on Hugging Face and worked with Google, Meta, Hugging Face teams to fix bugs in open-source models like Llama, Phi & Gemma models. I was previously working at NVIDIA making TSNE 2000x faster.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletter
TimestampsA Taxonomy for Next-gen Reasoning — Nathan Lambert, Allen Institute (AI2) & Interconnects.aiAI Engineer2025-07-19 | Current AI models are extremely skilled, which was seen as the step change in evaluation scores across the industry in the first half of 2025, but often fail when presented with even medium time-horizon tasks. This talk presents a taxonomy of 4 traits of reasoning models -- skills, calibration, strategy, and abstraction -- that will be crucial to creating the next generation of AI applications. With this, we focus on the latter two, strategy and abstraction, and discuss how these traits will enable long-horizon and reliable agents. The talk concludes with a scenario where these agentic behaviors are the foundation for RL continuing to scale in the coming years and post-training techniques reaching compute parity with pretraining methors sooner than later.
About Nathan Lambert Nathan Lambert is a Senior Research Scientist and post-training lead at the Allen Institute for AI focusing on building open language models. At the same time he founded and operates Interconnects.ai to increase transparency and understanding of current AI models and systems.
Previously, he helped build an RLHF research team at HuggingFace. He received his PhD from the University of California, Berkeley working at the intersection of machine learning and robotics. He was advised by Professor Kristofer Pister in the Berkeley Autonomous Microsystems Lab and Roberto Calandra at Meta AI Research. He was lucky to intern at Facebook AI and DeepMind during his Ph.D. Nathan was was awarded the UC Berkeley EECS Demetri Angelakos Memorial Achievement Award for Altruism for his efforts to better community norms.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletter
Timestamps:
[00:00] The Current State of Reasoning in AI Models
[01:06] Unlocking New Language Model Applications
[03:48] The Need for Advanced Planning in AI
[04:29] A Proposed Taxonomy for Next-Generation Reasoning
[06:16] Reinforcement Learning with Verifiable Rewards
[08:23] Current Challenges and Future Directions
[12:07] The Effort Required to Build New Capabilities
[16:20] A Research Plan for Training Reasoning Models
[17:36] The Shift in Compute Allocation from Pre-training to Post-trainingOpenThoughts: Data Recipes for Reasoning Models — Ryan Marten, Bespoke LabsAI Engineer2025-07-19 | Peel back the curtain on state of the art model post-training through the story of OpenThinker, a SOTA small reasoning model (outperforming DeepSeek distill), built in the open. Learn about the dataset recipe used to build the strongest reasoning models which you can apply to your own domain-specific specialized reasoning models. Hear about the strategies that scale (and that don't) based on our rigorous experimentation on the journey from thousands of data points (Bespoke-Stratos) to millions of data (OpenThinker3). Build upon our open source engineering solutions for large-scale synthetic data generation, training on multiple supercomputing clusters, and building out fast reliable evaluations.
About Ryan Marten Ryan Marten is co-lead of OpenThinker collaboration and a founding engineer at Bespoke Labs, working on data curation and model post-training. Previously, Ryan has been an AI researcher at the University of Illinois Urbana-Champaign, University of Toronto, University of Oxford, AI2, and Vector Institute. When he's not at the lab, he's probably out surfing.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletter
Timestamps:
0:00 - Introduction to the problem of open-source reasoning in AI models.
1:09 - The effectiveness of Supervised Fine-Tuning (SFT) for reasoning.
3:38 - Introduction to OpenThoughts 3 and its performance.
7:52 - Key learnings from the data recipe development.
11:34 - Guidance on adapting the dataset recipe to specific domains.
15:15 - Call for open collaboration and where to find the project's resourcesGoogle Photos Magic Editor: GenAI Under the Hood of a Billion-User App - Kelvin Ma, Google PhotosAI Engineer2025-07-19 | Go behind the scenes of Google Photos' Magic Editor. Explore the engineering feats required to integrate complex CV and cutting-edge generative AI models into a seamless mobile experience. We'll discuss optimizing massive models for latency/size, the crucial interplay with graphics rendering (OpenGL/Halide), and the practicalities of turning research concepts into polished features people actually use.
About Kelvin Ma I'm Kelvin Ma. A product engineer with 15 years of experience working across innovative consumer applications that is used by millions of consumers. I'm passionate about using technology to build tools that improves users lives by allowing greater expression, building skills, and fostering communication.
Recorded at the AI Engineer World's Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletter