Weights & BiasesIn this episode of Gradient Dissent, Lukas Biewald talks with Jarek Kutylowski, CEO and founder of DeepL, an AI-powered translation company. Jarek shares DeepL’s journey from launching neural machine translation in 2017 to building custom data centers and how small teams can not only take on big players like Google Translate but win.
They dive into what makes translation so difficult for AI, why high-quality translations still require human context, and how DeepL tailors models for enterprise use cases. They also discuss the evolution of speech translation, compute infrastructure, training on curated multilingual datasets, hallucinations in models, and why DeepL avoids fine-tuning for each individual customer. It’s a fascinating behind-the-scenes look at one of the most advanced real-world applications of deep learning.
Timestamps: [00:00:00] Introducing Jarek and DeepL’s mission [00:01:46] Competing with Google Translate & LLMs [00:04:14] Pretraining vs. proprietary model strategy [00:06:47] Building GPU data centers in 2017 [00:08:09] The value of curated bilingual and monolingual data [00:09:30] How DeepL measures translation quality [00:12:27] Personalization and enterprise-specific tuning [00:14:04] Why translation demand is growing [00:16:16] ROI of incremental quality gains [00:18:20] The role of human translators in the future [00:22:48] Hallucinations in translation models [00:24:05] DeepL’s work on speech translation [00:28:22] The broader impact of global communication [00:30:32] Handling smaller languages and language pairs [00:32:25] Multi-language model consolidation [00:35:28] Engineering infrastructure for large-scale inference [00:39:23] Adapting to evolving LLM landscape & enterprise needs
How DeepL Built a Translation Powerhouse with AI with CEO Jarek KutylowskiWeights & Biases2025-07-08 | In this episode of Gradient Dissent, Lukas Biewald talks with Jarek Kutylowski, CEO and founder of DeepL, an AI-powered translation company. Jarek shares DeepL’s journey from launching neural machine translation in 2017 to building custom data centers and how small teams can not only take on big players like Google Translate but win.
They dive into what makes translation so difficult for AI, why high-quality translations still require human context, and how DeepL tailors models for enterprise use cases. They also discuss the evolution of speech translation, compute infrastructure, training on curated multilingual datasets, hallucinations in models, and why DeepL avoids fine-tuning for each individual customer. It’s a fascinating behind-the-scenes look at one of the most advanced real-world applications of deep learning.
Timestamps: [00:00:00] Introducing Jarek and DeepL’s mission [00:01:46] Competing with Google Translate & LLMs [00:04:14] Pretraining vs. proprietary model strategy [00:06:47] Building GPU data centers in 2017 [00:08:09] The value of curated bilingual and monolingual data [00:09:30] How DeepL measures translation quality [00:12:27] Personalization and enterprise-specific tuning [00:14:04] Why translation demand is growing [00:16:16] ROI of incremental quality gains [00:18:20] The role of human translators in the future [00:22:48] Hallucinations in translation models [00:24:05] DeepL’s work on speech translation [00:28:22] The broader impact of global communication [00:30:32] Handling smaller languages and language pairs [00:32:25] Multi-language model consolidation [00:35:28] Engineering infrastructure for large-scale inference [00:39:23] Adapting to evolving LLM landscape & enterprise needs
Follow Weights & Biases: twitter.com/weights_biases linkedin.com/company/wandbWhy LLM benchmarks might be misleadingWeights & Biases2025-09-22 | Edwin Chen shares his perspective on how popular leaderboards like LMSYS may unintentionally reward surface-level outputs.Structured pruning with Weights & Biases | Model optimization made simpleWeights & Biases2025-09-18 | We dive into one of the most powerful techniques in model optimization: *pruning*. By zeroing out or removing low-magnitude weights, pruning reduces model size and improves efficiency. But it doesn’t stop there—*structured pruning* takes things further by eliminating entire rows, columns, or channels, making it far more hardware-friendly. This video explores not only how pruning works but also how tools like *Weights & Biases* completely transform the way we track and manage experiments.
#shortsWhy synthetic data isn’t enoughWeights & Biases2025-09-18 | A lab trained on 20 million synthetic math problems—only to watch model performance drop everywhere else. Surge AI’s Edwin Chen explains why even huge amounts of synthetic data can’t compete with a small set of high-quality human examples.Production monitoring for AI applications using W&B WeaveWeights & Biases2025-09-17 | Discover how to evaluate and monitor your AI applications in production using W&B Weave. In this video, we walk you through creating online evaluations, letting you catch failures instantly and maintain application quality over time. See how Weave’s simple interface lets technical and non-technical team members build and use monitors.
⏳Timestamps: 0:00 Monitoring is critical for productionizing AI 2:39 Monitors versus guardrails 4:05 Building Monitors in Weave 7:00 A support agent in action using online evaluations 7:27 Viewing online evaluations results in Weave 5:39 Our improved, friendlier support agent in action 9:25 Monitor trends over time using Weave Trace Plots 9:46 Conclusion: continue to track important metrics after deployment using WeaveWhy you can’t just scale with peopleWeights & Biases2025-09-16 | Surge AI’s founder explains why managing quality at scale takes more than headcount.The Startup Powering The Data Behind AGIWeights & Biases2025-09-16 | In this episode of Gradient Dissent, Lukas Biewald talks with the CEO & founder of Surge AI, the billion-dollar company quietly powering the next generation of frontier LLMs. They discuss Surge's origin story, why traditional data labeling is broken, and how their research-focused approach is reshaping how models are trained.
You’ll hear why inter-annotator agreement fails in high-complexity tasks like poetry and math, why synthetic data is often overrated, and how Surge builds rich RL environments to stress-test agentic reasoning. They also go deep on what kinds of data will be critical to future progress in AI—from scientific discovery to multimodal reasoning and personalized alignment.
It’s a rare, behind-the-scenes look into the world of high-quality data generation at scale—straight from the team most frontier labs trust to get it right.
00:00 – Intro: Who is Edwin Chen? 03:40 – The problem with early data labeling systems 06:20 – Search ranking, clickbait, and product principles 10:05 – Why Surge focused on high-skill, high-quality labeling 13:50 – From Craigslist workers to a billion-dollar business 16:40 – Scaling without funding and avoiding Silicon Valley status games 21:15 – Why most human data platforms lack real tech 25:05 – Detecting cheaters, liars, and low-quality labelers 28:30 – Why inter-annotator agreement is a flawed metric 32:15 – What makes a great poem? Not checkboxes 36:40 – Measuring subjective quality rigorously 40:00 – What types of data are becoming more important 44:15 – Scientific collaboration and frontier research data 47:00 – Multimodal data, Argentinian coding, and hyper-specificity 50:10 – What's wrong with LMSYS and benchmark hacking 53:20 – Personalization and taste in model behavior 56:00 – Synthetic data vs. high-quality human data
Follow Weights & Biases: twitter.com/weights_biases linkedin.com/company/wandbAccelerating ML workflows: IBM x Weights & BiasesWeights & Biases2025-09-16 | Training a state-of-the-art foundation model like IBM Granite is no small feat, and requires an intricate coordination of AI researchers working throughout a complex workflow. In this video, IBM leaders share how partnering with Weights & Biases enabled their teams to accelerate their ML workflows and improve the quality of their models -- all while improving their own quality of lives and workHow Woven by Toyota builds video AI agents for automated driving with W&B WeaveWeights & Biases2025-08-13 | Using W&B Weave for all video inputs and experiments tracking led to massive gains in productivity, accuracy, and also sanity and peace of mind for the Woven by Toyota team.
#shorts #woven #toyotaAI will give everyone what only CEOs have todayWeights & Biases2025-08-12 | What if everyone had the support of a CEO? AI is making it possible—giving every person their own dream team.
Full episode: youtu.be/lYz5MQvK3wUFully Connected London 2025: Agentic AI applications from prototype to productionWeights & Biases2025-08-11 | Join us in London on Nov 4–5 for Fully Connected 2025, hosted by @Weights & Biases at Convene Sancroft, St. Paul’s.
Day 1: Hands-on workshops, safe LLMs, multi-agent orchestration, and fine-tuning you can deploy the same day.
Day 2: AI Pioneer Series leadership conversations from teams building at scale. Speakers TBA. You’ll leave with code, playbooks, and new allies for getting AI to production.
🎟️ Tickets: wandb.me/FCLondonLIEveryone told me enterprise search was a graveyardWeights & Biases2025-08-07 | Arvind Jain - CEO & Founder of Glean
Full episode: youtu.be/lYz5MQvK3wUStreamline evaluation, monitoring, optimization of AI data flywheel with NVIDIA and Weights & BiasesWeights & Biases2025-08-07 | Deploying AI agents at scale introduces significant challenges, including high compute costs and latency bottlenecks—especially in performance-critical environments. Balancing model accuracy with efficiency often requires complex workflows and ongoing manual intervention.
The NVIDIA Data Flywheel Blueprint provides a systematic, automated solution to refine and redeploy optimized models that maintain accuracy targets while lowering resource demands. This blueprint establishes a self-reinforcing data flywheel, using production traffic logs and institutional knowledge to continuously improve model efficiency and accuracy.
Weights & Biases enhances the NVIDIA AI Blueprint for building data flywheels by providing advanced traceability for continuous model optimization with real-world data and user feedback. While the blueprint orchestrates automated evaluation and selection of models for optimal latency, cost, and enterprise-grade accuracy, Weights & Biases layers in native traceability, robust experiment tracking, version management, and visualization. The framework enables monitoring, automated retraining, and rapid iteration cycles, accelerating the transition from experimentation to production.
Chapters: 0:00 Introduction 1:01 What is an AI Agent? 2:33 Common Failures in AI Agents 3:25 The Role of Fine-Tuning 5:32 Creating Fine-Tuning Datasets 7:13 Data Filtering & Annotation 9:07 Fine-Tuning Execution 11:14 Comparative Evaluation 13:04 Complete Fine-Tuning Pipeline Recap 16:46 Final Takeaways & ClosingWhy enterprise search was brokenWeights & Biases2025-08-05 | Arvind Jain explains why enterprise teams spend 30% of their time just looking for information and how Glean fixes it.
Full episode: youtu.be/lYz5MQvK3wUArvind Jain on building Glean and the future of enterprise AIWeights & Biases2025-08-05 | In this episode of Gradient Dissent, Lukas Biewald sits down with Arvind Jain, CEO and founder of Glean. They discuss Glean's evolution from solving enterprise search to building agentic AI tools that understand internal knowledge and workflows. Arvind shares how his early use of transformer models in 2019 laid the foundation for Glean’s success, well before the term "generative AI" was mainstream.
They explore the technical and organizational challenges behind enterprise LLMs—including security, hallucination suppression—and when it makes sense to fine-tune models. Arvind also reflects on his previous startup Rubrik and explains how Glean’s AI platform aims to reshape how teams operate, from personalized agents to ever-fresh internal documentation.
Timestamps: 0:00 Intro 01:00 What Glean is and how it works 02:39 Starting Glean before the LLM boom 04:10 Using transformers early in enterprise search 06:48 Semantic search vs. generative answers 08:13 When to fine-tune vs. use out-of-box models 12:38 The value of small, purpose-trained models 13:04 Enterprise security and embedding risks 16:31 Lessons from Rubrik and starting Glean 19:31 The contrarian bet on enterprise search 22:57 Culture and lessons learned from Google 25:13 Everyone will have their own AI-powered "team" 28:43 Using AI to keep documentation evergreen 31:22 AI-generated churn and risk analysis 33:55 Measuring model improvement with golden sets 36:05 Suppressing hallucinations with citations 39:22 Agents that can ping humans for help 40:41 AI as a force multiplier, not a replacement 42:26 The enduring value of hard workCoreWeave infrastructure observability in W&B ModelsWeights & Biases2025-08-04 | Discover how to bring CoreWeave’s deep infrastructure observability directly into your W&B training workflows. When training on CoreWeave, W&B Models automatically captures key events from CoreWeave’s Mission Control, giving you powerful insights and expert-backed remediation hints to debug your training and fine-tuning runs faster.
⏳Timestamps: 0:00 W&B Models overview 0:56 Introducing CoreWeave observability in W&B Models 2:05 Analyzing LLM training runs in W&B Models 2:59 W&B Models without CoreWeave observability 3:23 W&B Models enhancements with CoreWeave observability 4:29 CoreWeave event annotations overlaid on training run metrics charts 5:39 Direct links from W&B Models to CoreWeave’s observability platform 6:25 Conclusion: build better models faster with CoreWeave and Weights & BiasesEvaluating AI applications using W&B WeaveWeights & Biases2025-08-01 | Why evaluating AI applications with Weave is a total game-changer for developers and builders in the AI space.Here we take a deep dive into how W&B Weave streamlines performance evaluation for AI applications—capturing every data point, visualizing metrics like latency, accuracy, and cost, and enabling developers to iterate confidently.
#shorts #shortW&B Inference: test open-source LLMs in SECONDSWeights & Biases2025-07-31 | W&B Inference lets you test open-source LLMs in SECONDS with zero setup, using the familiar OpenAI API format via the Weave SDK. In this walkthrough, we demo how users can seamlessly test popular models like GPT-4.1-mini, LLaMA 3.1.8B, and DeepSeek—all hosted on CoreWeave. See how LLM calls are visualized, debugged, and compared within Weave's Traces page and Playground. Learn how to switch models with just a dropdown, run side-by-side evaluations, and analyze accuracy, latency, and cost. With built-in scoring, datasets, and real-time streaming outputs, W&B Inference simplifies open-source LLM experimentation for developers, researchers, and teams. If you're looking to scale LLM testing without infrastructure headaches, this is your new secret weapon. #shorts #shortProtect your AI applications from risk and uncertainty with W&B Weave GuardrailsWeights & Biases2025-07-31 | AI agents powered by large language models (LLMs) are revolutionizing support, productivity, and customer engagement—but without proper **guardrails**, things can go sideways fast. In this eye-opening demo, we shows how **W&B Weave Guardrails** bring stability, safety, and control to AI-powered workflows.
#shorts #shortBuild agentic AI applications with W&B Weave: a financial research agent in action!Weights & Biases2025-07-31 | - How to create **agentic workflows** for financial research - Using the **Financial Research Agent** to analyze NASDAQ composite trends - Smart querying with the **Planner Agent** that autonomously generates research directions - Content summarization with the **Financial Writer Agent**, supported by domain experts like Fundamentals and Risk Analysts - Verification workflows that ensure trustworthy outputs with our **Verification Agent** - Navigating the **trace dashboard**, including total cost, latency, and token analysis - Exploring the **code composition view** and **flame view** to understand and debug agent behaviors
#shorts #shortUnified stack of Kubernetes, Ray, PyTorch, and vLLM.Weights & Biases2025-07-24 | AI workloads are growing fast and so are the operational challenges.
At Fully Connected 2025, Robert Nishihara, Co-founder of Anyscale, shared a clear path forward: a unified stack of Kubernetes, Ray, PyTorch, and vLLM.
His talk walks through live demos featuring autoscaling GPU pools and sub-10ms generation latency, showing what’s possible when these tools work together.
The full session is now available on demand. Worth a watch: https://lnkd.in/gUZHC8hJ
#shortsBuilding agentic AI workflows with W&B Weave: a hiring assistant case studyWeights & Biases2025-07-10 | Explore how to construct, evaluate, and monitor agentic AI workflows using W&B Weave in this comprehensive demo. Karan and Nico guide you through the development of a Hiring Assistant that assesses candidate applications for interview suitability.
-Advanced tracing of agentic workflow with hallucination guardrail - self-reflection or human-in-the-loop interaction for hallucinated decision reasoning -In-depth evaluation on quantitative and qualitative scores based on deterministic scorers and LLM judges -First part of a series on building an E2E Hiring Agent with consideration of the EU AI Act (incl. Fine-tuning of open-source models). Explore the public project, full code, and *EU AI Act whitepaper here:* wandb.ai/wandb-smle/hiring-agent-demo-public/reports/Hiring-Agent-E2E-AI-Evaluation-for-Fair-Auditable-Hiring-Decisions--VmlldzoxMjI0MjI0Mw
⏳Timestamps: 0:00 Introduction to Hiring Agent AI system 2:40 Demo: Hiring Agent prototype in action 4:47 Tracing AI Application with W&B Weave 14:15 Human-in-the-Loop: Expert Annotator View 17:05 Debugging AI Agents with the Weave Playground 24:31 Evaluate AI Agents in Weave 25:45 Quantitative Benchmark: Diving into Evaluation Results 29:15 Qualitative Drill-down: Reviewing Model Outputs 43:33 Conclusion and invitation to Explore W&B WeaveReal-time speech translation is changing everythingWeights & Biases2025-07-10 | DeepL CEO and founder Jarek Kutylowski explains how real-time speech translation is transforming international business. No more waiting on interpreters—just direct, immersive conversations.
Full episode: youtu.be/-ikvSn6xB1I?si=eCDQ5hnSY7EAB9k-Will AI replace human translators?Weights & Biases2025-07-08 | DeepL CEO & Founder Jarek Kutylowski explains why AI will take over the boring parts of translation—but humans are still essential for high-stakes work. From legal contracts to pharmaceutical documents, human oversight isn’t going anywhere.
#shortsLLMOps for eval-driven development at scaleWeights & Biases2025-07-08 | Mercari has invested heavily in DevOps, MLOps, and, recently, LLMOps with significant payoff to developer speed and quality, both in terms of software quality and quality of life. From prompt management, evaluation, and LLM application observability, this talk dives into major LLMOps focus areas and open-source software that were key to the success of our most ambitious projects, which have already delivered unparalleled customer value to over 23 million users of Japan's largest C2C e-commerce marketplace.The AI that solves the market: A new era in forecasting with natural language explainabilityWeights & Biases2025-07-07 | This Fully Connected 2025 San Francisco session introduces LG's groundbreaking Exaone Deep proprietary large language model and its revolutionary application in financial market forecasting. Attendees will discover how LG affiliates leverage this advanced AI technology to transform their decision-making processes through real-world use cases. The presentation will unveil our sophisticated forecasting framework that uniquely combines structured financial data with unstructured real-time news content to deliver superior market insights. The highlight of the session will be the announcement of our strategic partnership with a leading global financial data provider and the debut of our innovative product—an AI-powered US equities market forecasting score enhanced with natural language commentary that makes complex predictions accessible and actionable for investors.The future of multi-agents in enterprisesWeights & Biases2025-07-07 | Join us for 2025 Fully Connected San Francisco session led by Joao, CEO of CrewAI, as he explores the evolution and future of multi-agent systems. In this talk, Joao will share CrewAI’s pioneering approach to building agentic platforms, the challenges and opportunities in integrating multimodal capabilities, and what sets CrewAI apart in the rapidly growing agent ecosystem. He’ll also dive into real-world strategies for implementing agentic use cases and highlight how Crew is being used in production today.
Timestamps: 00:00 – Opening and Introduction to Crew AI 01:19 – 60 Million AI Agents Milestone 02:08 – Building the Crew AI Ecosystem 05:56 – 5 Levels of AI Agent Maturity 07:15 – Enterprise Stack for Agentic Resources 08:25 – Vendor Interoperability and Governance Needs 09:46 – Centralized vs Decentralized Agent Deployment 10:15 – Fortune 500 Implementation Case Study 13:12 – Common Pitfalls in AI Agent Adoption 17:54 – Real-World Use Cases: Code, Contracts, Support, OCRPhysical Intelligence Unleashed: Building robust AI at scale for Edge DevicesWeights & Biases2025-07-07 | Among the things that come to mind thinking of electric vehicles are clean air, butterflies, and sunshine. You also probably think of range anxiety and long charging times. Analog Devices (ADI) loves a big Edge AI challenge. As a company that works with car companies everywhere, we wanted to make things better.
EV batteries are multifaceted and complex. They involve many parameters, chemicals, and a bit of risk. That complexity demanded the use of machine learning models as the core of our solution. Given automotive requirements and regulations (ASPICE), we also looked for a tool to help us manage our experiments and reproducibility and provide an audit trail. That’s where W&B came in.
Yuval & Sheila share lessons learned collecting battery data, and the efforts to continuously track data, experiments, models, and results. The talk also explains the unique nuances and limitations of AI at the edge and how ADI overcomes them.Scaling GenAI inference: Techniques, optimizations, and real-world lessonsWeights & Biases2025-07-03 | Generative AI is transforming industries, but scaling models from research prototypes to production-grade systems presents significant challenges. This talk demystifies the core obstacles in Gen AI inference and offers practical strategies to reduce latency and control costs without compromising model performance. We'll explore advanced techniques including batching, model quantization, parallelism, KV cache management, and speculative decoding. Drawing from our hands-on experience, this session unpacks the trade-offs, pitfalls, and key lessons learned from scaling inference.Run the model, not the risk: Powering private inference for enterprise AI anywhereWeights & Biases2025-07-03 | Inference is the new frontier for enterprise AI—but using sensitive data securely is the bottleneck to production. When your model needs to run inference on sensitive/proprietary enterprise data, the default solution is to run on dedicated infrastructure. It’s seen as the only safe path, but it’s a trap: expensive to provision, difficult to manage, and impossible to scale without eroding your margins.In this talk we will introduce a new way forward: private inference anywhere. Using Protopia’s Stained Glass technology, your unmodified models can run without ever seeing raw data, eliminating the tradeoff between data privacy and AI use-case ROI.You’ll get a technical overview, real-world examples, and integration guidance that keeps your stack intact. Whether you're building AI inside the enterprise or selling it, this session shows how to escape the infrastructure trap—and run the model, not the risk.Synthetic data in medical device AI: Challenges and opportunitiesWeights & Biases2025-07-02 | Scarcity of same-domain training data is a major challenge faced by the developers of novel medical technologies. While deep transfer learning (DTL) with large foundation models has proven effective for a variety of applications, it is unclear the extent to which models trained on data from very disparate domains can be effectively repurposed for new types of medical data. SandboxAQ has developed a novel magnetocardiography (MCG) device for providing real-time decision support to cardiologists. Hailey and Geoff discuss how their team used large-scale synthetic data and deep transfer learning to address the data scarcity challenges associated with developing novel medical technologies.One size doesn’t fit all: Building AI agents specialized for your enterpriseWeights & Biases2025-07-02 | Generic AI agents consistently fall short in complex enterprise environments. Real-world success demands agents that are deeply specialized—able to retrieve, reason, and act within the unique systems, data, and workflows of each organization. You'll learn how we build domain-specific AI agents tailored to enterprises while giving customers full control over the specialization process, including training, evaluation, and continuous iteration. Discover how leading AI teams are moving beyond one-size-fits-all solutions to build AI agents purposefully designed for their use case, unlocking greater accuracy and real-world impact.A T cell foundation model for AI-powered target discovery and precision medicineWeights & Biases2025-07-02 | Identifying a potential drug candidate for a disease can take many years and tens of millions of dollars. AI has the potential to reduce overall time and cost to develop a drug. However, applications of AI to drug discovery have been hindered by the limited availability of datasets that are large in scale, yet sufficiently focused on specific biological systems of interest.
In this talk, we will describe how ArsenalBio’s automation lab and discovery platform enable large scale data generation for training and validating AI models of T cells, a core cell type of the immune system with key roles in cancer, autoimmunity, and infection. We will introduce the gx1 model, a foundation model of T cell biology. We will demonstrate how the gx1 model can be used for two applications: 1) virtual screens at a scale not possible through wet-lab experiments; 2) patient stratification to obtain biological insights inaccessible through other methods.Beyond RAG: Production-ready AI agents powered by enterprise-scale dataWeights & Biases2025-07-01 | Enterprise data isn’t always tidy—and AI agents need more than great retrieval to drive real value. In this session, we’ll share what we’ve learned at Snowflake about enabling agents that deeply understand and reason over structured business data, allowing our users to reliably "talk to their data." We’ll cover challenges like navigating messy schemas, generating trustworthy SQL, ensuring consistency in business semantics, and making the agent’s process visible to non-technical users. Whether you're scaling agent use across departments or starting to integrate them into core business workflows, you’ll leave with strategies to make agents effective, reliable, and trusted partners in the enterprise.Label factory: LLMs for training small language classification models at scaleWeights & Biases2025-07-01 | Social media content moderation at scale is a challenging task; beyond actually performing inference at scale, there are significant challenges to scaling up creation and evaluation of content moderation models. Many of these models have underlying policy that changes over time, and new models must be created on a routine basis.
Label Factory is Zefr’s patent-pending process of creating small, multimodal, multilingual classifiers for servicing Fortune 100 brands and advertisers. Rather than using labels that are annotated by humans first, it takes a knowledge-distillation approach to generate unsupervised classification labels for training purposes. These labels are over 90% accurate with their human counterparts and can be collected at scale.
Using our in-house infrastructure, as well as WandB, we are able to create/evaluate/deploy these small classifiers with a small fraction of the time/effort it used to take, without compromising on quality.W&B Inference: Access CoreWeave-hosted open-source models in W&B WeaveWeights & Biases2025-07-01 | W&B Inference provides API and playground access to leading open-source foundation models, without additional overhead of dealing with multiple model providers or hosting on their own. In this demo, we'll walk you through how to leverage W&B Inference for prototyping and evaluating AI applications using popular open-source LLMs.
⏳Timestamps: 0:00 Introduction to W&B Inference 1:54 How to access open-source LLMs in Weave 2:17 Tracing LLM calls using the Weave SDK 3:24 Testing an open-source LLM using W&B Inference in the playground 4:08 Comparing Weave traces 5:44 Setting up an AI application and Weave Evaluations using open-source LLMs 6:39 Exploring AI application evaluation results in Weave 7:51 Conclusion and invitation to try W&B WeaveToward zero traffic accidents: How we built a video AI agent for automated drivingWeights & Biases2025-07-01 | To build safer automated driving systems, we test our software by driving thousands of hours of real-world on-road vehicle testing. Manually checking these driving videos was slow and costly, so we developed AutoTriage—a video AI agent that finds system errors and identifies their root causes. We'll share three key takeaways behind our success: creating high-quality datasets, working closely with domain experts, and quickly adopting advances in generative AI. You'll learn practical steps to take your AI agent from prototype to production.
Chapters: 0:00 – Introduction: Woven by Toyota & Mission 1:41 – Why Automated Driving Matters 3:50 – Toyota’s Approach to Preventing Accidents 5:01 – Understanding Driving: Perception, Planning, Control 6:01 – Identifying Bottlenecks in AV Development 7:49 – Launching AUTO: Using GenAI for Bug Triage 9:18 – Challenges Building a Video AI Agent 12:21 – Overcoming Failures with Data, Prompts, and Models 14:26 – From Single Agent to Multi-Agent Judging 15:22 – Key Takeaways: Data, Collaboration, Persistence, Experiments 17:01 – Final Vision: Safer Roads Through AI-Driven AutomationFrom playground to production: Turbocharging GenAI innovation with AWS and Weights & BiasesWeights & Biases2025-06-30 | Discover how the AWS and Weights & Biases partnership accelerates enterprise GenAI development from experimentation to production. This session highlights key technical integrations between W&B's MLOps platform and AWS services, including testing Bedrock models in W&B Playground, evaluating LLMs on Bedrock, and monitoring Bedrock Agents with W&B Weave. Learn how these seamless workflows help teams iterate faster, maintain governance, and deliver production-ready GenAI applications with confidence.
Chapters: 00:00 – Welcome and Introduction of James Z 01:14 – The Generative AI Journey 03:35 – AWS’s Role in Supporting AI Development 03:51 – Amazon Bedrock and Foundation Models 05:10 – Amazon Nova and Content Generation Tools 06:49 – AWS Agent Options Overview 09:30 – Stran Agent SDK Deep Dive 13:32 – Quick Math Agent Demo 15:06 – AWS + Weights & Biases Integration 19:21 – Guardrails and Safety MeasuresFueling innovation at scale: Inside Pinterests machine learning platformWeights & Biases2025-06-30 | Explore how Pinterest's Machine Learning Platform drives innovation at scale, powering personalized experiences for millions of users worldwide. This talk offers a high-level overview of our approach to building ML platforms and the vibrant ecosystem that enables rapid experimentation and iteration of ML innovations. It will also highlight how Weights & Biases (Wandb) is seamlessly integrated into our ML lifecycle to support experiment tracking, model registry, and collaborative workflows. Discover how this integration streamlines our processes and empowers our teams to deliver state-of-the-art ML solutions at Pinterest's pace and scale.
Chapters: 0:00 – Introduction to ML Innovation at Pinterest 0:30 – ML Applications and Platform Overview 2:59 – Training & Inference Infrastructure 4:57 – ML Iteration Funnel: Model vs. Data 6:28 – Standardizing Model Development with MLM 9:59 – Pinterest’s Unified ML Platform Architecture 12:56 – Training Observability and Efficiency Gains 15:00 – Optimizing for Data Iteration with Ray 17:41 – New Developer Flow and Sampling Efficiency 19:13 – Key Takeaways and Closing ThoughtsTraining video models at scale with Adobe FireflyWeights & Biases2025-06-30 | Scaling generative video models poses unique challenges across architecture, data, optimization, and deployment. In this talk from Fully Connected 2025 San Francisco, Join Lior Shapira from Adobe to explore the key decisions involved in building these models—from designing architectures that balance quality and efficiency, to curating diverse and temporally coherent datasets, to managing large-scale training and inference-time constraints. Drawing on a real-world experience, the talk will offer practical insights into what it takes to train and deploy high-quality generative video models at scale.
Chapters: 0:00 – Introduction to Adobe Firefly Video GenAI 1:19 – Goals, Constraints, and Architectural Decisions 3:24 – Model Design: Diffusion Transformers & VAE Challenges 6:35 – Input Modalities and Premiere Pro Use Cases 8:40 – Training Curriculum and Responsible Data Strategy 11:01 – Scaling Challenges and Infrastructure Optimization 12:56 – Mixed Precision Training and Platform Monitoring 13:39 – Debugging Media Models: Confetti Bugs & Blinking People 16:21 – Inference Pipeline, Safety, and Deployment 18:59 – Performance Optimizations and Load Balancing 19:45 – Final Takeaways: Differentiation, Evaluation, and VisionEfficient Inference with Command A: Optimizing Speed and Cost for Enterprise AIWeights & Biases2025-06-27 | In the enterprise AI landscape, balancing speed, cost, and performance is critical. This talk explores the innovative techniques behind Command A's efficient inference pipeline, designed to deliver high-quality results at a low cost. We’ll delve into interleaved sliding window attention, which enhances both quality and speed, and discuss our optimizations like Speculative Decoding, sharing key insights from its training process. Join us to learn how Command A is redefining cost-effective AI for enterprise applications.
Chapters 0:00 – Introduction to Command R+ Inference Optimization 0:55 – Sparse Attention Architecture & Sliding Window 2:21 – Speculative Decoding Overview 4:32 – Using Medusa for Parallel Token Prediction 6:29 – Evaluation and Training with W&B 7:54 – Synthetic vs. Original Data in Speculative Training 9:00 – Final Gains and Performance Tradeoffs 11:44 – Guided Decoding with Speculative Inference 14:29 – Dynamic Guided Decoding and FSM Integration 19:03 – Combining Guided Decoding with Speculative TokensBuilding future-ready AI with agents & data flywheels: Insights from NVIDIA’s enterprise deploymentsWeights & Biases2025-06-27 | Santiago from NVIDIA shares insights, best practices, and lessons learned from building scalable, enterprise-ready AI agents using data flywheels. Drawing from our enterprise generative AI deployments, including chatbots, copilots, and ‘talk-to-your-data’ solutions, we’ll show how we implemented AI agents to orchestrate LLMs, APIs, and workflows for automating multi-step tasks, and data flywheels to drive continuous improvement of LLMs through user feedback. These architectural patterns are key to keeping enterprise AI solutions accurate, scalable, adaptable, and relevant in fast-paced business environments.
Chapters: 0:00 – Introduction and Audience Poll 0:36 – Framing Agents as Digital Employees 1:22 – NV Infobot: Architecture and Use Cases 3:00 – MAPE Framework for Self-Regulating Agents 3:34 – User Feedback Collection Challenges 5:42 – Root Cause Analysis with LLM Assistance 7:32 – Key Error Types and Prioritization 8:59 – Building the Data Flywheel 9:55 – NVIDIA Microservices for Agent Development 11:04 – Developer Workflow and Fine-Tuning 11:37 – Experiment 1: Router Optimization 13:06 – Experiment 2: Query Rephrasal Improvement 14:45 – Summary: Monitor, Analyze, Plan, Execute 15:32 – W&B Blueprint: Deploying Your Own Flywheel 16:44 – Q&A: Error Prioritization and Nemo Tools 19:30 – Closing Remarks and ApplauseThe open source AI compute tech stack: Kubernetes + Ray + PyTorch + vLLMWeights & Biases2025-06-26 | AI workloads are computationally demanding. They require scale for both compute and data, and they require unprecedented heterogeneity across workloads, models, data types, and hardware accelerators. As a consequence, the software stack for running compute-intensive AI workloads is fragmented and rapidly evolving. Companies that productionize AI end up building large AI platform teams to manage these workloads. However, within the fragmented landscape, common patterns are beginning to emerge. An emerging software stack combines Kubernetes, Ray, PyTorch, and vLLM. This talk describes the role of each of these frameworks, how they operate together, and illustrates this combination with case studies from Pinterest, Uber, and Roblox.
Chapters: 0:00 - Introduction to Ray and AnyScale 1:27 - Early Adoption and Growth of Ray 2:59 - Trend 1: Shift to Multimodal Data Processing 5:42 - Trend 2: Rise of Agentic AI and Multi-Agent Systems 7:48 - Trend 3: Post-Training and Reinforcement Learning Complexity 12:00 - Bridging the Gap Between AI Applications and Hardware 13:00 - Layered Tech Stack: Training/Inference, Distributed Compute, Orchestration 17:08 - Future of AI Tech Stack and Conference InvitationFrom research to reality in the age of (Gen)AIWeights & Biases2025-06-26 | The Generative AI era is fundamentally reshaping the pathway from research breakthrough to real-world impact and tangible products, moving us beyond an “Age of Data” into one of rapidly evolving models, emergent agentic capabilities, and the dawn of “superhuman data.” This presents unprecedented opportunities and unique challenges for product and engineering teams.
Join Xavi Amatriain, VP of Product, AI and compute enablement from Google who delves into the practical realities of this dynamic landscape. We will explore the shifting paradigm as powerful foundation models compress the research-to-product lifecycle, demanding new strategies for rapid iteration and “thinking ahead.” We’ll cover actionable insights for translating cutting-edge research—from novel architectures and advanced prompting techniques to agentic systems—into scalable, reliable product features. Critical to this journey are evolving best practices for robust evaluation beyond traditional metrics, mitigating hallucinations, ensuring responsible AI development, and understanding the evolving role of data. Furthermore, we’ll touch upon the increasing importance of comprehensive tools and platforms for managing experimentation, ensuring reproducibility, and accelerating the journey from idea to production-ready AI. Google’s full-stack infrastructure, from custom silicon to advanced frameworks, exemplifies how such capabilities enable the delivery of models at the Pareto frontier, ensuring production reliability, scale, and cost efficiency.
Chapters: 0:00 – From Research to Production: Setting the Stage 1:45 – The Early Days of AI: Linear Models to Deep Learning 4:07 – More Than Algorithms: UX, Domain Knowledge & Evaluation 7:12 – Enter GenAI: Acceleration in Research and Real-World Use 8:50 – Reinventing Evaluation: Creativity, Groundedness & Safety 11:02 – Rapid Innovation Cycles: Research to Production at Google 13:05 – Multimodal Agents: Astra and the Future of Evaluation 19:00 – Open Source Tooling & Agent Hackathons 21:09 – Beyond Today: Agents in Science and the Physical WorldShipping smart agents: lessons from the frontlinesWeights & Biases2025-06-26 | Join Alex Laubscher from Fully Connected San Francisco 2025 behind the scenes as of the first deployed engineers at Windsurf to explore how the company is revolutionizing enterprise software development. From launching a custom IDE to deploying proprietary LLMs like SUI-1, Windsurf is building the future of AI-assisted engineering.
🔍 Learn how Windsurf: -Built a flow-based human-AI collaboration model -Launched 10+ product releases in under a year -Developed Smart RAG and Riptide for advanced context retrieval -Integrated browser and PR tools for real-world dev workflows -Designed secure, enterprise-grade context ingestion with MCP and direct integrations -Trained and benchmarked their own models outperforming major competitors -iCreated a data flywheel to continuously improve model performance
📦 Whether you're an enterprise engineer, AI enthusiast, or just curious about what’s next in dev tooling—this session will give you real insight into building AI systems that work today while laying the groundwork for tomorrow.
Chapters: 0:00 - Introduction and Windsurf’s Mission 1:58 - Human-in-the-Loop and AI Collaboration 2:41 - Windsurf Editor and Product Velocity 4:00 - Beyond Codegen: Expanding the IDE Surface 7:00 - Extending into New Contexts: Reviews and Browsers 7:37 - Smarter Context: From Vectors to Riptide 9:55 - Enterprise-Ready Integrations and Data Privacy 12:05 - Frontier Models and the Data Flywheel 15:07 - Closing Vision and Product Flywheel StrategyPyTorch: The open language of AIWeights & Biases2025-06-25 | Over the past two years, the AI landscape has undergone a remarkable transformation. Large Language Models (LLMs) have moved to the forefront, powering applications like ChatGPT and driving an open revolution of models spearheaded by Llama. Now, we’re witnessing agentic systems entering the mainstream.
Despite these advances, significant challenges persist as we transition into a generative AI and agent-first world. These challenging questions require a collective community working towards common goals to address them effectively.
PyTorch has become that central place bringing together diverse perspectives with a community collectively building a comprehensive framework that integrates all layers of AI development. Joe Spisak, a long time leader of PyTorch and Llama open source, talks about the next phase of PyTorch, the growing ecosystem around Foundation 2.0 and how PyTorch continues to evolve as the open language of AI.
Chapters 00:00 - Introduction 01:50 - An abridged history of PyTorch 05:48 - The PyTorch ecosystem in a nutshell 08:28 - PyTorch adoption in AI research 10:12 - Major challenges in the AI space 12:40 - A broader vision for PyTorch 14:30 - What’s next for PyTorch: Planet-scale training and inference 17:10 - The difficulty of innovating SPMD 22:05 - What’s next for PyTorch: Can AI write the foundations of AI? 25:10 - Why are LLMs so bad at writing kernels? 28:49 - Closing thoughts and pointersBuilding super intelligent toolsWeights & Biases2025-06-25 | W&B Co-Founder Shawn Lewis built a programming agent that topped the SWE-Bench leaderboard for months. In this talk, Shawn talks about how his process, what he’s learned building agent experimentation loops, and how we’re bringing those lessons into Weights & Biases.
00:00 - The acceleration of AI viewed through programming agents 01:46 - Level-setting on AI agents: nomenclature, evals, and experimentation 03:18 - Experimentation types and processes 05:30 - Optimizing each step of the experiment loop 08:10 - Bringing W&B Launch back to solve knotty experimentation problems 11:50 - Optimizing the research phase in Weave 14:43 - Can we use AI to automatically improve the experimental loop? 17:39 - What changes with the resurgence of reinforcement learning 19:12 - OpenPipe’s Kyle Corbit on building reliable agents with RL 23:14 - Overcoming the limitation of evals with researcher agents 26:08 - How close are we to self-improving AI?Weights & Biases and CoreWeave: Fully Connected 2025 KeynoteWeights & Biases2025-06-25 | Join Weights & Biases Co-Founder Lukas Biewald and CoreWeave’s SVP of Engineering Chen Goldberg as they detail the first of many co-built projects to come.
After a brief introduction to the conference, Lukas gets into the state of the industry and how far AI has come in two short years. He touches on the massive improvements around all manner of generative models, how our kids will grow up in a new computing paradigm, and what new challenges come with agentic applications growing in both size and complexity. Lukas then runs through W&B’s newest releases for both W&B Models and Weave before Chen joins to showcase the three features we’ve been working on together.
00:00 - A brief history to the Fully Connected conference 02:44 - State of the industry: AI in 2025 06:55 - AI has come a long way in two short years 09:26 - DeepSeek and Jevon’s paradox 10:29 - Where is AI headed in 2025? 13:50 - New issues with AI in production 17:11 - Newest features: How W&B Models is evolving 20:53 - Newest features: How W&B Weave is evolving 24:36 - Introducing Chen Goldberg from CoreWeave 27:45 - CoreWeave’s purpose-built AI cloud services 30:50 - Demoing CoreWeave’s Mission Control in W&B 35:48 - Introducing W&B Inference powered by CoreWeaveAI’s $600B Question: Scaling for what comes nextWeights & Biases2025-06-20 | Join Sequoia’s David Cahn and Weights & Biases Lavanya Shukla from Fully Connected 2025 in San Francisco, CA for a wide-ranging interview about AI investing in 2025. They touch on everything from the AI talent wars and the future of search to how David thinks about investing in this market and where open-source fits into an increasingly competitive AI landscape.
00:00 - Introductions 00:45 - Why David invested in Weights & Biases 02:20 - David’s piece on AI’s $600B question 04:57 - Meta and the AI talent wars 08:45 - Can AI lead to a world with 1 billion developers? 12:05 - How founders can compete for AI talent 14:20 - The major difference between traditional and AI search 17:30 - What AI problems should people be working on today? 19:05 - Does founder-market fit matter as much as it once did? 21:10 - Why isn’t there an American Deepseek? 22:05 - Speed round inWhy developer happiness matters at GitHubWeights & Biases2025-06-12 | GitHub CEO Thomas Dohmke shares a core principle behind the company’s culture—and why happy developers build better products.