dbt LabsThere’s a good chance you’re an analytics engineer who just sort of landed in an analytics engineering career. Or made a murky transition from data science/data engineering/software engineering to full-time analytics person. When did you realize you fell into the wild world of analytics engineering?
In this session, Michael Chow (RStudio) draws upon his experience building open source data science tools and working with the data science community to discuss the early signs of a budding analytics engineer, and the small steps these folks can take to keep the best parts of Python and R, all while moving towards engineering best practices.
The accidental analytics engineerdbt Labs2022-10-25 | There’s a good chance you’re an analytics engineer who just sort of landed in an analytics engineering career. Or made a murky transition from data science/data engineering/software engineering to full-time analytics person. When did you realize you fell into the wild world of analytics engineering?
In this session, Michael Chow (RStudio) draws upon his experience building open source data science tools and working with the data science community to discuss the early signs of a budding analytics engineer, and the small steps these folks can take to keep the best parts of Python and R, all while moving towards engineering best practices.
Coalesce 2023 is coming! Register for free at coalesce.getdbt.com/.Live at dbt Summit: The parts of a data career that never make the job description (Episode 9)dbt Labs2026-09-25 | Recorded live at dbt Summit 2026, five women from dbt Labs, Faith McKenna, Erica Louie, Paige Berry, Grace Goheen, and Jerrie Kenney, get honest about the parts of a data career that never show up in a job description.
The group talks about the moment each of them first felt "technical enough," why curiosity matters more than having the right answer, what it actually looks like to bring your whole self to work while also parenting, managing a chronic condition, or setting hard boundaries, and how to name and advocate for the invisible "glue work" that keeps data teams running. They close out with live questions from the dbt Summit audience.
In this episode:
When each host first felt technical enough, and why curiosity mattered more than certainty What "bringing your whole self to work" looks like in practice, and why it requires trust and psychological safety on a team How to advocate for glue work and non-promotable labor in performance reviews, with real examples Live audience questions on staying passionate in a tough industry and finding community A lightning round: career advice each host wishes she'd heard earlier
0:00 Welcome and introductions 0:36 Meet the hosts 2:01 What's missing from the job description 2:17 Feeling technical enough 11:30 Bringing your whole self to work 23:10 The unseen "glue work," and advocating for it 28:52 Lightning round: career advice we wish we'd heard sooner 34:48 Audience Q&A begins 35:14 On losing (and finding) your passion 41:58 What's bringing us joy outside of work 46:56 Closing and creditsdbt v2 is heredbt Labs2026-09-16 | dbt v2 is here: a new foundation built for speed, rigor, and the agentic era.
dbt v2 adds a local compiler for SQL for the first time. Instead of finding out a query is broken only after running it against the warehouse, dbt now emulates your database's behavior and produces a logical plan before you deploy, catching mistakes early and saving time, money, and warehouse compute. It also gives agents a verifiable feedback loop instead of code they can only hope will work.
It's fast throughout. Moving to an optimized Rust binary and modern technologies like ADBC means a project that used to take 20 minutes now completes in under one, with everyday projects seeing meaningful speedups too.
Your whole project can now be described as standard Parquet files, cutting file size and enabling fast querying of your own project's metadata through tools like DuckDB. That means you can find dead columns, spot models with the least test coverage, enforce naming and ownership conventions as code, or let your agent look up lineage and blast radius instantly.
dbt docs also got a full rebuild: a new, scalable experience powered by an embedded DuckDB engine that reads project metadata on demand, so it holds up even on very large projects.
dbt v2 integrates seamlessly with dbt State, so you can skip or clone models that don't need to rerun.
Get started with the curl-based install for automatic updates, or use pip if you want dbt inside your existing Python virtual environment.
#dbt #AnalyticsEngineering #OpenSource #DataEngineeringdbt State: build only whats changed, skip the rest.dbt Labs2026-09-16 | Every job used to rebuild every model in the lineage, every single time, whether anything upstream actually changed or not. dbt State fixes that. On every run, it checks your metadata and model SQL to see what's changed. If nothing changed, it skips the build by reusing existing state at zero compute, cloning existing state at minimal cost, or auto-deferring to production state in development. If something did change, it builds. This works for full models, incremental models, snapshots, seeds, and tests.
The result: fresher data, lower warehouse costs, and no custom workflows or manual orchestration to maintain. Just turn dbt State on and run your jobs as often as you need.
Declare freshness and dependency rules directly in your project code, set a max staleness window so models rebuild only when they're due, and move orchestration logic from imperative scheduling into declarative, version-controlled configuration.
Data teams are already seeing the impact: CarGurus cut compute 9%, built 35% fewer models, and reduced Snowflake costs by 15%. Fanatics saw 25 to 30% model reuse and roughly 15% in Snowflake savings on their pilot. Obie Insurance is saving at least 30% on compute costs from model reuse alone.
dbt State works whether you're orchestrating in the dbt platform, running dbt Core locally, or using an external orchestrator.
#dbt #dbtState #AnalyticsEngineering #DataEngineering #DataOpsdbt Wizard: The agent purpose-built for analytics engineeringdbt Labs2026-09-16 | Meet dbt Wizard, the AI agent built for analytics engineers.
Generic coding agents can get the code right and still get the project wrong. dbt Wizard is different: it's grounded natively in your dbt project's lineage, compiled state, tests, and metric definitions, so every change comes with full context on what's upstream, downstream, and at stake.
With dbt Wizard, you can:
Refactor: rename or restructure models and Wizard updates every ref that needs to move, with a reviewable diff before anything ships.
Build: describe what you need and get back a validated model, tests, docs, and semantic definitions, all at once, all grounded in your existing project.
Investigate: tell Wizard what broke and it traces your lineage, proposes a fix, and validates it before surfacing an edit for review.
Migrate: move models without breaking dependencies, or fix Fusion conformance errors automatically.
Every change self-validates before you review it, and stays governed and auditable by default, so you always know what changed, why, and who approved it. Fewer production incidents, less time fixing what AI got wrong.
dbt Wizard is available as a terminal-native CLI for local development (works with or without a dbt platform subscription) and inside the dbt platform.
#dbt #dbtWizard #AnalyticsEngineering #AIAgent #DataEngineeringYou don’t have ICs anymore, you have managers (Emilie Schario)dbt Labs2026-09-09 | Emilie Schario is co-founder and head of product and engineering at Kilo Code, the open-source, model-agnostic coding agent platform that Anaconda acquired in July 2026. Before Kilo, she was an early member of GitLab's data team, ran the data org at Netlify, spent a year as data-strategist-in-residence at Amplify Partners, and founded Turbine, an ERP startup later acquired by Settle.
She joins Tristan Handy for Season 9 of The Analytics Engineering Podcast, our season on Analytics x Agents.
They cover "kilo speed" and the org design that makes it real (about 20 engineers, one product hire, every engineer owning their own area with a team of AI agents behind them), why Kilo bets on being model-agnostic across more than 500 models instead of tying itself to one lab, the Kilo founding story with Sid Sijbrandij and Scott Breitenother, and Emilie's argument that data people, already used to being pulled in a dozen directions at once, may be better prepared than software engineers for a world where everyone manages a portfolio of agents.
👤 Guest: Emilie Schario, co-founder and head of product and engineering at Kilo Code 🎙️ Host: Tristan Handy, co-founder of dbt Labs
Chapters:
0:00 Welcome, and Emilie's podded bio 2:05 What "kilo speed" means 5:55 The product-engineer model: one PM, 20 engineers 8:03 "You don't have a bunch of ICs anymore, you have managers" 8:28 Is this the death of the junior software engineer? 12:01 Emilie's path: Smile Direct Club, Doist, GitLab, Netlify 17:28 Founding Turbine, and the lesson about NetSuite 22:34 The Kilo founding story: Sid's call, and "temporary" 27:08 How Kilo differs from Claude Code or Codex 34:37 Kilo's open source roots: Cline, Roo, and a fork of a fork 38:11 Working with model labs before launch 44:53 What's the Kilo business model? 47:02 The right model for the task, not the cheapest one 52:09 The flip: data people as natural multi-threaded managers
The Analytics Engineering Podcast is sponsored by dbt Labs. Reach us at podcast@dbtlabs.com with comments and guest suggestions.
#analyticsengineering #dbt #aiagents #agenticengineering #dataengineeringRoundup: The post-AI data stack, physical AI, and the fight over data centersdbt Labs2026-09-03 | Tristan Handy and Jason Ganz are back with four stories: Accelerated Understanding's claims about simulating physical reality, the political fight over data centers, Ian Macomber's read on the shape of the post-AI data stack, and an update on the OpenAI/Hugging Face incident and what it means for how data teams think about evals.
👤 With: Jason Ganz; Host: Tristan Handy, co-founder of Fivetran + dbt Labs
Chapters: 0:00 Welcome back, episode three of the Roundup 1:13 Accelerated Understanding and physical AI 6:54 Is this the breakthrough, or something else? 9:10 The context window as an article of faith 11:16 The political fight over data centers 16:15 Why the backlash isn't irrational, even if the facts are wrong 21:13 Compute demand booked out through 2028 23:02 Ian Macomber on the post-AI data stack 25:32 Two jobs for the post-AI data team 30:51 "I want to use my agent to use your thing" 36:10 What data practitioners actually do next 40:07 An update on the Hugging Face incident 42:53 700 agents, 70,000 messages, and the Artifactory exploit 43:29 Astra, recurrent depth, and chain of thought monitoring 52:52 Where evals need to go next 55:52 Wrap-up and dbt Summit
The Analytics Engineering Podcast is sponsored by dbt Labs. Reach us at podcast@dbtlabs.com with comments and guest suggestions.
Guy Podjarny, co-founder of Tessl and founder of Snyk, on why instinct works for your own prompts and breaks the moment a skill gets shared.
He joined Tristan Handy this week on the Analytics Engineering Podcast.
Three ideas from the episode: 🔸 Context has a forcefulness dial. Rules load every turn. Skills load on a hint. Passive docs sit there until an agent goes looking. The front matter is the always-loaded part, which is why 100 skills end up competing for the same sliver of attention. 🔸 Skills are software, so they carry software risk. Malicious skills lifted wholesale from the open ecosystem. Negligent ones missing basic safety instructions. Vulnerable ones that walk an agent into pasting credentials somewhere they land in a log. 🔸 Most companies will own a harness, not rent one. Tristan thought building a dbt-specific harness was a bad idea, since Claude Code and Codex already work. It was neither hard to build nor a wash: better accuracy, meaningfully fewer tokens.
Listen wherever you get your podcasts 🎧You cant scale vibes (Guy Podjarny)dbt Labs2026-08-26 | Guy Podjarny has built developer tools three times over. He founded Blaze, a web performance company acquired by Akamai, and then founded Snyk, a developer-first security company. Now he's co-founder of Tessl, betting that software development is shifting from revolving around code to revolving around intent and instructions.
He joins Tristan Handy to discuss Guy's model of the agentic stack (models, tools, context, and harness), why he thinks skills need the same lifecycle as code or teams end up scaling nothing but vibes. They give into the security risks of treating skills as disposable markdown files, and why most companies will end up owning their own harness instead of renting one from a frontier lab, as cost pressure and open-weight models like GLM 5.1 change the math.
👤 Guest: Guy Podjarny, co-founder of Tessl, previously founder of Snyk and Blaze 🎙️ Host: Tristan Handy, co-founder of dbt Labs
Chapters: 0:09 Welcome 0:36 Guy's path: AppSec, Blaze, Akamai, Snyk, Tessl 3:46 "A glutton for punishment": what keeps him starting companies 7:00 Why Tessl chose to be co-located in London 11:16 The emerging agentic stack: models, tools, context, harness 15:42 The three buckets of context: policies, specs, workflows 21:19 What a skill.md file actually contains 23:59 "The inability to scale vibes" 27:23 Malicious, negligent, and vulnerable skills 28:56 From spec-centric development to speccing the programmer 31:49 Tessl as an agent enablement platform and composable factory 33:04 The new Tessl Agent: a factory-builder harness 38:48 Why the labs won't own developer tooling too 42:09 GLM 5.1, and talking about cost without sounding like a Luddite 44:02 Tristan on rethinking the dbt-specific harness 48:44 What an eval actually is 54:03 Wrap-up
The Analytics Engineering Podcast is sponsored by dbt Labs. Reach us at podcast@dbtlabs.com with comments and guest suggestions.
#analyticsengineering #dbt #aiagents #agenticdevelopment #dataengineeringRoundup: Why your sales team shouldnt query Gong directlydbt Labs2026-08-21 | "You can burn a lot of credits very quickly if you aren't careful. Make it an incremental dbt model so you don't blow through your credit budget."
Tristan Handy on what happens the moment you point a whole sales team at the same unstructured data.
He's joined by Jason Ganz for the Roundup: a recurring conversation covering data and AI news at the speed it moves.
Three ideas from the episode:
🔸 The five Exponential View gauges say boom, not bubble. AI capex is still under 1% of GDP, well below the 2.5% during the railroad build-out, and revenue is doubling every seven months. Funding quality is the one drifting toward red.
🔸 Two data engineering benchmarks landed in two weeks. Opus 5 cleared 70% on Snowflake's 103 dbt-specific tasks. Models handle bounded problems well. They're still bad at saying "I don't know" and at spotting a question that needs reframing.
🔸 Hook every rep up to the Gong MCP and everyone pays for the same analysis over and over. Britton Stamper took the queries people actually ask and aggregated them into an incremental dbt model.
Listen wherever you get your podcasts 🎧Roundup: Why your sales team shouldnt query Gong directlydbt Labs2026-08-21 | Tristan Handy and Jason Ganz are back for Roundup, a new format for The Analytics Engineering Podcast and ouur season on analytics and agents.
Four stories: Stripe's $7 billion acquisition of OpenRouter, two new benchmarks showing data engineering agents have gotten reliably good, Exponential View's latest bubble gauges (still no bubble, though funding quality is worsening), and how dbt Labs' Britton Stamper is turning Gong call data into something an agent can use without burning through the token budget.
Chapters: 0:00 Welcome back, episode two of the Roundup 0:55 Exponential View's bubble gauges 5:01 Is this cycle actually vibes-driven? 7:40 Real treasury yields as a hard-to-read signal 9:05 Does the bubble question even matter long-term? 12:23 Two new data engineering benchmarks 14:20 Snowflake's data-eng-bench, and what it tests 17:09 Why the harness matters as much as the model 19:48 Where the model scores land 21:02 Hex's benchmark on thorny analytical questions 22:00 The overthinking curve at max effort 25:06 The meta takeaway: what do we build now? 27:19 The fundamentals caveat 28:48 Stripe acquires OpenRouter for $7 billion 33:32 Locked into one lab versus fully model-agnostic 35:38 Will model workloads become portable? 40:55 50 trillion tokens a month, and what Stripe learns from it 41:21 Britton Stamper turns Gong calls into dbt models 45:12 Why Gong calls were impractical to analyze before 47:43 Britton's approach: aggregate into incremental dbt models 52:56 The dbt Summit fireside chat, and wrap-up
The Analytics Engineering Podcast is sponsored by dbt Labs. Reach us at podcast@dbtlabs.com with comments and guest suggestions.
#analyticsengineering #dbt #dataengineering #aiagents #openrouterdbt Core v2: faster parsing and AI-ready metadatadbt Labs2026-08-14 | We asked Penny for her take on dbt Core v2. 🐾
Faster parsing. AI-ready metadata. Woof. (Not bad for a dog)
If you want the full breakdown from actual humans on our product team, join the webinar on August 26th.
Save your seat: bit.ly/4wHlvJbDont hand a bazooka to an agent making a sandwich (Jeremiah Lowin)dbt Labs2026-08-13 | “Hey Claude, can you go make me a sandwich? And by the way, here's a bazooka in case you run into any zombies."
Jeremiah Lowin, founder and CEO of Prefect, on why a 15-step workflow in a single skill file hands agents a dangerous tool at step one.
He joined Tristan Handy this week on the Analytics Engineering Podcast. Three ideas from the episode:
🔸 A skill file is a polite note the agent is free to ignore. Fine for steering behavior. Less fine when steps 1-14 check that everything is correct and step 15 moves the money. 🔸 Agentic workflows are DAGs of outcomes. Traditional workflows are DAGs of implementations. Every workflow tool we have was built for the second kind. 🔸 You can't archive your way to a context layer. A stock price changes by the second, so the layer with real gravity governs access to context rather than storing it.
Listen wherever you get your podcasts 🎧Dont hand a bazooka to an agent making a sandwich (Jeremiah Lowin)dbt Labs2026-08-12 | Jeremiah Lowin built FastMCP as a weekend side project a few days after Anthropic introduced the Model Context Protocol. He contributed it into Anthropic's official SDK, archived his own repo, and figured that was the end of it. Months later, after OpenAI and Google both announced MCP support, he woke up to find that archived do-not-use-this repo sitting at number one trending on GitHub. So he changed his own job at Prefect for 60 days and went back to being a maintainer.
He joins Tristan Handy for Season 9 of The Analytics Engineering Podcast, our season on Analytics x Agents.
They cover why MCP's real product market fit is the enterprise rather than the individual developer, why a skill file is a polite note the agent is free to ignore, what breaks when you treat one as a 15-step workflow, why agentic workflows are DAGs of outcomes while traditional workflows are DAGs of implementations, and why Jeremiah stopped leading with the context layer when the enterprises calling him want security plumbing first.
👤 Guest: Jeremiah Lowin, founder and CEO of Prefect, creator of FastMCP 🎙️ Host: Tristan Handy, co-founder of dbt Labs
Chapters: 0:00 Welcome 0:35 A new dawn for Prefect 1:08 Why Prefect, FastMCP, and Horizon are one company 3:05 How FastMCP happened by accident 4:21 Everyone suddenly cares about automation 5:22 MCP was impossibly hard to use 6:28 Anthropic asks to bring FastMCP into the official SDK 7:48 OpenAI and Google announce support, and all hell breaks loose 8:23 The archived repo that hit number one on GitHub 10:03 Firing himself as CEO to go be a maintainer 11:34 Two FastMCPs, and the coming rename 14:09 What MCP actually is 15:15 MCP versus CLI 16:11 Why MCP's product market fit is the enterprise 17:41 Tristan on the dbt MCP server's adoption 18:48 Namespacing tools and the context window problem 20:42 Skills versus MCP: steering behavior vs. granting capability 23:20 Skill files that escalate into "do not skip this step" 24:31 Why skills aren't workflows 25:02 The bank: step 15 moves the money 25:52 The bazooka and the sandwich 26:23 Tristan's constrained, deterministic approach 27:49 The autonomy spectrum 29:30 "How do you know the agent did what you wanted?" 30:53 DAGs of outcomes vs. DAGs of implementations 31:16 The Ralph loop and goal-seeking behavior 33:31 Agent swarms: Tristan's eight-hour experiment 35:05 Can you cleanly define success? 37:07 Approving pull requests all day is not a job 38:14 How much harder this gets inside the enterprise 39:32 Horizon: an MCP gateway and governance layer 40:29 Shadow IT and the seven MCP servers nobody tracks 41:18 The four-step enterprise journey 42:18 MCP Apps: ship a UI, not a hundred thousand rows 43:41 What you solve at the FastMCP layer vs. Horizon 45:00 "We're going to stop bringing you to meetings" 46:31 What even is context? 48:03 Continuous context and the price of a stock 50:17 From context provider to context access layer 54:23 Is Prefect an AI company yet? 55:50 Wrap-up The Analytics Engineering Podcast is sponsored by dbt Labs. Reach us at podcast@dbtlabs.com with comments and guest suggestions. #analyticsengineering #dbt #mcp #fastmcp #aiagents #dataengineeringDo you have to be a dbt platform customer to use dbt State?dbt Labs2026-08-12 | Do you have to be a dbt platform customer to use dbt State? 🤔
Nope. David Macias breaks it down: dbt State works anywhere you run dbt, Core or platform, with any orchestrator you're already using.
Turn it on, set your lag tolerance based on how often you run jobs, and it only builds what's changed.
Join the webinar on September 1 to see how it works, with in-depth examples from real scenarios from Joel Labes and Reuben McCreanor. Save your seat: bit.ly/4hvioiGWhat does dbt Core v2 mean for dbt platform customers?dbt Labs2026-08-10 | We asked Aika Zikibayeva what dbt Core v2 means for platform customers. Her answer: not much. 😌
You're already running on the same dbt Fusion engine, so nothing really changes on your end.
If you want to move onto Fusion faster, it's all self-serve: dbt platform has an upgrade flow that shows you which jobs are eligible, and dbt Wizard 🪄 helps you find and fix whatever's standing between your project and running smoothly on Fusion.
Register for the webinar on August 26 to learn more bit.ly/4wHlvJbWhen does dbt Core v2 go GA?dbt Labs2026-08-06 | When does dbt Core v2 go GA?
You'll have to wait until dbt Summit for Jeremy Cohen to spill the tea. 🍵
#dbtSummit is September 15-18 in Las Vegas.
Join us bit.ly/4gefI7ZRoundup: a rogue agent, Kimi K3, and data teams in the AI eradbt Labs2026-08-06 | Something new. Tristan Handy is joined by Jason Ganz for the first Roundup, a recurring conversation about the stories moving the data ecosystem right now. Part of Season 9 of The Analytics Engineering Podcast, our season on Analytics x Agents.
Three stories: the OpenAI model that escaped its evaluation sandbox and spent days inside Hugging Face infrastructure, and the legal vacuum it leaves behind. Moonshot's Kimi K3, and why a 2.8 trillion parameter model under a restrictive license might be the first defensible open weights business. And Katie Bauer's read on what changes for data teams as agents arrive, including the return of the full stack data analyst.
👤 With: Jason Ganz, dbt Labs 🎙️ Host: Tristan Handy, co-founder of dbt Labs
Chapters: 0:00 Welcome, and why we are trying a new format 1:07 Katie Bauer on data teams in the AI era 4:32 Platform work versus distribution work 5:44 Is self-serve actually here? 6:30 Guardrails, context, and the dbt MCP server 8:10 Does the data team start to look like IT? 10:05 Humans are still the destination for the work 12:58 The case that long-running agents are the future 14:00 What agentic use cases are companies actually targeting? 15:20 The return of the full stack data analyst 18:20 The OpenAI agent that hacked Hugging Face 20:52 Reading the Hugging Face disclosure 24:07 The asymmetry between attacker and defender 26:36 Two camps on how to prioritize defense over offense 29:54 The AOL-ification of the internet 34:01 Strict operator liability, and why it does not apply 36:07 Are our legal institutions ready for rogue agents? 37:11 The Hugging Face CEO's two asks 41:20 So what should a data leader actually do? 45:07 How documentation shot up the priority list 46:38 Kimi K3, the most capable open weights model released 50:29 Can you actually run a 2.8 trillion parameter model? 52:38 The restrictive license, and why it matters 54:57 Is open core a good business model for a model company? 57:59 If openness is not about price, why care? 59:26 Open data infrastructure is about choice 1:01:08 The coding harness is the real lock-in 1:02:53 Wrap-up, and send us your feedback
#analyticsengineering #dbt #dataengineering #aiagents #openweightsHow to get ready for dbt Core v2dbt Labs2026-08-05 | We stopped Grace Goheen on a dog walk to ask about dbt Core v2.
Here's the fastest way to get there: -Get on dbt Core v1.12 first -Run your project through the new v2 parser to catch issues early -Hit something broken? dbt-autofix (plus the new agent skill) cleans up most of it automatically -Parse clean, and you're ready to upgrade
If you want to work through it with our team, we're staffing a power-up station built for v2 upgrades at dbt Summit in Vegas this September.
Penny's ready for v2. Your project can be too.
Check out the update guide bit.ly/4bvUIqD Register for the webinar on August 26 to learn more bit.ly/4wHlvJbData lessons from inside Meta (Shridhar Iyer)dbt Labs2026-07-30 | Shridhar Iyer spent more than 13 years inside Meta's data organization, most recently leading AI and the data stack. Now on a career break, he joins Tristan Handy for Season 9 of The Analytics Engineering Podcast, our season on Analytics x Agents.
They open with an unexpected detour into the hard problem of consciousness, then get practical: how truncating a single column called Extra saved Meta a few million dollars, why you almost never actually delete a column at Meta scale, how Meta's stack evolved from Hadoop and Hive into a schematized and unified system, and the two steps every company needs to become AI-native.
👤 Guest: Shridhar Iyer, formerly Senior Tech Lead and Director for AI and the data stack at Meta 🎙️ Host: Tristan Handy, co-founder of dbt Labs
Chapters: 0:00 Cold open: the column called Extra 1:23 Welcome: 13 years at Meta, and a career break 2:13 Why a data person takes philosophy seriously 2:40 The hard problem of consciousness 7:16 From consciousness to data and AI 8:02 Two reasons AI pulled him back to philosophy 14:18 Early career: JB Hunt, an MBA, and a pivot into data 18:18 A retail analytics startup and the earliest days of data engineering 20:34 Joining Meta in 2013 as one of the first data engineers 21:28 What Meta's data stack looked like: Hadoop, Hive, DataSwarm 24:33 Splitting 13 years into eras: centralized to embedded in product 27:14 The scale: tens of millions of tables 29:17 The column called Extra, and the "fix of the week" 31:29 How you delete a column at Meta scale 35:25 What big tech knows that will leak out to the rest of us 37:13 Schematize, unify, then add semantics 42:17 Onboarding at Meta, and why the hard problems are analytics problems 44:29 How hiring shifted from builders to analytics engineers 45:57 The data team's role in becoming AI-native 46:34 Two steps to AI-native: readiness and archetypes 50:03 Org structures, fewer management layers, and experimentation 54:13 Managing 40 direct reports and the end of the middle layer 56:09 Super pods, and building a knowledge layer for agents 57:42 Wrap-up
The Analytics Engineering Podcast is sponsored by dbt Labs. Reach us at podcast@dbtlabs.com with comments and guest suggestions.
#analyticsengineering #dbt #dataengineering #aiagents #metaThe Analytics Engineering Podcast: Data lessons from inside Meta (Shridhar Iyer)dbt Labs2026-07-30 | You almost never delete a column at Meta. Except once, when it saved a few million dollars.
Shridhar Iyer is this week's guest on the Analytics Engineering Podcast. He spent 13 years inside Meta's data org, most recently as Senior Tech Lead and Director for AI and the data stack, and joined Tristan Handy to unpack a decade of lessons.
Three ideas from the episode:
🔸 At Meta scale, deleting a column is a company-wide event. Sri once truncated a debug column called Extra and saved a few million dollars. But on core tables, you never delete or truncate. You create a new version, migrate people over with tooling, and just stop populating the old one.
🔸 Meta's real advantage was building abstractions in layers. Schematize first, enforcing strong types all the way up to the log statement. Unify next, one compiler and language across Spark and Presto, one catalog where every table and column is a typed asset with a URI. Only then do semantic and knowledge layers for AI become possible. Sri thinks most companies are behind on this sequence.
🔸 Becoming AI-native takes two steps, and most teams skip the first. Step one is AI-readiness: take a single workflow, learn to run it well, extract reusable primitives, and figure out the lowest-cost way to use agents with real guardrails. Skip that and you burn tokens for bad outcomes. Step two is reorganizing around three archetypes: the builder who automates workflows, the forward-deployment engineer who makes environments AI-ready, and the domain specialist embedded in a product.
Listen wherever you get your podcasts 🎧dbt MCP server in Claude: Query your governed data in plain Englishdbt Labs2026-07-21 | The dbt MCP server is now available as a remote connector in Claude, so you can query your governed data in plain English without installing anything. In this quick demo I connect dbt to the Claude web app and ask my semantic layer real questions about financial and economic data.
Everything runs on the dbt Semantic Layer, so the metrics Claude returns are the same governed numbers the data team owns. No hallucinated metrics, and no digging through five dashboards to piece an answer together. I walk through discovering what lives in the project, checking recent job runs so the data is current, and asking a couple of analytical questions that come back as charts and narratives I can trust.
Chapters: 0:00 Using the dbt MCP server inside Claude 0:06 What projects and semantic layer am I connected to? 0:21 Exploring metrics, models, and question types 0:34 Checking project state and recent job runs 0:54 Claude's overview of models, metrics, and domains 1:39 Semantic layer query: asset class volatility over five years 2:19 Why governed metrics mean answers you can trust 2:44 Follow-up: risk-adjusted returns and positive skew 3:19 Governed tables and charts, no hallucinated metrics 3:45 Try it yourself Connect dbt to Claude: docs.getdbt.com/docs/dbt-ai/integrate-mcp-claude Full list of dbt MCP tools: docs.getdbt.com/docs/dbt-ai/mcp-available-tools Questions? Join #tools-dbt-mcp in the dbt community Slack. If you already build with dbt, connect the MCP server and start with one semantic layer question today. #dbt #Claude #MCP #analyticsengineering #dataengineering #semanticlayerThe scarce resource is consensus (Ian Macomber)dbt Labs2026-07-15 | The last time Ramp’s Ian Macomber joined the show, the episode was titled, "Ramp's $8 billion data strategy." Ramp is now valued at $44 billion, and the data team's job has completely changed. Ian, who leads data at Ramp, returns to talk with Tristan Handy about how Ramp got self-serve analytics to 50x scale with an internal tool called Ramp Research, routing agents through the data lake, and why consensus is the scarce resource in a post-AI world.
👤 Guest: Ian Macomber, Head of Analytics Engineering and Data Science at Ramp 🎙️ Host: Tristan Handy, co-founder of dbt Labs
Chapters: 0:00 Cold open: consensus is the scarce resource 1:33 Reintroducing Ian and Ramp 2:06 From an $8 billion data strategy to a $44 billion one 4:26 Setting up the stack for text-to-SQL 5:34 The two jobs of a post-AI data team 7:02 Ramp Research and the questions that went unasked 8:20 From 10x to 50x questions in six months 10:00 A finance app built on top of Ramp Research 12:09 Ramp Research as the API to the data lake 13:46 Snowflake Summit and the word "agent" 17:37 Why agents are not built on the data lake 19:37 Permissions: the unit is still the human 22:29 If the warehouse was buttoned up, agents do not add new risk 25:43 "We were too empathetic" 27:33 The token maxing era and the leaderboard 28:54 Job two: building a singular reality 30:22 The 7-out-of-10 dashboard problem 33:08 Chasing slop vs. creating meaning 35:44 Building consensus in a world of agents 36:22 Canonical metrics moving into the dbt repo 38:29 Evals, skills, and paring back context 41:32 Consolidating the data team's job families 45:55 The AI Index and what Ramp is watching 49:27 Wrap-up
The Analytics Engineering Podcast is sponsored by dbt Labs. Reach us at podcast@dbtlabs.com with comments and guest suggestions.
#analyticsengineering #dbt #dataengineering #aiagents #fintechThe Analytics Engineering Podcast: The context engineering playbook (Claire Gouze)dbt Labs2026-07-07 | Context engineering is the new analytics engineering.
Claire Gouze, co-founder and CEO of nao Labs, took her agent from 40% to 90% reliability. The fix was cleaning up her data model and writing better docs.
She joined Tristan Handy on the Analytics Engineering Podcast to walk through her context engineering playbook. Three ideas from the episode:
- Context engineering is the new analytics engineering. Same job as always: turn tacit business knowledge into something structured and trustworthy. New medium: markdown and files, not just models. - The biggest reliability gains are unglamorous. Query logs and fancy context sources barely moved the needle. Cleaning up the data model and writing real documentation is what got her agent from 40% to 90%. - We're in the "just plug it into production" era of agents. Wiring an agent straight into every raw source is the same mistake as plugging a BI tool into a production database, a decade later. Context needs its own stack.
Listen wherever you get your podcasts 🎧 roundup.getdbt.com/p/the-context-engineering-playbookThe context engineering playbook (Claire Gouze)dbt Labs2026-07-01 | Context is everything in data right now. In this episode, Tristan Handy talks with Claire Gouze, co-founder and CEO of nao Labs, an open-source analytics agent. Claire and her team have done the hands-on work of context engineering: authoring a playbook, building evals, and growing a community that is figuring it out in the open. This is a pragmatic conversation about how to build a context layer your agents can rely on.
Chapters 00:00 Cold open: the missing piece is context 01:45 Welcome, and a few words on the French accent 02:52 Claire's path into data: business school to BCG Gamma 04:09 Joining sunday and building a data stack from scratch 06:39 Founding nao, an open-source analytics agent 10:42 The journey so far: 1,300 GitHub stars, 80 companies in production 13:45 Why pivot from "cursor for data" to the context layer 15:22 Context layer hype at Snowflake Summit 17:48 Is there a product in context, or just best practices? 21:23 The context engineering playbook: start small, iterate 24:45 Where the 50 eval questions come from 25:40 Ramp Research and the questions data teams never got asked 27:06 Is "context engineer" a real job title? 28:55 Flattening roles on the data team 29:56 Machine vs. human context, and where to keep it 32:43 The highest-signal context: clean data models and docs 35:36 Where nao is headed: a source of truth across Slack, MCP, and more 37:41 From the data lake for analytics to infrastructure for agents 39:05 The "just plug it into production" moment and the context stack 41:49 Roadmap: automating context creation 43:33 Context as organizational memory 46:00 File systems, MetricFlow, and the semantic layer question 49:05 Why open source, and the commercial model 51:39 Wrap-up
The Analytics Engineering Podcast is sponsored by dbt Labs. Learn why more than 80,000 data teams use dbt to accelerate their data development: getdbt.com
Questions, comments, and guest suggestions: podcast@dbtlabs.com
#analyticsengineering #dataengineering #AI #agents #contextengineering #dbtMotherDucks 3ms Query Time — Jordan Tigani | The Analytics Engineering Podcastdbt Labs2026-06-27 | "Our median query time is three milliseconds." "Oh, sorry. I just needed a minute to absorb that."
Jordan Tigani of MotherDuck on The Analytics Engineering Podcast. Listen and subscribe roundup.getdbt.com/p/duckdbs-agent-moment-jordan-tiganiDuckDB is Built for AI Agents — Jordan Tigani | The Analytics Engineering Podcastdbt Labs2026-06-26 | `brew install bigquery` is not a thing.
Jordan Tigani of MotherDuck on why DuckDB's local-first architecture turns out to be a perfect fit for agents.
Jordan Tigani (MotherDuck) said it. Tristan had thoughts.
Listen to the full episode of The Analytics Engineering Podcast roundup.getdbt.com/p/duckdbs-agent-moment-jordan-tiganiHow Anthropic hit 95% accuracy on analytics queries (it wasnt a smarter model)dbt Labs2026-06-23 | Most AI projects don't fail because the AI is bad. They fail because the data underneath isn't ready. Here's what it actually takes to get your dbt project and your data AI-ready, the difference between an agent you can trust and one that lets you be wrong faster.
I walk through the two halves of AI readiness for data teams: using AI to build and maintain your project (dbt Wizard, the dbt MCP server, and dbt agent skills), and getting your data ready for an agent to query reliably so a business user can ask "what was revenue last quarter" and get the right number every time.
We dig into how Anthropic gets around 95% accuracy on 95% of their internal analytics queries, and why it came from removing the agent's freedom to guess rather than from a smarter model. Then the practical part: structured context, a governed semantic layer, fewer and better models, data quality, and the governance an agent inherits for free.
The takeaway: data modeling, defined metrics, documentation, and data quality were good engineering before agents existed. Agents just raise the cost of skipping them.
Chapters 0:00 — Why most AI projects fail 0:48 — The two jobs of AI readiness 2:03 — The dbt stack: Wizard, MCP server, agent skills 3:00 — How Anthropic gets to 95% accuracy 3:27 — The three failure modes: ambiguity, staleness, retrieval 4:16 — What AI-ready data looks like in dbt 4:53 — The semantic layer, and why it's non-negotiable 5:35 — Fewer, better models and pruning zombies 5:50 — Data quality, freshness, and blast radius 6:12 — Why this beats generic text-to-SQL 6:36 — Trust is the product we deliver 7:08 — Two paths: where to start
Try it: install dbt agent skills or dbt Wizard, point it at your project, and watch where the agent gets confused. That's a free map of where your context is thin.
#dbt #analyticsengineering #dataengineering #AI #semanticlayer #dataqualityDuckDBs agent moment (Jordan Tigani)dbt Labs2026-06-18 | Jordan Tigani helped build BigQuery, then left to bet that most data isn't big. Three years on, agents are proving him right.
The MotherDuck CEO joins Tristan Handy to open Season 9 of The Analytics Engineering Podcast, our season on Analytics x Agents. They get into why local-first, single-node databases fit the agent era, how MotherDuck stays faster and cheaper than the incumbents, where data lakes and Iceberg fit in, and what an "agent swarm for data management" would actually do.
👤 Guest: Jordan Tigani, co-founder and CEO of MotherDuck 🎙️ Host: Tristan Handy, co-founder of dbt Labs
Chapters: 0:00 Why DuckDB is having an agent moment 1:08 Reintroducing Jordan and MotherDuck 3:10 The MotherDuck and DuckDB Labs relationship 6:50 Where the line gets drawn: DuckDB vs. MotherDuck 9:44 Quack and the enterprise layer 11:05 How widely is DuckDB actually used? 13:25 Who uses MotherDuck, and what they migrate from 14:41 Hyper-tenancy and read scaling 18:05 Why it's faster and cheaper: latency vs. throughput 21:35 "Big Data is Dead," revisited 25:45 Iceberg, Duck Lake, and DuckDB on the lake 31:19 "ETL is highly vibe codable" 32:36 Tristan pushes back 35:00 From MCP server to Dives 38:35 What do agents want from an analytical database? 41:37 Two kinds of agents 43:09 The local-first advantage 46:02 The agent swarm and "Water-Town" 52:06 A thousand data analysts and the cost of inference 54:53 Wrap-up
The Analytics Engineering Podcast is sponsored by dbt Labs. Reach us at podcast@dbtlabs.com with comments and guest suggestions.
#analyticsengineering #dbt #dataengineering #duckdb #aiagentsdbt Wizard CLI demo: An AI agent that knows your datadbt Labs2026-06-17 | One curl command. Your own AI key. No dbt platform account required.
dbt Wizard CLI 🧙 is now in public beta, and it’s open to every dbt user. It’s an AI agent designed for analytics engineering.
Here's why it feels like magic: It ships with a native metadata engine that indexes your entire project before your first prompt. Lineage, tests, metrics, contracts, run history. It already has the map. Tell it a model broke. It knows which tool to call, traces lineage to the root cause, writes the fix, and proactively validates it works — before you ever review it. And when it builds, it only builds what changed.
That means less compute, faster iterations, lower warehouse bills.
Your keys, your terminal, your project. Available for self-hosted or platform users. BYOK with OpenAI, Anthropic, Bedrock, and Snowflake Cortex.
🪄 Try it today: docs.getdbt.com/docs/dbt-ai/wizard-quickstartHow Virgin Media O2 scaled their data platform with dbtdbt Labs2026-06-10 | Linh Perkins of Virgin Media O2 shares how adopting dbt transformed the way their data team works. The team moved from fragmented tools and siloed teams to a unified, standardized platform built on a SQL-first principle.
Linh walks through why dbt was one of the best decisions Virgin Media O2 made: how it created visibility into business logic, built trust in data across the organization, and gave the team a way to enforce best practice at scale without sacrificing speed or flexibility.
Key themes: — Breaking down silos and creating code visibility — Building data trust through transparency — Templated projects that enforce best practice across teams — Reducing duplication and freeing up engineering time — A secure, scalable, sustainable data environment that enables innovation
Chapters 0:00 Why dbt was one of the best decisions Virgin Media O2 made 0:32 Building trust through code visibility and business logic 1:10 Enforcing best practice at scale with templated projects 1:25 Reducing duplication and freeing up engineering time 1:54 Security, sustainability, and making space for innovation
See how other data teams are using dbt to build trust, scale faster, and cut technical debt: getdbt.com/customersWant to use AI with dbt? Start here.dbt Labs2026-06-09 | Wondering how to actually bring AI into your dbt workflow? The dbt MCP server is one of the fastest ways to get started. In this video, I break down what it is, how it works, and three things people are already doing with it. The dbt MCP server is dbt Labs' implementation of the Model Context Protocol — an open standard that lets AI tools like Claude, Cursor, and custom agents securely access your dbt project's structured context: your models, metrics, lineage, job runs, and more. Whether you're an analytics engineer wondering where AI fits into your day-to-day, or a data leader evaluating how to bring AI into your stack — this is the place to start. Chapters: 0:00 What you'll learn 0:15 What is MCP? 0:55 What the dbt MCP server actually does 1:20 Three real use cases 2:20 Local vs. remote: control vs. convenience 2:45 How to get started
dbt State is in Preview: intelligent orchestration that checks warehouse metadata and model SQL on every run and only builds a model if the result would actually change.
If data hasn't changed and code hasn't changed, the model doesn't run. On every run, dbt State decides whether to build, skip, clone existing state, or auto-defer to production. On average, that means 30% less warehouse compute. In development, it means fast iteration without the risk of costly mistakes.
It also eliminates the custom workflows most teams have built to avoid rebuilding everything on every run. No more job-by-job schedules, sub-selectors, or manifest scripts. Turn dbt State on and let it build.
Available as a plugin for dbt Core or out-of-the-box in the dbt platform.
Learn more about dbt State getdbt.com/product/dbt-stateWhat every analytics engineer needs to know about dbt Wizarddbt Labs2026-06-03 | dbt Wizard is a coding agent built for analytics engineering. It runs in your terminal or in the dbt platform, and the difference is what it's grounded in: instead of grepping your files, it reads your project's actual structure—the DAG, column-level lineage, your tests, your compiled state. So when you ask it to change something, it already knows the blast radius.
In this video Developer Experience Advocate Alex Noonan walks you through three things you can do today:
— Rename a column and propagate it downstream, with a reviewable diff before anything saves — Hand it a failing test and watch it trace the lineage to the root cause and fix it — Explore an unfamiliar model and see what feeds into it
Wizard validates its own work before you ever see it. It runs modified models against dev, defers unchanged models to prod state, compares row counts, and hands you an impact report. It edits files only—it never writes to your tables—and you stay in the loop on every change.
Chapters: 0:00 The problem with generic agents 0:25 What is dbt Wizard 1:05 Demo: rename and propagate a column 1:45 How the validation loop works 2:25 Demo: investigate a failing test 2:55 Where it runs and what teams are seeing 3:20 Get starteddbt Wizard: the agent purpose-built for analytics engineeringdbt Labs2026-06-02 | dbt Wizard: the agent purpose-built for analytics engineering
Introducing dbt Wizard. An AI agent built from the ground up for analytics engineers. Not just code generation, but the entire data lifecycle. It knows your lineage, your contracts, your tests, and your metric definitions before it writes a single line.
dbt Wizard knows how dbt actually works. It's optimized for dbt workflows out of the box. It always calls the right dbt-specific tool, reasons from your complete dbt artifact set, validates by default, and shows you every change as a visual DAG so you understand how the logic actually evolved, not just what code changed. No MCP to maintain. No DIY stack. And it's available wherever analytics engineering work actually happens, in the dbt platform or CLI.
Learn more about dbt Wizard getdbt.com/product/dbt-wizardFivetran + dbt Labs complete merger to create the data infrastructure for trusted AI agentsdbt Labs2026-06-01 | Swipe right if you move fast 🔥
It's a match, and it's official. Fivetran and dbt Labs have completed our merger to build the data infrastructure for agents.
The numbers were already there: 16,000 dbt projects run on Fivetran data every week, 1,500 joint customers, and 100,000+ data teams combined.
Reliable data movement plus governed, trusted transformation. That's exactly what agents need to work, and it’s what Fivetran + dbt Labs delivers.
Fivetran moves and manages data across any source. dbt ensures it's defined, tested, and trusted with shared business logic, governed context, and software engineering best practices built into the data lifecycle. Open standards. No lock-in. Any cloud, engine, or tool.
Thiago Baldim, Senior Manager, Data Engineering, and Yuna Tang, Senior Analytics Engineer, walk through how dbt transformed their workflow, from fragmented pipelines to a fully documented, fully tested, AI-ready data platform.
The results: -Processing time dropped from 14 hours to 90 minutes -Self-serve data access expanded across the entire business -AI product adoption is growing every day -New team members can onboard and contribute without the guesswork
"Without dbt, we would never have been AI-ready."
Learn more about how dbt helps data teams build trusted data products at getdbt.com/case-studiesThe dbt Developer Agent is now in Preview: the coding agent for analytics engineeringdbt Labs2026-05-06 | The dbt Developer Agent is now in preview.
Most coding agents see one file. The dbt Developer Agent sees your whole project.
Available in dbt Studio IDE, it’s a coding agent built for analytics engineering, grounded in your lineage, tests, docs, and semantics. That means faster, safer changes and fewer downstream surprises.
• It understands context beyond the file (lineage, tests, docs, semantics) • It keeps related files in sync (SQL + YAML + docs) • And it ships with reviewable diffs + human-in-the-loop control
It’s built for the work analytics engineers actually do: • Multi-file refactors • Migrations • Keeping tests and docs aligned with SQL changes • Faster iteration, fewer broken dashboards
Read the blog to learn more getdbt.com/blog/the-dbt-developer-agent-is-now-in-previewCounting what matters: How dbt Labs tracks Fusion success (ft. Paige Berry)dbt Labs2026-04-16 | Anders Swanson, Senior DX Advocate at dbt Labs, sits down with Paige Berry, Lead Data Analyst at dbt Labs, to talk about how the team is measuring dbt Fusion engine adoption and what the data is actually showing.
They cover: -How dbt Labs defines weekly active projects and why that metric translates to real-world Fusion adoption -What it means for a paying customer to be "onboarded" to Fusion -How ARR from Fusion-onboarded customers is becoming a key business signal -Why retention is the core lens for measuring developer experience success -What it looks like when VS Code extension users keep coming back week after week, month after month -How telemetry infrastructure makes all of this measurable in the first place
Learn more about the dbt Fusion engine: getdbt.com/product/fusionInside the rewrite: How Fusion is reimagining static analysis and the DAG (ft. Chenyu Li)dbt Labs2026-04-16 | Anders Swanson, Senior DX Advocate at dbt Labs, sits down with Chenyu Li, Staff Software Engineer at dbt Labs, to talk about one of the most intellectually demanding parts of the dbt Fusion engine rewrite: what happens to familiar dbt concepts when you introduce static analysis.
Chenyu has been in the Fusion rewrite since the beginning, and the problem that has occupied most of that time is schema.
They cover: -Why "what is a schema" becomes a genuinely hard question when static analysis enters the picture -How deferral works in dbt Core and why Fusion's static analysis makes it significantly more complex -Why you can't just point refs to the production schema when you're propagating from sources locally -How the analyze phase adds a new layer to the task graph that didn't exist in dbt Core -What DAG optimization could look like when Fusion can identify shared operations across models and collapse them into a single executor node
Learn more about the dbt Fusion engine: getdbt.com/product/fusiondbt Core vs. the dbt Fusion engine: A Support team deep dive (ft. Jeremy Yeo & Shelli White)dbt Labs2026-04-16 | Anders Swanson, Senior DX Advocate at dbt Labs, sits down with Jeremy Yeo, Staff Customer Solutions Engineer, and Shelli White, Manager of Customer Solutions Engineering at dbt Labs, to get the support team's perspective on what's actually different about the dbt Fusion engine.
Nobody sees the real-world impact of an engine rewrite quite like the people fielding support tickets every day. Jeremy and Shelli have been close to both dbt Core and Fusion, and they have a lot to say.
They cover: -What made dbt Core simple and what Fusion's added complexity actually buys you -How Fusion's built-in debugging tools change the support experience -Why state modified explainability is a meaningful quality-of-life improvement for users -How the new system report feature cuts down on support back-and-forth -What detailed Fusion logging unlocks compared to dbt Core's SQL-only output -Why switching between Fusion versions is a one-liner instead of a Python virtual environment headache
Learn more about the dbt Fusion engine: getdbt.com/product/fusionWhats coming after dbt Fusion engine GA (ft. Alex Bogdanowicz)dbt Labs2026-04-16 | Anders Swanson, Senior DX Advocate at dbt Labs, sits down with Alex Bogdanowicz, Sr. Manager, Software Engineering at dbt Labs, to talk about where the dbt Fusion engine goes after GA and what analytics engineers should be thinking about right now.
Alex has a hot take: stop using union relations. From there, the conversation opens up into a broader discussion about where the analytics engineering craft is headed.
They cover: -Why union relations creates serious performance problems at scale and what's driving it -How Jinja's synchronous nature pushes teams toward patterns that work but don't scale -The evolution from "unsafe" to "baseline" and why that framing shift matters -What the Fusion team can build now that parity with dbt Core is no longer the primary focus -What's on the roadmap: checks, governance, classifiers, local schema sourcing, and compute
A lot of the complexity in analytics engineering projects today exists because better abstractions didn't exist yet. Fusion is where those abstractions get built.
Learn more about the dbt Fusion engine: getdbt.com/product/fusionHow Fusion enhances local and warehouse compute with Rust and Arrow (ft. Diego Fernandez)dbt Labs2026-04-16 | Anders Swanson, Senior DX Advocate at dbt Labs, sits down with Diego Fernandez, Senior Software Engineer II at dbt Labs, to talk about his path from the Semantic Layer team to Fusion, and what he sees as the opportunity on the other side of GA.
Diego spent years building the Semantic Layer gateway, a query execution engine built on Arrow, before making the move to Fusion. That background gives him a unique vantage point on where the stack is headed.
They cover: -How the Arrow ecosystem spans languages from Kotlin to Python to Rust, and why that matters -What the learning curve looks like going from Kotlin and Python to Rust -How AI tooling has changed the experience of picking up a new language -The parallelization work Diego did to unify file parsing and Jinja rendering logic across Fusion -Why Python's parallelization story created limitations that Rust opens back up -What becomes possible when local Fusion compute and warehouse compute start working together
The part worth watching for: Diego's take on what Fusion makes possible that wasn't even on the table with dbt Core.
Learn more about the dbt Fusion engine: getdbt.com/product/fusionState modified: Solving the multi-language engine puzzle (ft. Nick Othieno)dbt Labs2026-04-16 | Anders Swanson, Senior DX Advocate at dbt Labs, sits down with Nick Othieno, Senior Software Engineer at dbt Labs, to talk about nine months spent on one of Fusion's most deceptively complex features: state modified.
State modified lets teams run only the subset of their DAG that actually changed, which translates directly to warehouse cost savings. The concept is straightforward. The implementation was not.
They cover: -What state modified is and why it matters for teams running large dbt projects -How data representation differences between Python and Rust created unexpected edge cases -The YAML Norway problem and why strictness in Rust surfaces issues that Python quietly ignored -Why Fusion moved from string comparisons to checksums and what that change exposed -What it's like to ship a feature that depends on basically all of Fusion, while 30 other engineers are shipping at the same time
Learn more about the dbt Fusion engine: getdbt.com/product/fusionReimagining Jinja: How we brought dbts superpower to the Fusion engine (ft. Zhong Xu)dbt Labs2026-04-16 | Anders Swanson, Senior DX Advocate at dbt Labs, sits down with Zhong Xu, Senior Staff Software Engineer at dbt Labs, to talk about one of the more technically demanding parts of building the dbt Fusion engine: making Jinja work in a Rust-based engine.
Jinja is what makes dbt so expressive. It's also what made the Fusion migration hard. Zhong breaks down the work that went into getting it right.
They cover: -What Jinja is and why it's both dbt's superpower and a double-edged sword -How the team moved from Python's Jinja to MiniJinja, a Rust-native implementation -Why re-implementing pure Python functions in Rust was the hard but necessary path -How a side-by-side testing system ensured migration parity between Python and Rust rendering -What's next: bringing a typing system to Jinja so bugs surface at compile time, not runtime
That last part is worth paying attention to. A typed Jinja means hover-over types, go-to-definition for macros, and errors that tell you exactly what's wrong and where.
Learn more about the dbt Fusion engine: getdbt.com/product/fusionHow the Fusion team is owning their destiny with Arrow ADBC (ft. Mila Page)dbt Labs2026-04-16 | Anders Swanson, Senior DX Advocate at dbt Labs, sits down with Mila Page, Senior Software Engineer II at dbt Labs, to talk about one of the bigger mindset shifts behind the dbt Fusion engine: owning the driver stack.
Mila spent years on the dbt Core team shipping adapter work before moving to Fusion. That context makes her perspective on the transition worth hearing.
They cover: -What it felt like to go from "that's a driver problem" to owning the drivers outright -How Arrow ADBC provides a unified standard that made taking on that ownership possible -Why controlling the driver stack means faster fixes, fewer blockers from upstream vendors -How contributions from the Fusion team are getting upstreamed into open standards, benefiting the broader data ecosystem
The shift from depending on warehouse providers to move first has been a meaningful one. Bugs get fixed on the Fusion team's timeline, not someone else's.
Learn more about the dbt Fusion engine: getdbt.com/product/fusionThe future of dbt adapters: Building with Apache Arrow & Fusion verticalization (ft Felipe Carvalho)dbt Labs2026-04-16 | Anders Swanson, Senior DX Advocate at dbt Labs, sits down with Felipe Carvalho, Senior Staff Software Engineer at dbt Labs, to talk about two of the more consequential architectural decisions behind the dbt Fusion engine: adopting Apache Arrow as a standard and rethinking how adapters are built.
Felipe is a direct contributor to the Arrow ecosystem, and brings that depth to the conversation. They cover: -What Arrow is and why standardizing on it as an in-memory data format matters -How Fusion uses Arrow throughout, including inside the Agate library in Jinja -Why row-based database APIs created so much overhead and what Arrow changes -How the new adapter architecture (verticalization) centralizes shared logic across data platforms -What ADBC is and why it sits at the foundation of the Fusion adapter stack
The goal of verticalization: instead of defining a hundred functions per adapter, you contribute to one shared implementation and add your special case only when you need to. Every new adapter makes the next one easier to build.
Learn more about the dbt Fusion engine: getdbt.com/product/fusionBridging the gap from dbt Core to Fusion: Package compatibility and dbt autofix (ft. Chaya Carey)dbt Labs2026-04-16 | Six months in, Chaya Carey has been deep in the work that makes dbt Fusion engine adoption possible for real-world projects.
Anders Swanson, Senior DX Advocate at dbt Labs, sits down with Chaya, Software Engineer II at dbt Labs to talk about what it takes to get the package ecosystem ready for Fusion. With 350+ packages and 4,200+ package versions in Package Hub, the range of how teams use dbt is enormous.
They cover: -Why Jinja flexibility is both a superpower and a migration challenge -How Package Hub evolved to surface Fusion compatibility information -What dbt autofix handles automatically so you don't have to -How the team moved from manual package reviews to automated Fusion parsing at scale
The goal: no more waiting to hit problems yourself. The tooling does that work upfront.
Learn more about the dbt Fusion engine getdbt.com/product/fusionHow AI is reshaping the way data practitioners workdbt Labs2026-04-06 | Hot take from this week's pod: the hardest questions AI is surfacing in data work aren't new. The outcome vs. process tension has always been there. AI just raises the stakes.
🎧 Listen to The View on Data wherever you get your podcasts.Ready for the dbt Fusion engine? A practical framework for modern data teamsdbt Labs2026-04-03 | The dbt Fusion engine changes how dbt works under the hood, bringing development, orchestration, governance, and insights into a single engine. For dbt Core and dbt platform users, the question is no longer if Fusion makes sense, but how prepared their existing projects and workflows are to adopt it.
But adopting Fusion isn’t just a technical upgrade. It is a chance to step back and examine how your data team builds, governs, and delivers trusted data products today.
In this session, Brooklyn Data joins us to walk through their Fusion Readiness Assessment, a structured framework designed to help dbt Core and dbt platform users understand what Fusion readiness really looks like and how to evaluate their own starting point. You’ll learn what Fusion enables, why readiness matters, and how teams can prepare their eligible dbt projects to realize value from Fusion without unnecessary risk or disruption.
What you’ll gain from watching: -What Fusion is and why it matters: How the dbt Fusion engine changes the way teams build, orchestrate, govern, and analyze data -How ready you really are: How project, process, and people readiness affect adoption, using Brooklyn Data’s Fusion Readiness Assessment and real-world examples -What to do next: A practical framework to prioritize next steps and accelerate value after Fusion is implemented gaps
Who should watch this workshop: -Analytics engineers -Analytics engineering leaders -Data engineering and platform teams -Heads of data, analytics, or data platforms -Data architects and technical leads -Organizations using dbt Core or dbt Platform who are evaluating or planning for Fusion
Meet the speakers: Michael Carlone, Director, Analytics Engineering at Brooklyn Data Aika Zikibayeva, Senior Product Marketing Manager at dbt Labs