Uploaded June 2026 | Updated September 2026, 3 weeks ago
From open-sourcing the layer above coding agents to rethinking databases for the agent era, Databricks cofounders Matei Zaharia and Reynold Xin are pushing the company beyond the lakehouse into a full data-and-AI operating system. In this episode, Matei and Reynold join swyx after Data + AI Summit to unpack Omnigent, LTAP, Lakebase, agent security, open formats, Mosaic, and why databases may matter more than ever once AI agents start doing real work.
We go deep on Omnigent: Databricks’ open-source meta-harness for combining, controlling, and sharing agents across Claude Code, Codex, Cursor, Pi, custom agents, and internal tools. Matei explains why coding agents and enterprise agents run into the same problems: portability, collaboration, session history, security, spend controls, and the need for a common API above every harness.
Then Reynold walks through Databricks’ database dream: why CDC is brittle enough to joke that it means “continuous data corruption,” why HTAP has been the holy grail of database engineering, and why Databricks thinks LTAP gets most of the benefits by unifying the storage layer instead of collapsing every query engine. We also cover Databricks’ infrastructure scale, the culture behind rapid prototyping, the difference between tech and enterprise customers, Databricks vs Snowflake, whether vector databases should have ever existed, the Mosaic model strategy, Genie, AI Runtime, RL fine-tuning, and the thesis that traditional software gets rewritten once the data is in the right place and agents sit on top.
We discuss:
• Why Databricks built Omnigent as a meta-harness above existing AI agents
• Why coding agents and custom enterprise agents need the same infrastructure
• The common API for agent sessions, files, streams, tool calls, and cancellation
• Why persistent sessions, cloud sandboxes, sharing, search, and collaboration matter
• Why Databricks open-sourced Omnigent instead of keeping it proprietary
• Databricks’ internal agent usage, cloud sandboxes, and coding workflows
• The scale of Databricks: 50–60 million virtual machines a day and exabytes before breakfast
• Why agent security needs contextual and stateful policies
• How an agent could read confidential docs, install a compromised npm package, and leak data
• Why spend control matters when an agent can burn $500 reading logs
• Startup opportunities around coding-agent analytics, quality, skills, and spend
• LTAP, Lakebase, and why Databricks wants to rethink the database stack
• OLTP vs OLAP, CDC, and why data pipelines break at 3 a.m.
• Why HTAP has historically been the holy grail of database engineering
• Why Databricks thinks LTAP is “HTAP done right”
• How writing transactional data into column-oriented formats changes analytics
• Why agents need live operational context from databases, not just telemetry
• How Databricks prototypes strategic systems without endless process
• Enterprise vs tech customers, governance, procurement, and DIY culture
• The “second system syndrome” risk of rewriting a database engine
• Building a database engine from a decade of traces and quadrillions of data points
• Why vector databases should never have been a separate category
• Why Databricks thinks open formats and AI changed the race with Snowflake
• The Mosaic story, DBRX, Genie, document parsing models, and specialized model training
• Why model customization and RL fine-tuning may become mainstream
• Why “get the data there, slap some agent on top” may rewrite traditional software
—
Matei Zaharia
• LinkedIn: linkedin.com/in/mateizaharia
• X: https://x.com/matei_zaharia
Reynold Xin
• LinkedIn: linkedin.com/in/rxin
• X: https://x.com/rxin
Databricks
• Website: databricks.com
• X: https://x.com/databricks
Timestamps
00:00:00 Hook
00:01:13 Introduction
00:03:35 Omnigent and the Agent Infrastructure Layer
00:09:52 Agent Clouds, Common APIs, and Open Source
00:18:05 Databricks Scale and Internal AI Workflows
00:19:16 Agent Security, Governance, and Spend Controls
00:28:47 LTAP and the Database Dream
00:31:43 CDC, HTAP, and Why Data Pipelines Break
00:35:18 Lakebase, Parquet, and Live Data for Agents
00:38:00 Databricks’ Culture of Fast Prototyping
00:44:53 The Dream Engine and Rewriting the Database Stack
00:52:15 Vector Databases, Query Engines, and LTAP
00:53:49 Databricks vs Snowflake
00:59:01 Mosaic, DBRX, Genie, and Specialized Models
01:04:24 Context, AI Runtime, and RL Fine-Tuning
01:07:28 Why Data + Agents May Rewrite Software
01:08:22 Closing Thoughts
From open-sourcing the layer above coding agents to rethinking databases for the agent era, Databricks cofounders Matei Zaharia and Reynold Xin are pushing the company beyond the lakehouse into a full data-and-AI operating system. In this episode, Matei and Reynold join swyx after Data + AI Summit to unpack Omnigent, LTAP, Lakebase, agent security, open formats, Mosaic, and why databases may matter more than ever once AI agents start doing real work.
We go deep on Omnigent: Databricks’ open-source meta-harness for combining, controlling, and sharing agents across Claude Code, Codex, Cursor, Pi, custom agents, and internal tools. Matei explains why coding agents and enterprise agents run into the same problems: portability, collaboration, session history, security, spend controls, and the need for a common API above every harness.
Then Reynold walks through Databricks’ database dream: why CDC is brittle enough to joke that it means “continuous data corruption,” why HTAP has been the holy grail of database engineering, and why Databricks thinks LTAP gets most of the benefits by unifying the storage layer instead of collapsing every query engine. We also cover Databricks’ infrastructure scale, the culture behind rapid prototyping, the difference between tech and enterprise customers, Databricks vs Snowflake, whether vector databases should have ever existed, the Mosaic model strategy, Genie, AI Runtime, RL fine-tuning, and the thesis that traditional software gets rewritten once the data is in the right place and agents sit on top.
We discuss:
• Why Databricks built Omnigent as a meta-harness above existing AI agents
• Why coding agents and custom enterprise agents need the same infrastructure
• The common API for agent sessions, files, streams, tool calls, and cancellation
• Why persistent sessions, cloud sandboxes, sharing, search, and collaboration matter
• Why Databricks open-sourced Omnigent instead of keeping it proprietary
• Databricks’ internal agent usage, cloud sandboxes, and coding workflows
• The scale of Databricks: 50–60 million virtual machines a day and exabytes before breakfast
• Why agent security needs contextual and stateful policies
• How an agent could read confidential docs, install a compromised npm package, and leak data
• Why spend control matters when an agent can burn $500 reading logs
• Startup opportunities around coding-agent analytics, quality, skills, and spend
• LTAP, Lakebase, and why Databricks wants to rethink the database stack
• OLTP vs OLAP, CDC, and why data pipelines break at 3 a.m.
• Why HTAP has historically been the holy grail of database engineering
• Why Databricks thinks LTAP is “HTAP done right”
• How writing transactional data into column-oriented formats changes analytics
• Why agents need live operational context from databases, not just telemetry
• How Databricks prototypes strategic systems without endless process
• Enterprise vs tech customers, governance, procurement, and DIY culture
• The “second system syndrome” risk of rewriting a database engine
• Building a database engine from a decade of traces and quadrillions of data points
• Why vector databases should never have been a separate category
• Why Databricks thinks open formats and AI changed the race with Snowflake
• The Mosaic story, DBRX, Genie, document parsing models, and specialized model training
• Why model customization and RL fine-tuning may become mainstream
• Why “get the data there, slap some agent on top” may rewrite traditional software
—
Matei Zaharia
• LinkedIn: linkedin.com/in/mateizaharia
• X: https://x.com/matei_zaharia
Reynold Xin
• LinkedIn: linkedin.com/in/rxin
• X: https://x.com/rxin
Databricks
• Website: databricks.com
• X: https://x.com/databricks
Timestamps
00:00:00 Hook
00:01:13 Introduction
00:03:35 Omnigent and the Agent Infrastructure Layer
00:09:52 Agent Clouds, Common APIs, and Open Source
00:18:05 Databricks Scale and Internal AI Workflows
00:19:16 Agent Security, Governance, and Spend Controls
00:28:47 LTAP and the Database Dream
00:31:43 CDC, HTAP, and Why Data Pipelines Break
00:35:18 Lakebase, Parquet, and Live Data for Agents
00:38:00 Databricks’ Culture of Fast Prototyping
00:44:53 The Dream Engine and Rewriting the Database Stack
00:52:15 Vector Databases, Query Engines, and LTAP
00:53:49 Databricks vs Snowflake
00:59:01 Mosaic, DBRX, Genie, and Specialized Models
01:04:24 Context, AI Runtime, and RL Fine-Tuning
01:07:28 Why Data + Agents May Rewrite Software
01:08:22 Closing Thoughts






![[State of Research Funding] Beyond NSF, Slingshots, Open Frontiers — Andy Konwinski, Laude Institute
From co-founding *Databricks* and *Perplexity* to launching the *Laude Institute*—a dual venture fund and nonprofit designed to turbocharge the path from *research breakthrough to breakout company*—*Andy Konwinski* is building the infrastructure to recreate the Databricks motion at scale: fund researchers doing open work, help them ship products that matter, and turn paradigm-shifting ideas into trillion-dollar companies. We caught up with Andy live at *NeurIPS 2025* to dig into the origin story of Laud (right resource, right researcher, right time), why the *Databricks founding model* (eight co-founders, deep research scars, years of collaboration) is becoming the gold standard for AI startups (not an anomaly), how Laudes *venture arm* backs researchers-turned-founders with 50+ professors and PhDs as LPs (Jeff Dean, top Berkeley/Stanford faculty, Databricks and Perplexity co-founders), why the *nonprofit arm* does no-strings-attached grants to fund open research before incorporation (the upstream funnel that feeds the next generation of companies), the *slingshot program* funding breakthrough projects like *DSPy, Terminal Bench, LMArena, and continual learning research,* why *NSF isnt broken but insufficient* (its $1B/year for computer science when we need $10-100B, and Silicon Valleys picker model can deploy capital more effectively), how the *post-post-training layer* (prompt optimization, context management, RAG, memory curation, tool usage) is becoming the new frontier above pre-training and post-training, why *Chinese labs are outpublishing Western labs* in open research (Moonshot, DeepSeek shipping twice as many interesting papers as American startups because OpenAI and the frontier labs stopped publishing), the launch of *Open Frontiers*—a live-streamed conference in San Francisco bringing together the 100 most influential open researchers (Yann LeCun, François Chollet, Jan Leike, Percy Liang, Berkeley AI Research, Allen Institute, and more) to share roadmaps and unify the ecosystem, why the *Laude Lounge at NeurIPS* became the VIP gathering spot (Starlink WiFi, free food, couches, and the gods of AI hanging out because conferences need a place for the VVIPs to actually sit down), and his thesis that *open research is the path to world-changing impact*—and Laud is the bridge from grant to company, from paper to product, and from researcher to billionaire founder. We discuss:
* What *Laude Institute* does: dual structure with a *venture fund* (backing researchers-turned-founders post-incorporation) and a *nonprofit* (no-strings-attached grants for open research pre-incorporation)
* The *slingshot program:* funding *DSPy, Terminal Bench, LMArena, continual learning, and Jepa-style prompt optimization* projects across Berkeley, Stanford, MIT, CMU, Wisconsin, Caltech, UI Urbana-Champaign, Toronto, Waterloo, and beyond
* The *post-post-training layer:* compound systems, prompt optimization, context management, RAG, memory curation, tool usage—the layer above pre-training and post-training where innovation is exploding
* *GEPA and DSPy:* evolutionary prompt optimization (genetic algorithms reinvented by PhD student Laxia) and the DSPy framework (a reverse compiler that takes code and compiles natural language)
* Why *NSF isnt broken but insufficient:* $1B/year for computer science (and theyre trying to cut it in half) when we need $10-100B for frontier AI research—Laud complements NSF with Silicon Valleys picker model and high-velocity grant writing
* The *PhD entrepreneurship clubs:* Computer Science Grad Entrepreneurs (CSGE) at Berkeley (started 2012), Agent at University of Washington, Research to Impact at Wisconsin, Saplings at Stanford, and more clubs forming at CMU, MIT, UI Urbana-Champaign
* *Open Frontiers:* a live-streamed conference in San Francisco (next five months) bringing together the 100 most influential open researchers (Yann LeCun, François Chollet, Jan Leike, Percy Liang, Berkeley AI Research, Allen Institute, and more) to share roadmaps and unify the ecosystem
* The vision: *open research as the path to world-changing impact,* and Laud as the bridge from grant to company, from paper to product, and from researcher to trillion-dollar founder
— Andy Konwinski
* Laude Institute: https://www.laude.org/lounge
* X: https://x.com/andykonwinski
00:00:00 Introduction: Andy Konwinski and the Laud Institute Vision
00:01:17 The Databricks Motion: From PhD Research to Billion-Dollar Companies
00:02:15 Lauds Two-Sided Model: Venture Fund and Philanthropic Grants
00:06:37 Slingshot Program: Funding the Layer Above Foundation Models
00:07:56 JEPA and DSPy: Evolutionary Prompt Optimization
00:10:29 Beyond Berkeley and Stanford: Expanding the Research Network
00:13:22 NSF Complementarity: Not Broken, Just Insufficient
00:17:03 The Laud Lounge: Creating a VIP Experience at NeurIPS
00:18:41 Open Frontiers: Reclaiming Leadership in Open AI Research
00:19:06 The Open Research Crisis: Why [State of Research Funding] Beyond NSF, Slingshots, Open Frontiers — Andy Konwinski, Laude Institute](https://i.ytimg.com/vi/ZagdY6UJYL4/mqdefault.jpg)



