Uploaded April 2026 | Updated September 2026, 2 weeks ago
From building Applied Intuition from YC-era autonomy tooling into a $15B physical AI company, Qasar Younis and Peter Ludwig have spent the last decade living through the full arc of autonomy: from simulation and data infrastructure for robotaxi companies, to operating systems for safety-critical machines, to deploying AI onto cars, trucks, mining equipment, construction vehicles, agriculture, defense systems, and driverless L4 trucks running in Japan today. They join us to explain why “physical AI” is not just LLMs on wheels, why the real bottleneck is no longer model intelligence but deployment onto constrained hardware, and why the future of autonomy may look less like one-off demos and more like Android for every moving machine.
We discuss:
• Applied Intuition’s mission: building physical AI for a safer, more prosperous world, powering cars, trucks, construction and mining equipment, agriculture, defense, and other moving machines
• Why physical AI is different from screen-based AI: learned systems can make mistakes in chat or coding, but safety-critical machines like driverless trucks, autonomous vehicles, and robots need much higher reliability
• The evolution from autonomy tooling to a broad physical AI platform: starting with simulation and data infrastructure for robotaxi companies, then expanding into 30+ products across simulation, operating systems, autonomy, and AI models
• The three core buckets of Applied Intuition’s technology: simulation and RL infrastructure, true operating systems for vehicles and machines, and fundamental AI models for autonomy and world understanding
• Why vehicles need a real AI operating system: real-time control, sensor streaming, latency, memory management, fail-safes, reliable updates, and why “bricking a car” is much worse than bricking an iPad
• How open the platform is: customers can use Applied’s autonomy stack, operating system, developer tools, or mix and match with their own systems
• Coding agents inside Applied Intuition: Cursor, Claude Code, internal adoption leaderboards, and how AI tools are changing engineering workflows even in embedded systems and safety-critical software
• Cruise, Waymo, and public trust: Qasar and Peter discuss why autonomy failures are not just technical issues, how companies interact with regulators, and why Waymo is setting a high bar for the industry
• Simulation vs. reality: why no simulator perfectly represents the real world, how sim-to-real validation works, and why real-world testing will never disappear
• World models for physical AI: hydroplaning, construction equipment, visual cues, cause-and-effect learning, and where world models help versus where they are not enough
• Why robotics demos are not production: the brittle last 1%, humanoid reliability, China’s humanoid marathon, DARPA Grand Challenge-style prize policy, and the advanced engineering gap between research and deployment
• Applied Intuition’s hard-earned lessons: after nearly a decade, Peter says they can look at a robotics demo and predict the next 20 problems the company will hit
• Qasar’s advice to founders: constrain the commercial problem, avoid copying mature-company strategies too early, and remember that compounding technology only matters if you survive long enough to see it compound
Applied Intuition:
• YouTube: youtube.com/@AppliedIntuitionInc
• X: https://x.com/AppliedInt
• LinkedIn: linkedin.com/company/applied-intuition-inc
Qasar Younis:
• X: https://x.com/qasar
• LinkedIn: linkedin.com/in/qasar
Peter Ludwig:
• LinkedIn: linkedin.com/in/peterwludwig
00:00:00 Cold Open: Physical Machines Before Android
00:01:52 Introduction: Applied Intuition’s Founders
00:02:28 What Applied Intuition Builds Today
00:03:23 Physical AI Beyond Screens
00:04:25 From YC Autonomy Tooling to 30+ Products
00:09:40 Simulation, Operating Systems, and AI Models
00:13:55 Sensors, Lidar, and Production Hardware
00:16:12 Why Vehicles Need a Real AI Operating System
00:19:27 The Android Analogy for Physical Machines
00:23:29 Coding Agents Inside Applied Intuition
00:25:43 How AI Changes Engineering Hiring
00:28:27 Evals, RL, and Neural Simulation
00:31:05 From Binary Tests to Statistical Safety
00:34:19 Cruise, Waymo, and Public Trust
00:37:21 Sim-to-Real Gaps and Robot Overheating
00:42:05 World Models and Hydroplaning
00:45:05 Onboard vs. Offboard AI Models
00:46:49 Why Deployment Is the Bottleneck
00:50:04 Local AI, RTK GPS, and Legacy Autonomy
00:52:43 Plan Mode for Physical Autonomy
00:54:39 Why Robotics Demos Aren’t Production
00:58:46 Founder Advice: Constraints and Compounding Tech
01:04:13 Why 2014 YC Advice Doesn’t Apply in 2026
01:06:09 Open Problems: Efficient Models and Safety Evals
01:07:26 Hiring Engineers at Applied Intuition
01:11:53 The Engineering Mindset
01:13:54 Closing
From building Applied Intuition from YC-era autonomy tooling into a $15B physical AI company, Qasar Younis and Peter Ludwig have spent the last decade living through the full arc of autonomy: from simulation and data infrastructure for robotaxi companies, to operating systems for safety-critical machines, to deploying AI onto cars, trucks, mining equipment, construction vehicles, agriculture, defense systems, and driverless L4 trucks running in Japan today. They join us to explain why “physical AI” is not just LLMs on wheels, why the real bottleneck is no longer model intelligence but deployment onto constrained hardware, and why the future of autonomy may look less like one-off demos and more like Android for every moving machine.
We discuss:
• Applied Intuition’s mission: building physical AI for a safer, more prosperous world, powering cars, trucks, construction and mining equipment, agriculture, defense, and other moving machines
• Why physical AI is different from screen-based AI: learned systems can make mistakes in chat or coding, but safety-critical machines like driverless trucks, autonomous vehicles, and robots need much higher reliability
• The evolution from autonomy tooling to a broad physical AI platform: starting with simulation and data infrastructure for robotaxi companies, then expanding into 30+ products across simulation, operating systems, autonomy, and AI models
• The three core buckets of Applied Intuition’s technology: simulation and RL infrastructure, true operating systems for vehicles and machines, and fundamental AI models for autonomy and world understanding
• Why vehicles need a real AI operating system: real-time control, sensor streaming, latency, memory management, fail-safes, reliable updates, and why “bricking a car” is much worse than bricking an iPad
• How open the platform is: customers can use Applied’s autonomy stack, operating system, developer tools, or mix and match with their own systems
• Coding agents inside Applied Intuition: Cursor, Claude Code, internal adoption leaderboards, and how AI tools are changing engineering workflows even in embedded systems and safety-critical software
• Cruise, Waymo, and public trust: Qasar and Peter discuss why autonomy failures are not just technical issues, how companies interact with regulators, and why Waymo is setting a high bar for the industry
• Simulation vs. reality: why no simulator perfectly represents the real world, how sim-to-real validation works, and why real-world testing will never disappear
• World models for physical AI: hydroplaning, construction equipment, visual cues, cause-and-effect learning, and where world models help versus where they are not enough
• Why robotics demos are not production: the brittle last 1%, humanoid reliability, China’s humanoid marathon, DARPA Grand Challenge-style prize policy, and the advanced engineering gap between research and deployment
• Applied Intuition’s hard-earned lessons: after nearly a decade, Peter says they can look at a robotics demo and predict the next 20 problems the company will hit
• Qasar’s advice to founders: constrain the commercial problem, avoid copying mature-company strategies too early, and remember that compounding technology only matters if you survive long enough to see it compound
Applied Intuition:
• YouTube: youtube.com/@AppliedIntuitionInc
• X: https://x.com/AppliedInt
• LinkedIn: linkedin.com/company/applied-intuition-inc
Qasar Younis:
• X: https://x.com/qasar
• LinkedIn: linkedin.com/in/qasar
Peter Ludwig:
• LinkedIn: linkedin.com/in/peterwludwig
00:00:00 Cold Open: Physical Machines Before Android
00:01:52 Introduction: Applied Intuition’s Founders
00:02:28 What Applied Intuition Builds Today
00:03:23 Physical AI Beyond Screens
00:04:25 From YC Autonomy Tooling to 30+ Products
00:09:40 Simulation, Operating Systems, and AI Models
00:13:55 Sensors, Lidar, and Production Hardware
00:16:12 Why Vehicles Need a Real AI Operating System
00:19:27 The Android Analogy for Physical Machines
00:23:29 Coding Agents Inside Applied Intuition
00:25:43 How AI Changes Engineering Hiring
00:28:27 Evals, RL, and Neural Simulation
00:31:05 From Binary Tests to Statistical Safety
00:34:19 Cruise, Waymo, and Public Trust
00:37:21 Sim-to-Real Gaps and Robot Overheating
00:42:05 World Models and Hydroplaning
00:45:05 Onboard vs. Offboard AI Models
00:46:49 Why Deployment Is the Bottleneck
00:50:04 Local AI, RTK GPS, and Legacy Autonomy
00:52:43 Plan Mode for Physical Autonomy
00:54:39 Why Robotics Demos Aren’t Production
00:58:46 Founder Advice: Constraints and Compounding Tech
01:04:13 Why 2014 YC Advice Doesn’t Apply in 2026
01:06:09 Open Problems: Efficient Models and Safety Evals
01:07:26 Hiring Engineers at Applied Intuition
01:11:53 The Engineering Mindset
01:13:54 Closing





![[State of Context Engineering] Agentic RAG, Context Rot, MCP, Subagents — Nina Lopatina, Contextual
From neuroscience PhD research on reward learning and decision making to building the infrastructure for *context engineering at scale,* *Nina Lopatina* has spent the last year watching a brand-new category emerge from prototype to production—and now shes leading the charge to turn context engineering from a collection of design patterns into a *full-stack discipline* with benchmarks, tooling, and real-world deployment at enterprise scale. We caught up with Nina live at *NeurIPS 2025* (her fifth!) to dig into the state of context engineering heading into 2026: why this year felt like *six months compressed into a year* (the category only really took hold in mid-2024), how *agentic RAG is now the baseline* (query reformulation into subqueries improved performance so dramatically it became the new standard), why *context rot is cited in every blog* but industry benchmarks at real scale (100k+ documents, billions of tokens) are still rare, how *MCP is both a driver and a flaw* for context engineering (giant JSON tool definitions stuff the context window, but MCP servers unlock rapid prototyping before you optimize down to direct API calls), the rise of *sub-agents with turn limits and explicit constraints* (unlimited agency degrades performance and causes hallucinations), why *instruction-following re-rankers* are critical for scaling retrieval across massive databases (more recall up front, more precision in the final context window), how *benchmarks are being saturated faster than ever* (Claude Code just saturated a Princeton benchmark released in October, with solutions so good the gold dataset had errors), the *KV cache decision-making framework* for multi-turn agents (stuff that doesnt change goes up front, stuff that changes a lot goes at the bottom), why shes *embodied-evaling frontier models as a snowboarding coach* (training for a 25-lap mogul race over 3–4 months, and why she had to close the window and restart because the model lost training context), and her thesis that 2026 will be the year context engineering moves from *component-level innovation to full-system design patterns*—where the conversation shifts from how do I optimize my re-ranker to what does the end-to-end architecture look like for reasoning over billions of tokens in production?
We discuss:
* What Contextual does: *end-to-end platform for context engineering across domains* (code, legal, retail, e-commerce, support), with multimodal ingestion, hybrid search, re-rankers, and dynamic agents
* The *first instruction-following re-ranker* (launched March 2024): latency is the biggest complaint, but for dynamic agents (where latency is less sensitive), its a game-changer for reasoning over large databases
* Why *agentic RAG is now the baseline:* query reformulation into subqueries improved performance so dramatically it became the new standard (normal RAG is dead)
* The *context engineering hackathon* (Retail Universe, ~100k documents, PDFs/CSVs/logs): Ninas team used a dynamic agent with turn limits and explicit constraints to avoid infinite sub-agent loops
* *Context rot:* everyone cites it, but Anthropics work putting numbers on it (e.g., at 700k tokens in a 1M context window, retrieval drops to 30%) is what made it actionable
* *Sub-agents with turn limits:* unlimited agency degrades performance and causes hallucinations, so explicit constraints (turn limits, validation loops) are critical for scale
* The need for *industry benchmarks at real scale:* most benchmarks use toy datasets, but the Retail Universe hackathon dataset (100k+ documents, billions of tokens) is closer to production reality
* *KV cache decision-making:* stuff that doesnt change (system prompt, early turns) goes up front, stuff that changes a lot (recent turns, dynamic context) goes at the bottom—critical for multi-turn agents
* Why *intentional context compression* matters: models arent great at compaction yet, so Nina proactively limits turns (even in Cursor, she opens a new window mid-conversation to avoid context loss)
—
Nina Lopatina
* Contextual AI: https://contextual.ai
* X: https://x.com/ninalopatina
* LinkedIn: https://linkedin.com/in/ninalopatina
00:00:00 Introduction: Nina Lopatina on Context Engineering at NeurIPS
00:04:34 The Death of Normal RAG: Rise of Agentic RAG and Query Reformulation
00:06:20 Sub-Agents and Turn Limits: Lessons from the Retail Universe Hackathon
00:09:07 Context Engineering in 2024: Design Patterns and the Prototyping Stage
00:10:17 Benchmarks and Scale: From Princeton HOW to Saturated Research Tasks
00:12:52 Context Rot, MCP, and Tool Selection Challenges
00:17:28 Prompt Optimization: Jeppa, ACE, and Evolutionary Approaches
00:19:42 KV Cache Strategy and Multi-Turn Agent Stability
00:22:30 Domain Generalization: Code, Legal, Retail, and Beyond
00:23:59 Predictions and Full System Design: The Future of Context Engineering [State of Context Engineering] Agentic RAG, Context Rot, MCP, Subagents — Nina Lopatina, Contextual](https://i.ytimg.com/vi/tSRqTerZrH8/mqdefault.jpg)




