Uploaded November 2025 | Updated September 2026, 2 weeks ago
In this episode of Latent Space, Matthias Wagner, CEO & co-founder of Flux, reveals how they're revolutionizing hardware design with AI agents that can transform product briefs into manufacturable PCB designs in under 30 minutes. Building what he calls "the AI hardware engineer," Flux addresses a glaring gap: while software development tooling has transformed dramatically over decades, hardware design tools remained stuck in time - until now.
*From Burning Man to Building the Future of Hardware* Matthias's journey to founding Flux began unconventionally - taking a summer off after leaving Meta in 2019 to work on Burning Man projects reignited his hardware passion and exposed the stark tooling gap between software and hardware development. Despite supply chains evolving to enable individual makers to manufacture almost anything, the design tools hadn't kept pace. Flux set out to change that, initially building a browser-based, collaborative CAD tool from scratch - similar to Figma's approach - but architected from day one as a reinforcement learning environment for AI agents.
*The LLM Revolution and Tool Calling Breakthrough* While Flux started with machine learning approaches in 2019, the emergence of LLMs in 2022 turbocharged their vision. Matthias reveals they were likely the first engineering design tool to ship AI chat capabilities, even before GPT-4's public release. The real breakthrough came when tool calling became reliable about a year ago, enabling their agents to search component libraries, check pricing and availability across distributors, and execute complex design tasks autonomously.
*Live Demo: From Voice Assistant Brief to PCB in Real-Time* During the episode, Matthias demonstrates Flux's capabilities by designing a custom Alexa-like device from scratch - complete with ESP32 microcontroller, beamforming microphones, OLED display, speaker, battery management, and Wi-Fi connectivity. The agent autonomously searches through millions of components, checks real-time availability and pricing from distributors like DigiKey and Arrow, ensures compatibility, and generates a manufacturable design - all while explaining its decisions and accepting user feedback.
*The AI Engineering Stack and Iteration Process* Flux's technical approach layers multiple specialized agents using LangGraph and LangChain, with prompts managed externally in Langsmith for rapid iteration. Matthias candidly discusses the challenges of prompt management across numerous sub-agents, the balance between evals and "vibe checking," and why they prioritize iteration speed over premature abstraction. Their average user session runs 25 minutes, with agents handling everything from component selection to routing optimization.
*7,000 Paying Customers and 26x Growth* Flux has achieved remarkable traction with 7,000 paying customers - a 26x year-over-year growth - all through organic channels. Their users range from hobbyists to Fortune 10 companies, with everyone from vending machine manufacturers to traffic light companies adopting the platform. The vision extends far beyond PCBs: Matthias envisions a future where you can "prompt a smartphone into existence," fundamentally disrupting the OEM model by making custom hardware as accessible as generating text with ChatGPT. The conversation also explores how Flux integrates with manufacturing partners, manages complex supply chain data, and why the shift from mass production to on-demand custom hardware is becoming economically viable when AI eliminates design costs, leaving only material expenses. 00:00:00 Introduction and Building the AI Hardware Engineer
00:00:43 From Meta to Flux: The Hardware Tooling Problem
00:02:16 Pre-LLM Vision to AI-Powered Design
00:04:14 Early AI Chat Implementation and Tool Calling Evolution
00:06:28 User Expectations Shift: From Features to Agents
00:07:50 Live Demo: Building an Alexa-like Device
00:09:40 Supply Chain Integration and Component Pricing
00:11:04 AI Agent Finding Component Alternatives
00:15:08 Knowledge Base and Personalization System
00:17:33 Creating a Voice Assistant from Scratch
00:19:35 Agent Planning and Execution Process
00:23:17 Devon Integration and Computer Use Discussion
00:28:33 Component Library and User-Generated Content
00:31:14 AI Engineering Stack: LangChain and LangGraph
00:33:50 Prompt Management Challenges and Solutions
00:40:46 Business Growth and Market Vision
00:43:02 The Future of Personalized Manufacturing
00:45:40 Project Review and Next Steps
In this episode of Latent Space, Matthias Wagner, CEO & co-founder of Flux, reveals how they're revolutionizing hardware design with AI agents that can transform product briefs into manufacturable PCB designs in under 30 minutes. Building what he calls "the AI hardware engineer," Flux addresses a glaring gap: while software development tooling has transformed dramatically over decades, hardware design tools remained stuck in time - until now.
*From Burning Man to Building the Future of Hardware* Matthias's journey to founding Flux began unconventionally - taking a summer off after leaving Meta in 2019 to work on Burning Man projects reignited his hardware passion and exposed the stark tooling gap between software and hardware development. Despite supply chains evolving to enable individual makers to manufacture almost anything, the design tools hadn't kept pace. Flux set out to change that, initially building a browser-based, collaborative CAD tool from scratch - similar to Figma's approach - but architected from day one as a reinforcement learning environment for AI agents.
*The LLM Revolution and Tool Calling Breakthrough* While Flux started with machine learning approaches in 2019, the emergence of LLMs in 2022 turbocharged their vision. Matthias reveals they were likely the first engineering design tool to ship AI chat capabilities, even before GPT-4's public release. The real breakthrough came when tool calling became reliable about a year ago, enabling their agents to search component libraries, check pricing and availability across distributors, and execute complex design tasks autonomously.
*Live Demo: From Voice Assistant Brief to PCB in Real-Time* During the episode, Matthias demonstrates Flux's capabilities by designing a custom Alexa-like device from scratch - complete with ESP32 microcontroller, beamforming microphones, OLED display, speaker, battery management, and Wi-Fi connectivity. The agent autonomously searches through millions of components, checks real-time availability and pricing from distributors like DigiKey and Arrow, ensures compatibility, and generates a manufacturable design - all while explaining its decisions and accepting user feedback.
*The AI Engineering Stack and Iteration Process* Flux's technical approach layers multiple specialized agents using LangGraph and LangChain, with prompts managed externally in Langsmith for rapid iteration. Matthias candidly discusses the challenges of prompt management across numerous sub-agents, the balance between evals and "vibe checking," and why they prioritize iteration speed over premature abstraction. Their average user session runs 25 minutes, with agents handling everything from component selection to routing optimization.
*7,000 Paying Customers and 26x Growth* Flux has achieved remarkable traction with 7,000 paying customers - a 26x year-over-year growth - all through organic channels. Their users range from hobbyists to Fortune 10 companies, with everyone from vending machine manufacturers to traffic light companies adopting the platform. The vision extends far beyond PCBs: Matthias envisions a future where you can "prompt a smartphone into existence," fundamentally disrupting the OEM model by making custom hardware as accessible as generating text with ChatGPT. The conversation also explores how Flux integrates with manufacturing partners, manages complex supply chain data, and why the shift from mass production to on-demand custom hardware is becoming economically viable when AI eliminates design costs, leaving only material expenses. 00:00:00 Introduction and Building the AI Hardware Engineer
00:00:43 From Meta to Flux: The Hardware Tooling Problem
00:02:16 Pre-LLM Vision to AI-Powered Design
00:04:14 Early AI Chat Implementation and Tool Calling Evolution
00:06:28 User Expectations Shift: From Features to Agents
00:07:50 Live Demo: Building an Alexa-like Device
00:09:40 Supply Chain Integration and Component Pricing
00:11:04 AI Agent Finding Component Alternatives
00:15:08 Knowledge Base and Personalization System
00:17:33 Creating a Voice Assistant from Scratch
00:19:35 Agent Planning and Execution Process
00:23:17 Devon Integration and Computer Use Discussion
00:28:33 Component Library and User-Generated Content
00:31:14 AI Engineering Stack: LangChain and LangGraph
00:33:50 Prompt Management Challenges and Solutions
00:40:46 Business Growth and Market Vision
00:43:02 The Future of Personalized Manufacturing
00:45:40 Project Review and Next Steps



![Scaling Past Informal AI - Carina Hong, Axiom Math
Carina Hong, founder and CEO of Axiom Math, joins the AI for Science podcast right after closing a $200M Series A to argue that the road to superintelligence runs through formal verification — not as a bug fix, but as the only way to compound and scale AI brilliance. Her company, seven months old and 30 people strong, scored a perfect 120/120 on the 2024 Putnam exam, beating the best human and every other AI system at the time. We dig into the Lean theorem prover, why verified generation gives better training signal than informal RL, the hard specification problem, and why Carina believes an informal system alone can never reach math AGI.
00:00 — [INTRO — spliced from final take at 01:47:28]
00:52 — The $200M Series A and the Math Startup Thesis
04:52 — Verified AI: Scaling Brilliance, Not Fixing Lousiness
13:42 — Axioms System: Lean Data, RL, and the Putnam Perfect Score
22:12 — Mathematical Discovery — Before the Conjecture
25:12 — Rices Theorem, Incompleteness, and Practical Limits
30:42 — Code With Proof — The Verina Benchmark
37:57 — Proof Trees, Context Windows, and Scaling Limits
43:57 — Markets, Moat, and the Business Case ($1.6B valuation)
55:27 — Personal Origin Story: Oxford, UCL Gatsby, Stanford Law
01:00:57 — The Erdos Controversy and the Difficulty of Search
01:06:02 — AlphaZero for Math, Self-Improvement
01:08:47 — Startup Advantage and the OpenAI GPTF Thread
01:13:17 — Axle API — Open Infrastructure for Lean at Scale
01:20:47 — Collaboration, Polymath, and Human Attention as the Bottleneck
01:22:21 — Founding Story — Obsession, Law School, and Julie Zhuo
01:26:17 — The Bigger Vision — AGI, Science, and Transfer Learning
01:35:02 — Bottlenecks, Fragmentation, and the Fields Future Scaling Past Informal AI - Carina Hong, Axiom Math](https://i.ytimg.com/vi/abYcV5LHMG4/mqdefault.jpg)


![[State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
From pre-training data curation to shipping *GPT-4o,* *o1,* *o3,* and now *GPT-5 thinking* and the *shopping model,* *Josh McGrath* has lived through the full arc of OpenAIs post-training evolution—from the PPO vs DPO debates of 2023 to todays RLVR era, where the real innovation isnt optimization methods but *data quality, signal trust, and token efficiency.* We sat down with Josh at *NeurIPS 2025* to dig into the state of post-training heading into 2026: why RLHF and RLVR are both just policy gradient methods (the difference is the input data, not the math), how *GRPO* from DeepSeek Math was underappreciated as a shift toward more trustworthy reward signals (math answers you can verify vs. human preference you cant), why *token efficiency* matters more than wall-clock time (GPT-5 to 5.1 bumped evals _and_ slashed tokens), how *Codex* has changed his workflow so much he feels trapped by 40-minute design sessions followed by 15-minute agent sprints, the infrastructure chaos of scaling RL (way more moving parts than pre-training), why *long context* will keep climbing but agents + graph walks might matter more than 10M-token windows, the *shopping model* as a test bed for interruptability and chain-of-thought transparency, why *personality toggles* (Anton vs Clippy) are a real differentiator users care about, and his thesis that the education system isnt producing enough people who can do *both distributed systems and ML research*—the exact skill set required to push the frontier when the bottleneck moves every few weeks.
We discuss:
* Joshs path: *pre-training data curation → post-training researcher at OpenAI,* shipping GPT-4o, o1, o3, GPT-5 thinking, and the shopping model
* Why he switched from pre-training to post-training: Do I want to make 3% compute efficiency wins, or change behavior by 40%?
* The *RL infrastructure challenge:* way more moving parts than pre-training (tasks, grading setups, external partners), and why babysitting runs at 12:30am means jumping into unfamiliar code constantly
* How *Codex* has changed his workflow: 40-minute design sessions compressed into 15-minute agent sprints, and the strange trapped feeling of waiting for the agent to finish
* The *RLHF vs RLVR debate:* both are policy gradient methods, the real difference is *data quality and signal trust* (human preference vs. verifiable correctness)
* Why *GRPO* (from DeepSeek Math) was underappreciated: not just an optimization trick, but a shift toward reward signals you can actually trust (math answers over human vibes)
* The *token efficiency revolution:* GPT-5 to 5.1 bumped evals _and_ slashed tokens, and why thinking in tokens (not wall-clock time) unlocks better tool-calling and agent workflows
* *Personality toggles:* Anton (tool, no warmth) vs Clippy (friendly, helpful), and why Josh uses custom instructions to make his model just a tool
* The *router problem:* having a router at the top (GPT-5 thinking vs non-thinking) _and_ an implicit router (thinking effort slider) creates weird bumps, and why the abstractions will eventually merge
* *Long context:* climbing Graph Blocks evals, the dream of 10M+ token windows, and why agents + graph walks might matter more than raw context length
* Why the education system isnt producing enough people who can do *both distributed systems and ML research,* and why thats the bottleneck for frontier labs
* The 2026 vision: *neither pre-training nor post-training is dead,* were in the fog of war, and the bottleneck will keep moving (so emotional stability helps)
—
Josh McGrath
* OpenAI: https://openai.com
* https://x.com/j_mcgraph
00:00:00 Introduction: Josh McGrath on Post-Training at OpenAI
00:04:37 The Shopping Model: Black Friday Launch and Interruptability
00:07:11 Model Personality and the Anton vs Clippy Divide
00:08:26 Beyond PPO vs DPO: The Data Quality Spectrum in RL
00:01:40 Infrastructure Challenges: Why Post-Training RL is Harder Than Pre-Training
00:13:12 Token Efficiency: The 2D Plot That Matters Most
00:03:45 Codex Max and the Flow Problem: 40 Minutes of Planning, 15 Minutes of Waiting
00:17:29 Long Context and Graph Blocks: Climbing Toward Perfect Context
00:21:23 The ML-Systems Hybrid: Whats Hard to Hire For
00:24:50 Pre-Training Isnt Dead: Living Through Technological Revolution [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI](https://i.ytimg.com/vi/botHQ7u6-Jk/mqdefault.jpg)



