Uploaded February 2026 | Updated September 2026, 2 weeks ago
This is a clip from https://www.latent.space/p/paid-anthropic-distillation-and-how?utm_source=youtube_shorts
See the full video: youtube.com/watch?v=7EBQ04OL-is
#shorts #substack
This is a clip from https://www.latent.space/p/paid-anthropic-distillation-and-how?utm_source=youtube_shorts
See the full video: youtube.com/watch?v=7EBQ04OL-is
#shorts #substack

![Scaling Past Informal AI - Carina Hong, Axiom Math
Carina Hong, founder and CEO of Axiom Math, joins the AI for Science podcast right after closing a $200M Series A to argue that the road to superintelligence runs through formal verification — not as a bug fix, but as the only way to compound and scale AI brilliance. Her company, seven months old and 30 people strong, scored a perfect 120/120 on the 2024 Putnam exam, beating the best human and every other AI system at the time. We dig into the Lean theorem prover, why verified generation gives better training signal than informal RL, the hard specification problem, and why Carina believes an informal system alone can never reach math AGI.
00:00 — [INTRO — spliced from final take at 01:47:28]
00:52 — The $200M Series A and the Math Startup Thesis
04:52 — Verified AI: Scaling Brilliance, Not Fixing Lousiness
13:42 — Axioms System: Lean Data, RL, and the Putnam Perfect Score
22:12 — Mathematical Discovery — Before the Conjecture
25:12 — Rices Theorem, Incompleteness, and Practical Limits
30:42 — Code With Proof — The Verina Benchmark
37:57 — Proof Trees, Context Windows, and Scaling Limits
43:57 — Markets, Moat, and the Business Case ($1.6B valuation)
55:27 — Personal Origin Story: Oxford, UCL Gatsby, Stanford Law
01:00:57 — The Erdos Controversy and the Difficulty of Search
01:06:02 — AlphaZero for Math, Self-Improvement
01:08:47 — Startup Advantage and the OpenAI GPTF Thread
01:13:17 — Axle API — Open Infrastructure for Lean at Scale
01:20:47 — Collaboration, Polymath, and Human Attention as the Bottleneck
01:22:21 — Founding Story — Obsession, Law School, and Julie Zhuo
01:26:17 — The Bigger Vision — AGI, Science, and Transfer Learning
01:35:02 — Bottlenecks, Fragmentation, and the Fields Future Scaling Past Informal AI - Carina Hong, Axiom Math](https://i.ytimg.com/vi/abYcV5LHMG4/mqdefault.jpg)


![[State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
From pre-training data curation to shipping *GPT-4o,* *o1,* *o3,* and now *GPT-5 thinking* and the *shopping model,* *Josh McGrath* has lived through the full arc of OpenAIs post-training evolution—from the PPO vs DPO debates of 2023 to todays RLVR era, where the real innovation isnt optimization methods but *data quality, signal trust, and token efficiency.* We sat down with Josh at *NeurIPS 2025* to dig into the state of post-training heading into 2026: why RLHF and RLVR are both just policy gradient methods (the difference is the input data, not the math), how *GRPO* from DeepSeek Math was underappreciated as a shift toward more trustworthy reward signals (math answers you can verify vs. human preference you cant), why *token efficiency* matters more than wall-clock time (GPT-5 to 5.1 bumped evals _and_ slashed tokens), how *Codex* has changed his workflow so much he feels trapped by 40-minute design sessions followed by 15-minute agent sprints, the infrastructure chaos of scaling RL (way more moving parts than pre-training), why *long context* will keep climbing but agents + graph walks might matter more than 10M-token windows, the *shopping model* as a test bed for interruptability and chain-of-thought transparency, why *personality toggles* (Anton vs Clippy) are a real differentiator users care about, and his thesis that the education system isnt producing enough people who can do *both distributed systems and ML research*—the exact skill set required to push the frontier when the bottleneck moves every few weeks.
We discuss:
* Joshs path: *pre-training data curation → post-training researcher at OpenAI,* shipping GPT-4o, o1, o3, GPT-5 thinking, and the shopping model
* Why he switched from pre-training to post-training: Do I want to make 3% compute efficiency wins, or change behavior by 40%?
* The *RL infrastructure challenge:* way more moving parts than pre-training (tasks, grading setups, external partners), and why babysitting runs at 12:30am means jumping into unfamiliar code constantly
* How *Codex* has changed his workflow: 40-minute design sessions compressed into 15-minute agent sprints, and the strange trapped feeling of waiting for the agent to finish
* The *RLHF vs RLVR debate:* both are policy gradient methods, the real difference is *data quality and signal trust* (human preference vs. verifiable correctness)
* Why *GRPO* (from DeepSeek Math) was underappreciated: not just an optimization trick, but a shift toward reward signals you can actually trust (math answers over human vibes)
* The *token efficiency revolution:* GPT-5 to 5.1 bumped evals _and_ slashed tokens, and why thinking in tokens (not wall-clock time) unlocks better tool-calling and agent workflows
* *Personality toggles:* Anton (tool, no warmth) vs Clippy (friendly, helpful), and why Josh uses custom instructions to make his model just a tool
* The *router problem:* having a router at the top (GPT-5 thinking vs non-thinking) _and_ an implicit router (thinking effort slider) creates weird bumps, and why the abstractions will eventually merge
* *Long context:* climbing Graph Blocks evals, the dream of 10M+ token windows, and why agents + graph walks might matter more than raw context length
* Why the education system isnt producing enough people who can do *both distributed systems and ML research,* and why thats the bottleneck for frontier labs
* The 2026 vision: *neither pre-training nor post-training is dead,* were in the fog of war, and the bottleneck will keep moving (so emotional stability helps)
—
Josh McGrath
* OpenAI: https://openai.com
* https://x.com/j_mcgraph
00:00:00 Introduction: Josh McGrath on Post-Training at OpenAI
00:04:37 The Shopping Model: Black Friday Launch and Interruptability
00:07:11 Model Personality and the Anton vs Clippy Divide
00:08:26 Beyond PPO vs DPO: The Data Quality Spectrum in RL
00:01:40 Infrastructure Challenges: Why Post-Training RL is Harder Than Pre-Training
00:13:12 Token Efficiency: The 2D Plot That Matters Most
00:03:45 Codex Max and the Flow Problem: 40 Minutes of Planning, 15 Minutes of Waiting
00:17:29 Long Context and Graph Blocks: Climbing Toward Perfect Context
00:21:23 The ML-Systems Hybrid: Whats Hard to Hire For
00:24:50 Pre-Training Isnt Dead: Living Through Technological Revolution [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI](https://i.ytimg.com/vi/botHQ7u6-Jk/mqdefault.jpg)





