Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs @aiDotEngineer
Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs  @aiDotEngineer
Uploaded July 2026 | Updated September 2026, 3 weeks ago
Mahesh Sathiamoorthy's pitch is to stand in the researcher's shoes: the hard part of post-training is not the algorithm but the data and the environments that feed it. As agents get pushed to run autonomously for hours, something eventually falls over, and reinforcement learning is the tool for stretching that reliability, but RL environments are really just data in a different shape. Bespoke Labs works on curating both, from supervised fine-tuning sets to the environments models learn in.

He grounds it in OpenThoughts, the widely used reasoning dataset his team built, and the counterintuitive lessons that came out of curating it: diversity of reasoning traces matters, keeping multiple answers per question helps, and the obvious recipe often is not the best one. A favorite example is teaching a model to reason about credit card compliance, where fine-tuning on the right tagged data lifted the compliance metrics that a raw model kept getting wrong. The through line, supported by their Curator tooling, is that a disciplined curation stack, not just more compute, is what turns a base model into a capable post-trained one.

Speaker info:
- https://x.com/madiator
- linkedin.com/in/smaheswaran
- smahesh.com

Timestamps:
0:00 - Standing in the researcher's shoes
1:30 - Post-training at Bespoke Labs
3:13 - When agents fall over on long tasks
4:44 - RL environments as data
6:29 - Building OpenThoughts
7:36 - Finding a curation recipe
10:27 - Counterintuitive lessons
13:49 - A credit card compliance example
16:13 - Curating reasoning data with Curator
17:16 - The full curation stack
Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke LabsPerceptual Evaluations: Evals for Aesthetics — Diego Rodriguez, Krea.aix402 isn’t good (yet) — Jan Curn, ApifyAI Agents Are Just Distributed Systems Now — Salman Munaf, TikTokWhy Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, GoogleVoiceVision RAG - Integrating Visual Document Intelligence with Voice Response — Suman Debnath, AWSWe Vetted 2000 AI Skills Before They Reached Developers — Lucas Palma, NubankRealtime multiplayer, automation, and you! — Idan Gazit, GitHubShip Production Software in Minutes, Not Months — Eno Reyes, FactoryFull Workshop: Setting Yourself Up for Success —Jason Liu, OpenAI CodexBeyond Static Intelligence: Evaluating Continual Learning — Parth Asawa, UC BerkeleyDesigning Agents (The Floor Is the Frontier) — Ben Hylak, Raindrop
AI Engineer |

Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER