humans& with Alexis Ross, Manya Bansal, and Niloofar Mireshghallah - Weaviate Podcast #145! @Weaviate
humans& with Alexis Ross, Manya Bansal, and Niloofar Mireshghallah - Weaviate Podcast #145!  @Weaviate
Uploaded September 2026 | Updated September 2026, 35 minutes ago
Alexis, Manya, and Niloofar from humans& join the Weaviate Podcast to introduce Persimmon, a user model built to simulate how humans actually behave in multi-turn, multi-party conversations. Persimmon is explicitly not an assistant, a companion, or a Character AI-style stand-in, it is a research preview aimed at faithfully capturing the distribution of human behavior. This includes the natural friction of frustration, excitement, and group dynamics that assistant chatbots trained to be helpful never exhibit.

Alexis, Manya, and Niloofar bring a striking mix of backgrounds to the problem: AI tutoring and student modeling, programming languages for high-performance computing, and privacy and information-flow research at Carnegie Mellon. The conversation opens with whether the Turing test is solved. Humans& runs a distributionally grounded, multi-turn version where the judge sees many examples of human and AI behavior. Frontier models fool it less than 5% of the time, while Persimmon reaches roughly 20% against a 50% ceiling. From there, the discussion dives into training for non-verifiable tasks: why rubrics-as-rewards approaches invite reward hacking, why the team refuses to impose its own theory of human behavior, and how distribution matching, with the multi-turn Turing test as a North Star metric operating in an implicit feature space, rather than Earth mover's distance over hand-picked features anchors both training and evaluation. They walk through evaluating on real human interaction data like the TIDES meeting transcripts and the TutorMoments tutoring dataset, and why role-played or scripted dialogue doesn't count.

The discussion then moves into theory of mind and world models as twin goals, with Persimmon enabling multi-agent environments where assistants get realistic human feedback at training time. The podcast further covers the choice of NVIDIA's Nemotron 3 Ultra and why starting from a base model matters: post-training causes mode collapse, you can't prompt-optimize your way out of it, and injected randomness drifts away over long rollouts. The conversation lands on what excites each guest next: personalized tutors, models that balance overlapping human goals, and training paradigms with long-term social pressures.

Links:
Persimmon: persimmon.humansand.ai/blog
Persimmon launch thread from humans& on Twitter: https://x.com/humansand/status/2098115046438215791
Niloofar Mireshghallah: mireshghallah.github.io
Manya Bansal: manya-bansal.github.io
Alexis Ross: alexisjihyeross.github.io

Chapters
0:00 Welcome Niloofar, Manya, and Alexis!
3:42 An Overview of Persimmon
7:28 Solving the Turing Test
14:56 User Models and AGI
17:20 RL with Non-Verifiable Rewards
21:55 Distribution Matching
29:20 Collecting Human Data
33:52 Theory of Mind in AI
39:00 NVIDIA Nemotron 3 Ultra
45:09 Prompt Optimization
47:20 Exciting Directions for AI
humans& with Alexis Ross, Manya Bansal, and Niloofar Mireshghallah - Weaviate Podcast #145!Range IndexAI that optimizes itself? 🤔ANN Algorithms in a nutshellWeaviate 1.24 | Release AnnouncementUX? DX? Meet AX!AI-Powered Resumes with Super People & WeaviateExploring the AI Renaissance with Jodie: Weaviate, vector databases, and the future of AISWE-bench with John Yang and Carlos E. Jimenez - Weaviate Podcast #107!Embedding model evaluation & selection guideWhat does the future of AI look like in 2025?Unstructured Data Objects
Weaviate vector database |

humans& with Alexis Ross, Manya Bansal, and Niloofar Mireshghallah - Weaviate Podcast #145!

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER