Uploaded July 2024 | Updated September 2026, 1 hour ago
Hey everyone! Thank you so much for watching the 101st episode of the Weaviate Podcast with Devin Petersohn! Devin is the creator of Modin, one of the world's most advanced systems for scaling Pandas! Devin then went onto co-found Ponder, which was acquired by Snowflake in early 2023. This was one of my favorite podcasts of all time, I learned so much about the internals of Data Systems and I hope you do as well!
Links:
Modin: github.com/modin-project/modin
Towards Scalable Dataframe Systems: arxiv.org/pdf/2001.00888
Additional publications from Devin: scholar.google.com/citations?user=lMAcwtwAAAAJ&hl=en
Ponder: ponder.io/blog
Devin at the CMU Database Group, Beyond SQL: Dataframes in the Database: https://db.cs.cmu.edu/events/spring-2024-beyond-sql-dataframes-in-the-database-devin-petersohn/
Chapters
0:00 Welcome Devin!
0:28 What lead you Scaling Dataframes?
3:55 What makes Pandas slower than SQL?
6:32 Separating the API from the Execution Engine
13:11 What is a Task Execution Engine?
15:50 Query Optimization
19:15 Materialized Views
22:45 File Formats
25:20 How to read CSVs faster?
29:58 gRPC, Serialization, and Apache Arrow
33:50 The Separation of Compute and Storage
36:50 CUDA Dataframes and RAPIDS
39:08 Ponder
43:50 Large Language Models
47:15 Thank you Devin!
Hey everyone! Thank you so much for watching the 101st episode of the Weaviate Podcast with Devin Petersohn! Devin is the creator of Modin, one of the world's most advanced systems for scaling Pandas! Devin then went onto co-found Ponder, which was acquired by Snowflake in early 2023. This was one of my favorite podcasts of all time, I learned so much about the internals of Data Systems and I hope you do as well!
Links:
Modin: github.com/modin-project/modin
Towards Scalable Dataframe Systems: arxiv.org/pdf/2001.00888
Additional publications from Devin: scholar.google.com/citations?user=lMAcwtwAAAAJ&hl=en
Ponder: ponder.io/blog
Devin at the CMU Database Group, Beyond SQL: Dataframes in the Database: https://db.cs.cmu.edu/events/spring-2024-beyond-sql-dataframes-in-the-database-devin-petersohn/
Chapters
0:00 Welcome Devin!
0:28 What lead you Scaling Dataframes?
3:55 What makes Pandas slower than SQL?
6:32 Separating the API from the Execution Engine
13:11 What is a Task Execution Engine?
15:50 Query Optimization
19:15 Materialized Views
22:45 File Formats
25:20 How to read CSVs faster?
29:58 gRPC, Serialization, and Apache Arrow
33:50 The Separation of Compute and Storage
36:50 CUDA Dataframes and RAPIDS
39:08 Ponder
43:50 Large Language Models
47:15 Thank you Devin!










- Star us on GitHub https://github.com/weaviate/weaviate
- Stay updated and subscribe to our newsletter: https://newsletter.weaviate.io/
- Try out Weaviate Cloud Services for free here: https://console.weaviate.cloud/
Got a question?
- Forum: https://forum.weaviate.io/
- Slack: https://weaviate.io/slack
Connect with us on
- Twitter: https://twitter.com/weaviate_io
- LinkedIn: https://www.linkedin.com/company/weaviate-io/
Tuana Socials
- Twitter/X: https://x.com/tuanacelik
- LinkedIn: https://www.linkedin.com/in/tuanacelik/ Weaviate Tech Hands-On: Query Agent](https://i.ytimg.com/vi/tqTFTS_E4i0/mqdefault.jpg)
