Data dictionaries, parquet, & Claude | Hadley Wickham | Data Science Lab @PositPBC
Data dictionaries, parquet, & Claude | Hadley Wickham | Data Science Lab  @PositPBC
Uploaded June 2026 | Updated September 2026, 3 weeks ago
The Data Science Lab is a live weekly call. Register at pos.it/dslab! Discord invites go out each week on lives calls. We'd love to have you!

The Lab is an open, messy space for learning and asking questions. Think of it like pair coding with a friend or two. Learn something new, and share what you know to help others grow.

On this call, Libby Heeren is joined by Hadley Wickham, who walks through using data dictionaries with Claude Code to clean and document datasets effectively.

Hadley demonstrates a workflow using three files: a data cleaning script, a data dictionary in YAML format, and the final cleaned data as a Parquet file. He shows how Claude Code and MCP REPL can help generate and maintain data dictionaries that document what you know about your data, making it easier for both humans and AI agents to work with your datasets. Using the NYC elevators dataset as an example, he walks through data cleaning tasks like normalizing whitespace, handling missing values, fixing date formats, and investigating geocoding issues - all while keeping the data dictionary, cleaning script, and Parquet file in sync through git.

Hosting crew from Posit: Libby Heeren, Isabella Velasquez

Hadley Wickham's GitHub: github.com/hadley
Hadley Wickham's Bluesky: https://bsky.app/profile/hadley.nz
Hadley Wickham's LinkedIn: linkedin.com/in/hadleywickham


Resources mentioned in the video and chat:
πŸ”— MCP REPL: github.com/posit-dev/mcp-repl
πŸ”— Data Dictionary YAML Format Specification: github.com/hadley/data-dict.yaml
πŸ”— Parquet files in R for Data Science: r4ds.hadley.nz/arrow.html#sec-parquet
πŸ”— Elevators Dataset Used in Demo: github.com/EmilHvitfeldt/elevators
πŸ”— Pointblank Package for Data Validation: posit-dev.github.io/pointblank
πŸ”— Arrow R Book: arrowrbook.com
πŸ”— Monaspace Font Family (with ligatures): monaspace.githubnext.com
πŸ”— YAML Multiline Strings Reference: yaml-multiline.info
πŸ”— UBC Course on Shiny with RAG and Parquet: ubc-mds.github.io/DSCI_532_vis-2_book/060-03-rag.html
πŸ”— Tom Scott Video on Timezones: youtube.com/watch?v=-5wpm-gesOY
πŸ”— Falsehoods Programmers Believe About Names: kalzumeus.com/2010/06/17/falsehoods-programmers-believe-about-names
πŸ”— UTC Is Enough for Everyone Right: zachholman.com/talk/utc-is-enough-for-everyone-right
πŸ”— Daniel Chen's Cherry Blossom Analysis: chendaniely.github.io/posts/2026/2026-03-30-yvr-cherry-blossoms-marathon
πŸ”— Project Drawdown Climate Impact Explorer: drawdown.org/explorer
πŸ”— Green Coding Bookdown Resource: bookdown.org/content/d1e53ac9-28ce-472f-bc2c-f499f18264a3
πŸ”— IBM Green Coding Topics: ibm.com/think/topics/green-coding
πŸ”— Secret Elevator in Central Park Article: https://undercovernyc.home.blog/2021/02/08/a-secret-elevator-hidden-in-central-park/
πŸ”— Artificial Cave Beneath Central Park: gizmodo.com/an-artificial-cave-200-beneath-central-park-with-micha-1446538828

β–Ί Subscribe to Our Channel Here: bit.ly/2TzgcOu
Follow Us Here:
Website: posit.co
Hangout: https://pos.it/dsh
The Lab: https://pos.it/dslab
LinkedIn: linkedin.com/company/posit-software
Bluesky: https://bsky.app/profile/posit.co

Thanks for hanging out with us! πŸ’›

Timestamps of Questions / Topics:
00:00 Introduction
03:38 "Can you talk a little bit about what MCP is?"
05:42 Introducing the elevators dataset
07:00 Creating initial data dictionary with Claude
09:03 "Are there any cases where the CSV format is actually a better choice than Parquet?"
09:50 "What font do you use?"
12:02 Adding context from the readme to the data dictionary
14:02 "What is a good way to store data dictionaries along with datasets?"
14:38 "Is this particular data dict YAML format useful for projects with only one table of data?"
17:45 "Is Claude also going to decide when it's a good time to make a commit?"
24:02 "Does MCP REPL work on a Windows machine and can one use other AI LLM for example ChatGPT with it?"
26:14 Converting date columns to proper date types
27:16 "Do you see a meaningful distinction between data dicts and data contracts?"
29:36 "How do you view your approach to data dictionaries and the development of pointblank?"
34:16 Eliminating placeholder values and using proper missing values
36:25 "Could you show off a diff of a Parquet file?"
38:25 Investigating geocodes and creating a map of elevators
40:20 Using a leaflet map to explore Central Park elevators
42:03 "Which model is Hadley using?"
43:55 Discussion of cost consciousness and environmental impact of LLMs
46:30 "Is there a way to quantify environmental and electrical costs?"
48:50 The mystery elevator in Central Park
54:23 "How do you know this is actually faster or more productive than just writing the code?"
55:45 The importance of deep knowledge of data in qualitative work
Data dictionaries, parquet, & Claude | Hadley Wickham | Data Science LabWes McKinney: Part 2 β€” The open source hustle and an insider view of PositronEasiest way to get started on a Python project?!Posit + Snowflake Product Overview | Enterprise R & Python Data SciencePositron workflows that make life easier | Andrew Heiss | Data Science LabEnd-to-End Data Science Workflow with the Posit Team Snowflake Native AppWidgets Are Lego Bricks (and Other Things People Are Sleeping On) β€” with Vincent WarmerdamExploring Extensions from the Posit Connect Gallery (Toph Allen) | posit::conf(2025)Alenka Frim: What yoga teaches us about discipline and collaboration in data scienceBuilding healthy data communities | Rachael Dempsey & Libby Heeren | Data Science HangoutReal-World Evidence Workflows with PositShiny is... *too* easy? πŸ˜… Uh oh 🫣
Posit PBC |

Data dictionaries, parquet, & Claude | Hadley Wickham | Data Science Lab

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER