The Pre-trainers toolkit: From dataset construction to model scaling @allenai
The Pre-trainers toolkit: From dataset construction to model scaling  @allenai
Uploaded May 2024 | Updated September 2026, 9 hours ago
Abstract: Recent breakthroughs in machine learning rely heavily on pre-training techniques, harnessing larger datasets, models, and computational resources to create base-models for subsequent fine-tuning. In this talk, we develop a pre-training toolkit. Drawing from empirical findings, we present methodologies for dataset construction and de-risking large-scale model training. Our discussion touches on both multimodal and language modeling domains. By addressing the entire pre-training pipeline, from dataset creation to downstream evaluation, we aim to create better, more reliable models.

Bio: Samir Yitzhak Gadre (Samir) is an NSF graduate research fellow and PhD student at Columbia University working with Shuran Song and Ludwig Schmidt. He studies the empirical foundations of pre-training. Samir also serves as a core maintainer of OpenLM, a minimal but performative open-source language modeling library.
The Pre-trainers toolkit: From dataset construction to model scalingFrom Compression to Convection: A Latent Variable PerspectiveMovement Primitives as Action Sequence Models for Efficient Robot LearningRobot learning and perception for contact-rich manipulationAppWorld: Reliable Evaluation of Interactive Agents in a Controllable World of Apps and PeopleYou Cant Have AI Safety Without InclusionUsing MolmoWeb as a Claude Code SkillAutomated Scientific Discovery of Mind and BehaviorText Modular Networks: Learning to Decompose Tasks in the Language of Existing ModelsBeyond Test Accuracies for Studying Deep Neural NetworksSPOKE: A massive biomedical knowledge graph for precision health and drug discoveryAMA with AI Pioneers Raj Reddy and Andries Andy van Dam
Ai2 |

The Pre-trainer's toolkit: From dataset construction to model scaling

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER