Deduplication of Large-scale Text Datasets for Pretraining of Language Models @allenai
Deduplication of Large-scale Text Datasets for Pretraining of Language Models  @allenai
Uploaded August 2024 | Updated September 2026, 1 day ago
In this talk, I'll cover the newly released DataComp for Language Models project, in which we generate a testbed for controlled experiments of building better datasets for pretraining language models in a compute-limited regime. From here I'll pivot to discussing one particular aspect of building better datasets: removing duplicates and near-duplicates from large corpuses of text, explaining several key techniques as well as our findings from extensive deduplication ablations. Finally, I'll raise some several open questions and future directions regarding deduplication of pretraining datasets, including some unpublished (but interesting!) results.
Deduplication of Large-scale Text Datasets for Pretraining of Language ModelsRobot Learning with Sparsity and ScarcityTowards Data-Driven Scientific Discovery with Generative AI: From Mathematical Modeling to LLMsWildDet3D - an open model for monocular 3D detectionDeepEarth: Multimodal Probabilistic World Model with 4D Spacetime EmbeddingEnhancing Reasoning in Smaller Models through Self-TrainingMaking Health Knowledge Accessible Through Personalized Language ProcessingTask Planning and Reinforcement Learning for General-Purpose Service Robots | AI2CorpusStudio and AbstractExplorer: Reading and Writing Scientific Works at ScaleWhy Ai2? | Nathan Lambert & Kyle WiggersBrain-Body Co-Optimization of Embodied Machines | Embodied AI Lecture Series at AI2Interaction Informed Design of Trustworthy AI
Ai2 |

Deduplication of Large-scale Text Datasets for Pretraining of Language Models

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER