Uploaded August 2024 | Updated September 2026, 1 day ago
In this talk, I'll cover the newly released DataComp for Language Models project, in which we generate a testbed for controlled experiments of building better datasets for pretraining language models in a compute-limited regime. From here I'll pivot to discussing one particular aspect of building better datasets: removing duplicates and near-duplicates from large corpuses of text, explaining several key techniques as well as our findings from extensive deduplication ablations. Finally, I'll raise some several open questions and future directions regarding deduplication of pretraining datasets, including some unpublished (but interesting!) results.
In this talk, I'll cover the newly released DataComp for Language Models project, in which we generate a testbed for controlled experiments of building better datasets for pretraining language models in a compute-limited regime. From here I'll pivot to discussing one particular aspect of building better datasets: removing duplicates and near-duplicates from large corpuses of text, explaining several key techniques as well as our findings from extensive deduplication ablations. Finally, I'll raise some several open questions and future directions regarding deduplication of pretraining datasets, including some unpublished (but interesting!) results.




![Enhancing Reasoning in Smaller Models through Self-Training
Abstract: Smaller language models can develop robust reasoning capabilities through pre-training, fine-tuning, or knowledge distillation from large language models (LLMs). However, unlike LLMs that employ a diverse array of reasoning strategies, smaller models typically rely on a single dominant approach. This limitation restricts their effectiveness in handling different multi-step reasoning tasks, which require a wide range of strategies in order to solve them. To address this challenge, self-training leverages the model’s own generated data, enabling smaller models to autonomously learn and adapt their reasoning strategies for improved performance across diverse tasks.
I will talk about a self-guided iterative distillation framework (SIKeD [1]), which combines multi-strategy outputs from LLMs with self-generated data from the smaller model to identify the most effective strategy for a given task in an on-policy manner.
Later, I will talk about how this self-training approach can be extended to improve refinement in models, where a model can learn to iteratively refine its output, eventually learning to pick the right strategy in its first attempt (SMART [2]).
[1] https://arxiv.org/abs/2410.18574
[2] https://arxiv.org/abs/2410.16128
Bio: Kumar Shridhar is a final-year Ph.D. candidate at ETH Zürich, Switzerland, under the supervision of Prof. Mrinmaya Sachan from ETH and Dr. Nicholas Monath from Google DeepMind. Prior to his doctoral studies, he spent summers interning at FAIR, Microsoft Research, and Alexa AI, and improving conversational AI at different startups.
His research focuses on advancing the reasoning capabilities of large language models (LLMs) and developing efficient distillation methods to impart these skills to smaller models. Moreover he is also working model alignment, autonomous agents, and model refinement. He is also a member of Swiss AI initiative, where the team is training foundational models across various domains. Enhancing Reasoning in Smaller Models through Self-Training](https://i.ytimg.com/vi/SS59gCT2KKs/mqdefault.jpg)





