Uploaded November 2023 | Updated September 2026, 3 days ago
Abstract:
There is growing evidence that pretraining on high quality, carefully thought-out tokens such as code or mathematics plays an important role in improving the reasoning abilities of large language models. For example, Minerva, a PaLM model finetuned on billions of tokens of mathematical documents from arXiv and the web, reported dramatically improved performance on problems that require quantitative reasoning. However, because all known open source web datasets employ preprocessing that does not faithfully preserve mathematical notation, the benefits of large scale training on quantitive web documents are unavailable to the research community. We introduce OpenWebMath, an open dataset inspired by these works containing 14.7B tokens of mathematical webpages from Common Crawl. We describe in detail our method for extracting text and LaTeX content and removing boilerplate from HTML documents, as well as our methods for quality filtering and deduplication. Additionally, we run small-scale experiments by training 1.4B parameter language models on OpenWebMath, showing that models trained on 14.7B tokens of our dataset surpass the performance of models trained on over 20x the amount of general language data. We hope that our dataset, openly released on the Hugging Face Hub, will help spur advances in the reasoning abilities of large language models.
Bio:
Keiran Paster is a fifth year PhD student supervised by Jimmy Ba and Sheila McIlraith at the University of Toronto. Keiran’s research lies in the intersection of sequence modeling and decision-making. His work has formalized how sequence models can be made to make sequential decisions, even in stochastic environments through specially optimized prompts. This work lead to the creation of a groundbreaking agent called STEVE-1 that can play Minecraft from pixels with keyboard and mouse actions while following both text and visual instructions through the fine-tuning of foundation models. Keiran also works on LLM reasoning, including work on automatic prompt engineering and its application for optimizing prompts for performant and general reasoning, as well as work on a web-scale math dataset and corresponding powerful foundation model for reasoning. Keiran also worked on data efficiency for frontier models within Google Research.
Abstract:
There is growing evidence that pretraining on high quality, carefully thought-out tokens such as code or mathematics plays an important role in improving the reasoning abilities of large language models. For example, Minerva, a PaLM model finetuned on billions of tokens of mathematical documents from arXiv and the web, reported dramatically improved performance on problems that require quantitative reasoning. However, because all known open source web datasets employ preprocessing that does not faithfully preserve mathematical notation, the benefits of large scale training on quantitive web documents are unavailable to the research community. We introduce OpenWebMath, an open dataset inspired by these works containing 14.7B tokens of mathematical webpages from Common Crawl. We describe in detail our method for extracting text and LaTeX content and removing boilerplate from HTML documents, as well as our methods for quality filtering and deduplication. Additionally, we run small-scale experiments by training 1.4B parameter language models on OpenWebMath, showing that models trained on 14.7B tokens of our dataset surpass the performance of models trained on over 20x the amount of general language data. We hope that our dataset, openly released on the Hugging Face Hub, will help spur advances in the reasoning abilities of large language models.
Bio:
Keiran Paster is a fifth year PhD student supervised by Jimmy Ba and Sheila McIlraith at the University of Toronto. Keiran’s research lies in the intersection of sequence modeling and decision-making. His work has formalized how sequence models can be made to make sequential decisions, even in stochastic environments through specially optimized prompts. This work lead to the creation of a groundbreaking agent called STEVE-1 that can play Minecraft from pixels with keyboard and mouse actions while following both text and visual instructions through the fine-tuning of foundation models. Keiran also works on LLM reasoning, including work on automatic prompt engineering and its application for optimizing prompts for performant and general reasoning, as well as work on a web-scale math dataset and corresponding powerful foundation model for reasoning. Keiran also worked on data efficiency for frontier models within Google Research.
![From F to A on the N.Y. Regents Science Exams: An Overview of the Aristo Project | AI2
AI has achieved remarkable mastery over games such as Chess, Go, and Poker, and even Jeopardy!, but the rich variety of standardized exams has remained a landmark challenge. Even as recently as 2016, the best AI system could achieve merely 59.3% on an 8th Grade science exam.
This talk reports success on the Grade 8 New York Regents Science Exam, where for the first time a system scores more than 90% on the exams non-diagram, multiple choice (NDMC) questions. In addition, our Aristo system, building upon the success of recent language models, exceeded 83% on the corresponding Grade 12 Science Exam NDMC questions. The results, on unseen test questions, are robust across different test years and different variations of this kind of test. They demonstrate that modern Natural Language Processing (NLP) methods can result in mastery on this task. While not a full solution to general question-answering (the questions are limited to 8th Grade multiple-choice science) it represents a significant milestone for the field. [ Paper at AI Magazine 41 (4), Winter 2020, https://arxiv.org/pdf/1909.01958.pdf ] From F to A on the N.Y. Regents Science Exams: An Overview of the Aristo Project | AI2](https://i.ytimg.com/vi/CR3aICkhCJM/mqdefault.jpg)









