Were bringing training text out in the open, Introducing OLMoTrace @allenai
Were bringing training text out in the open, Introducing OLMoTrace  @allenai
Uploaded April 2025 | Updated September 2026, 4 days ago
For years it’s been an open question — how much is a language model learning and synthesizing information, and how much is it just memorizing and reciting?

Introducing OLMoTrace, a new feature in the Ai2 Playground that begins to shed some light.
OLMoTrace connects phrases or even whole sentences in the language model’s output back to verbatim matches in its training data. It does this by searching billions of documents and trillions of tokens in real time and highlighting where it finds compelling matches.

You can try OLMoTrace for yourself at playground.allenai.org/!

Blog post: allenai.org/blog/olmotrace
Training and toolkit code: github.com/allenai/infinigram-api
Questions? Ask on Discord: discord.com/invite/NE5xPufNwu
Were bringing training text out in the open, Introducing OLMoTraceSide Effects May Include Homogenization and Overusing ClichesGooAQ: Open Question Answering with Diverse Answer Types | AI2Language AI for RNA Virus and RNA VaccineDoing for our robots what nature did for us | Embodied AI Lecture series at AI2Optimal Transport Posterior Alignment for Cross-lingual Semantic ParsingFiguring out how the world works: causality in a world full of real peopleAutomated Hypothesis Validation with Agentic Sequential FalsificationsOpenWebMath: An Open Dataset of High-Quality Mathematical Web TextFrom F to A on the N.Y. Regents Science Exams: An Overview of the Aristo Project | AI2Meet an AI Using ScholarReliability and interactive debugging for language models
Ai2 |

We're bringing training text out in the open, Introducing OLMoTrace

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER