olmOCR: an open-source tool to extract clean plain text from PDFs! @allenai
olmOCR: an open-source tool to extract clean plain text from PDFs!  @allenai
Uploaded February 2025 | Updated September 2026, 2 days ago
Introducing olmOCR, our open-source tool to extract clean plain text from PDFs!

Built for scale, olmOCR handles many document types with high throughput. Run it on your own GPU for free—at over 3000 token/s, equivalent to $190 per million pages, or 1/32 the cost of GPT-4o.

You can try olmOCR for yourself at https://olmocr.allenai.org!

Blog post: olmocr.allenai.org/blog
Training and toolkit code: github.com/allenai/olmocr
Hugging Face collection:huggingface.co/collections/allenai/olmocr-67af8630b0062a25bf1b54a1
Questions? Ask on Discord: discord.com/invite/NE5xPufNwu
olmOCR: an open-source tool to extract clean plain text from PDFs!Molmo 2 | Counting objects and actionsBLADE: Benchmarking Language Model Agents for Data-Driven ScienceJust-DREAM-about-it: Figurative Language Understanding with DREAM-FLUTEHelping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward HackingTowards the Age of Computation and AI for High Performance ClimateWhen Not to Trust Language Models: Investigating Effectiveness of Parametric&Non-Parametric MemoriesOpen-Ended Learning Leads to Generally Capable Agents | Embodied AI Lecture Series at AI2From LLMs to Agents: Generalizability from the Inside OutData-Centric Approaches to Adapting Foundation ModelsThe University of Washington eScience Institute: a Home for Data-Intensive DiscoveryGeneralization for Robot Learning In The Wild | Embodied AI Lecture series at AI2
Ai2 |

olmOCR: an open-source tool to extract clean plain text from PDFs!

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER