Uploaded February 2025 | Updated September 2026, 2 days ago
Introducing olmOCR, our open-source tool to extract clean plain text from PDFs!
Built for scale, olmOCR handles many document types with high throughput. Run it on your own GPU for free—at over 3000 token/s, equivalent to $190 per million pages, or 1/32 the cost of GPT-4o.
You can try olmOCR for yourself at https://olmocr.allenai.org!
Blog post: olmocr.allenai.org/blog
Training and toolkit code: github.com/allenai/olmocr
Hugging Face collection:huggingface.co/collections/allenai/olmocr-67af8630b0062a25bf1b54a1
Questions? Ask on Discord: discord.com/invite/NE5xPufNwu
Introducing olmOCR, our open-source tool to extract clean plain text from PDFs!
Built for scale, olmOCR handles many document types with high throughput. Run it on your own GPU for free—at over 3000 token/s, equivalent to $190 per million pages, or 1/32 the cost of GPT-4o.
You can try olmOCR for yourself at https://olmocr.allenai.org!
Blog post: olmocr.allenai.org/blog
Training and toolkit code: github.com/allenai/olmocr
Hugging Face collection:huggingface.co/collections/allenai/olmocr-67af8630b0062a25bf1b54a1
Questions? Ask on Discord: discord.com/invite/NE5xPufNwu










