Skip to content

Entry

olmOCR

Appears in 5 awesome lists

Toolkit for linearizing academic PDFs into LLM-ready text with high accuracy and structure preservation, optimized for scientific literature extraction

Open github.comallenai/olmocr

Found in these lists

Awesome Ai For Science

Section: High-Performance Document Processing · Toolkit for linearizing academic PDFs into LLM-ready text with high accuracy and structure preservation, optimized for scientific literature extraction

FreshScore 86

Awesome LLM Resources

Section: 数据 Data · A toolkit for training language models to work with PDF documents in the wild.

FreshScore 87

Awesome OCR

Section: Optical Character Recognition Engines and Frameworks

FreshScore 73

Awesome Production Machine Learning

Section: Industry Strength Natural Language Processing · olmOCR is a toolkit for training language models to work with PDF documents in the wild.

FreshScore 92

awesome-python

Section: Other · Toolkit for linearizing PDFs for LLM datasets/training

FreshScore 81

MarkItDown

Python tool for converting files and office documents to Markdown. Supports PDF, PowerPoint, Word, Excel, images, audio, HTML, and more with OCR and transcription capabilities. MIT licensed.

In 9 listsDetails

Kaldi

Kaldi is a toolkit for speech recognition written in C++ and licensed under the Apache License v2.0. Kaldi is intended for use by speech recognition researchers.

In 5 listsDetails

opendatalab/MinerU

High-accuracy document parsing for LLM and RAG workflows. Converts PDFs, Word, PPTs, and images into structured Markdown/JSON with VLM+OCR dual engine.

In 5 listsDetails

PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

In 5 listsDetails

TensorZero

(label: good-first-issue) TensorZero creates a feedback loop for optimizing LLM applications — turning production data into smarter, faster, and cheaper models.

In 4 listsDetails

EasyOCR

Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.

In 4 listsDetails

gosseract

Go package for OCR (Optical Character Recognition), by using Tesseract C++ library.

In 4 listsDetails

franc

Detect the language of text.

In 4 listsDetails