Skip to content

Entry

PaddleOCR

Appears in 5 awesome lists

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

Open github.compaddlepaddle/paddleocr

Found in these lists

Awesome Ai For Science

Section: High-Performance Document Processing · Advanced OCR with PP-StructureV3 document parsing, 13% accuracy improvement, supports 80+ languages

FreshScore 86

Awesome Data Analysis

Section: Additional Resources and Tools · Production-ready OCR toolkit with multilingual and document AI support.

FreshScore 80

Awesome OCR

Section: Optical Character Recognition Engines and Frameworks

FreshScore 73

Awesome Open Source AI

Section: 5. Retrieval-Augmented Generation (RAG) & Knowledge · Large-scale OCR suite with detection, recognition, and layout analysis, used widely for document digitization and downstream RAG pipelines.

FreshScore 89

awesome-python

Section: Other · Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

FreshScore 81

Contributors

🟪 - A complete study plan for a computer science education.

In 7 listsDetails

Sketch

A design toolkit built to help you create your best work from your earliest ideas, through to final artwork (for macOS).

In 7 listsDetails

Kaldi

Kaldi is a toolkit for speech recognition written in C++ and licensed under the Apache License v2.0. Kaldi is intended for use by speech recognition researchers.

In 5 listsDetails

Arxiv

arXiv is a free distribution service and an open-access archive for 2,142,712 scholarly articles in the fields of physics, mathematics, computer science, quantitative biology, quantitative finance, statistics, electrical engineering and systems science, and economics

In 6 listsDetails

olmOCR

Toolkit for linearizing academic PDFs into LLM-ready text with high accuracy and structure preservation, optimized for scientific literature extraction

In 5 listsDetails

opendatalab/MinerU

High-accuracy document parsing for LLM and RAG workflows. Converts PDFs, Word, PPTs, and images into structured Markdown/JSON with VLM+OCR dual engine.

In 5 listsDetails

A To Z Resources For Students

This is a dynamic list for everything you need to know for coders and non-coders. Make yourself comfortable updating the list and using any of the resources listed.

In 2 listsDetails

EasyOCR

Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.

In 4 listsDetails