Skip to content
The Internet Compass

Repository

PaddlePaddle/PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

ai4sciencechineseocrdocument-parsingdocument-translationkieocrpaddleocr-vlpdf-extractor-rag
★ 89k11.3k forksPythonView on GitHub ↗

Homepage

https://www.paddleocr.com ↗

Sourced from GitHub · Updated September 6, 2026

Data from the GitHub REST API. Not affiliated with or endorsed by GitHub.

View original source ↗Spot an error on this page? Let us know →

FAQ

Common questions

What is PaddlePaddle/PaddleOCR?

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

What license does PaddlePaddle/PaddleOCR use?

PaddlePaddle/PaddleOCR is licensed under Apache-2.0.

How popular is PaddlePaddle/PaddleOCR?

PaddlePaddle/PaddleOCR has 89k stars and 11.3k forks on GitHub, with 242 open issues.

Is PaddlePaddle/PaddleOCR actively maintained?

The last push to PaddlePaddle/PaddleOCR was July 22, 2026.