Repository
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
ai4sciencechineseocrdocument-parsingdocument-translationkieocrpaddleocr-vlpdf-extractor-rag
Homepage
https://www.paddleocr.com ↗Sourced from GitHub · Updated September 6, 2026
Data from the GitHub REST API. Not affiliated with or endorsed by GitHub.
View original source ↗Spot an error on this page? Let us know →FAQ
Common questions
What is PaddlePaddle/PaddleOCR?
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
What license does PaddlePaddle/PaddleOCR use?
PaddlePaddle/PaddleOCR is licensed under Apache-2.0.
How popular is PaddlePaddle/PaddleOCR?
PaddlePaddle/PaddleOCR has 89k stars and 11.3k forks on GitHub, with 242 open issues.
Is PaddlePaddle/PaddleOCR actively maintained?
The last push to PaddlePaddle/PaddleOCR was July 22, 2026.