πŸ† #5 overall#1 of 313 in OCR & Documents

PaddlePaddle /PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

$ git clone https://github.com/PaddlePaddle/PaddleOCR.git
GitHub social preview for PaddlePaddle/PaddleOCR
Stars
90.1K
90,142
Forks
11.4K
11,404
Language
Python
License
Apache-2.0
Created
May 8, 2020
6.4 years old
Last push
Sep 16, 2026

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

NanoNets/docstrange

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

PythonMITupdated Oct 31, 2025
GitHub β†—β˜… 1.6Kβ‘‚ 139
OCR & Documents

Topdu/OpenOCRπŸ”₯ active

OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.

PythonApache-2.0updated Sep 21, 2026
GitHub β†—β˜… 1.5Kβ‘‚ 147
OCR & Documents

opendatalab/MinerUπŸ”₯ active

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

PythonOtherupdated Sep 24, 2026
GitHub β†—β˜… 80.6Kβ‘‚ 6.7K
OCR & Documents

run-llama/liteparseπŸ”₯ active

A fast, helpful, and open-source document parser

RustApache-2.0updated Sep 22, 2026
GitHub β†—β˜… 12.6Kβ‘‚ 860
OCR & Documents

bytedance/Dolphin

The official repo for β€œDolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

PythonOtherupdated Mar 25, 2026
GitHub β†—β˜… 9.1Kβ‘‚ 778

Data from GitHub Β· snapshot Sep 24, 2026