๐Ÿ† #814 overall#127 of 313 in OCR & Documents๐Ÿ”ฅ active this week

Topdu /OpenOCR

OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.

$ git clone https://github.com/Topdu/OpenOCR.git
GitHub social preview for Topdu/OpenOCR
Stars
1.5K
1,459
Forks
147
147
Language
Python
License
Apache-2.0
Created
May 31, 2024
2.3 years old
Last push
Sep 21, 2026
๐Ÿ”ฅ this week

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

PaddlePaddle/PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

PythonApache-2.0updated Sep 16, 2026
GitHub โ†—โ˜… 90.1Kโ‘‚ 11.4K
OCR & Documents

enoch3712/ExtractThinker

ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.

๐Ÿง  Natural Language Processing
PythonApache-2.0updated Sep 16, 2026
GitHub โ†—โ˜… 1.6Kโ‘‚ 153
OCR & Documents

opendatalab/MinerU๐Ÿ”ฅ active

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

PythonOtherupdated Sep 24, 2026
GitHub โ†—โ˜… 80.6Kโ‘‚ 6.7K
OCR & Documents

run-llama/liteparse๐Ÿ”ฅ active

A fast, helpful, and open-source document parser

RustApache-2.0updated Sep 22, 2026
GitHub โ†—โ˜… 12.6Kโ‘‚ 860

Data from GitHub ยท snapshot Sep 24, 2026