πŸ† #8 overall#2 of 313 in OCR & DocumentsπŸ”₯ active this week

opendatalab /MinerU

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

$ git clone https://github.com/opendatalab/MinerU.git
GitHub social preview for opendatalab/MinerU
Stars
80.6K
80,582
Forks
6.7K
6,726
Language
Python
License
Other
Created
Feb 29, 2024
2.6 years old
Last push
Sep 24, 2026
πŸ”₯ this week

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

opendatalab/MinerU-Diffusion

[ECCV 2026] A diffusion-based framework for document OCR that replaces autoregressive decoding with block-level parallel diffusion decoding.

PythonMITupdated Jun 18, 2026
GitHub β†—β˜… 619β‘‚ 42
OCR & Documents

bytedance/Dolphin

The official repo for β€œDolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

PythonOtherupdated Mar 25, 2026
GitHub β†—β˜… 9.1Kβ‘‚ 778
OCR & Documents

DocumindHQ/documind

Open-source platform for extracting structured data from documents using AI.

JavaScriptOtherupdated May 15, 2025
GitHub β†—β˜… 1.5Kβ‘‚ 62
OCR & Documents

pymupdf/PyMuPDFπŸ”₯ active

PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.

PythonAGPL-3.0updated Sep 24, 2026
GitHub β†—β˜… 10.8Kβ‘‚ 804
OCR & Documents

lucasrla/remarks

Extract annotations (highlights and scribbles) from PDF, EPUB, and notebooks marked with reMarkable tablets. Export to Markdown, PDF, PNG, SVG

PythonGPL-3.0updated May 26, 2024
GitHub β†—β˜… 399β‘‚ 34
OCR & Documents

PaddlePaddle/PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

PythonApache-2.0updated Sep 16, 2026
GitHub β†—β˜… 90.1Kβ‘‚ 11.4K

Data from GitHub Β· snapshot Sep 24, 2026