πŸ† #147 overall#22 of 313 in OCR & DocumentsπŸ”₯ active this week

pymupdf /PyMuPDF

PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.

$ git clone https://github.com/pymupdf/PyMuPDF.git
GitHub social preview for pymupdf/PyMuPDF
Stars
10.8K
10,765
Forks
804
804
Language
Python
License
AGPL-3.0
Created
Oct 6, 2012
14.0 years old
Last push
Sep 24, 2026
πŸ”₯ this week

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

lucasrla/remarks

Extract annotations (highlights and scribbles) from PDF, EPUB, and notebooks marked with reMarkable tablets. Export to Markdown, PDF, PNG, SVG

PythonGPL-3.0updated May 26, 2024
GitHub β†—β˜… 399β‘‚ 34
OCR & Documents

opendatalab/MinerUπŸ”₯ active

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

PythonOtherupdated Sep 24, 2026
GitHub β†—β˜… 80.6Kβ‘‚ 6.7K
OCR & Documents

ocrmypdf/OCRmyPDFπŸ”₯ active

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

PythonMPL-2.0updated Sep 22, 2026
GitHub β†—β˜… 34.9Kβ‘‚ 2.4K
OCR & Documents

bytedance/Dolphin

The official repo for β€œDolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

PythonOtherupdated Mar 25, 2026
GitHub β†—β˜… 9.1Kβ‘‚ 778
OCR & Documents

axa-group/Parsr

Transforms PDF, Documents and Images into Enriched Structured Data

🧠 Natural Language Processing
JavaScriptApache-2.0updated Mar 20, 2026
GitHub β†—β˜… 6.2Kβ‘‚ 316

Data from GitHub Β· snapshot Sep 24, 2026