OCR & Documents
pymupdf/PyMuPDF-Utilities
Demos, examples and utilities using PyMuPDF
Jupyter NotebookAGPL-3.0updated Jan 8, 2026
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
$ git clone https://github.com/pymupdf/PyMuPDF.gitDemos, examples and utilities using PyMuPDF
Extract annotations (highlights and scribbles) from PDF, EPUB, and notebooks marked with reMarkable tablets. Export to Markdown, PDF, PNG, SVG
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
The official repo for βDolphin: Document Image Parsing via Heterogeneous Anchor Promptingβ, ACL, 2025.
Transforms PDF, Documents and Images into Enriched Structured Data
Data from GitHub Β· snapshot Sep 24, 2026