pymupdf/PyMuPDFπ₯ active
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Extract annotations (highlights and scribbles) from PDF, EPUB, and notebooks marked with reMarkable tablets. Export to Markdown, PDF, PNG, SVG
$ git clone https://github.com/lucasrla/remarks.gitPyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
The official repo for βDolphin: Document Image Parsing via Heterogeneous Anchor Promptingβ, ACL, 2025.
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
A search engine that "just works" for Obsidian. Supports OCR and PDF indexing.
SumatraPDF fork: Chinese EPUB/MOBI, smart PDF dark mode, OCR, TTS, offline dictionary.
Data from GitHub Β· snapshot Sep 24, 2026