opendatalab/MinerUπ₯ active
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
[ECCV 2026] A diffusion-based framework for document OCR that replaces autoregressive decoding with block-level parallel diffusion decoding.
$ git clone https://github.com/opendatalab/MinerU-Diffusion.gitTransforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
The official repo for βDolphin: Document Image Parsing via Heterogeneous Anchor Promptingβ, ACL, 2025.
Open-source platform for extracting structured data from documents using AI.
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Data from GitHub Β· snapshot Sep 24, 2026