OCR & Documents
opendatalab/MinerUπ₯ active
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
PythonOtherupdated Sep 24, 2026
Download your resume from resume.io as PDF
$ git clone https://github.com/felipeall/resumeio-to-pdf.gitTransforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
A fast, helpful, and open-source document parser
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
The official repo for βDolphin: Document Image Parsing via Heterogeneous Anchor Promptingβ, ACL, 2025.
Data from GitHub Β· snapshot Sep 24, 2026