opendatalab/MinerUπ₯ active
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical structure, tables, and meta information from textual electronic documents. (Parse document; Document content extraction; Logical structure extraction; PDF parser; Scanned document parser; DOCX parser; HTML parser
$ git clone https://github.com/ispras/dedoc.gitTransforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
The official repo for βDolphin: Document Image Parsing via Heterogeneous Anchor Promptingβ, ACL, 2025.
Scan, index, and archive all of your paper documents (acquired by Mayan EDMS)
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
Data from GitHub Β· snapshot Sep 24, 2026