opendataloader-project/opendataloader-pdf๐ฅ active
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.
$ git clone https://github.com/NanoNets/docstrange.gitPDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.
Turn any document into clean, AI-ready Markdown. Local-first desktop app: reads scanned PDFs, batches folders, runs offline, and uses far fewer tokens than vision models.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
A community-supported supercharged document management system: scan, index and archive all your documents
A Model Context Protocol server for converting almost anything to Markdown
Data from GitHub ยท snapshot Sep 24, 2026