NanoNets/docstrange
Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
$ git clone https://github.com/opendataloader-project/opendataloader-pdf.gitExtract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.
A Model Context Protocol server for converting almost anything to Markdown
ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
A community-supported supercharged document management system: scan, index and archive all your documents
OCR model that handles complex tables, forms, handwriting with full layout.
Data from GitHub ยท snapshot Sep 24, 2026