PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.
$ git clone https://github.com/Topdu/OpenOCR.gitTurn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.
TextBoxes++: A Single-Shot Oriented Scene Text Detector
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
A fast, helpful, and open-source document parser
Data from GitHub ยท snapshot Sep 24, 2026