OCR & Documents
aiptimizer/TurboOCR
TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
C++MITupdated Sep 8, 2026
PDF to markdown using vision LLMs โ tables, layouts, and structure preserved
$ git clone https://github.com/yigitkonur/api-llm-ocr.gitTurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
ParseBench - A Document Parsing Benchmark for AI Agents
๐จ Ready-to-use DeepSeek-OCR Web UI | Modern Interface | 7 Recognition Modes | Batch Processing | Real-time Logging | Fully Responsive
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
A fast, helpful, and open-source document parser
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Data from GitHub ยท snapshot Sep 24, 2026