๐Ÿ† #1,442 overall#217 of 313 in OCR & Documents๐Ÿ”ฅ active this week

run-llama /ParseBench

ParseBench - A Document Parsing Benchmark for AI Agents

$ git clone https://github.com/run-llama/ParseBench.git
GitHub social preview for run-llama/ParseBench
Stars
592
592
Forks
109
109
Language
Python
License
Apache-2.0
Created
Apr 10, 2026
0.5 years old
Last push
Sep 23, 2026
๐Ÿ”ฅ this week

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

enoch3712/ExtractThinker

ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.

๐Ÿง  Natural Language Processing
PythonApache-2.0updated Sep 16, 2026
GitHub โ†—โ˜… 1.6Kโ‘‚ 153
OCR & Documents

paperless-ngx/paperless-ngx๐Ÿ”ฅ active

A community-supported supercharged document management system: scan, index and archive all your documents

PythonGPL-3.0updated Sep 24, 2026
GitHub โ†—โ˜… 46Kโ‘‚ 3.2K
OCR & Documents

Zipstack/unstract๐Ÿ”ฅ active

LLM-Driven Extraction of Unstructured Data โ€” Built for API Deployments & ETL Pipeline Workflows

PythonAGPL-3.0updated Sep 24, 2026
GitHub โ†—โ˜… 7.3Kโ‘‚ 720
OCR & Documents

yobix-ai/extractous

Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.

๐Ÿง  Natural Language Processing
RustApache-2.0updated Dec 21, 2024
GitHub โ†—โ˜… 1.8Kโ‘‚ 96
OCR & Documents

NanoNets/docstrange

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

PythonMITupdated Oct 31, 2025
GitHub โ†—โ˜… 1.6Kโ‘‚ 139
OCR & Documents

aiptimizer/TurboOCR

TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC

C++MITupdated Sep 8, 2026
GitHub โ†—โ˜… 1.1Kโ‘‚ 112

Data from GitHub ยท snapshot Sep 24, 2026