enoch3712/ExtractThinker
ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.
ParseBench - A Document Parsing Benchmark for AI Agents
$ git clone https://github.com/run-llama/ParseBench.gitExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.
A community-supported supercharged document management system: scan, index and archive all your documents
LLM-Driven Extraction of Unstructured Data โ Built for API Deployments & ETL Pipeline Workflows
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.
TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
Data from GitHub ยท snapshot Sep 24, 2026