🏆 #420 overall#67 of 313 in OCR & Documents

CatchTheTornado /text-extract-api

Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown

$ git clone https://github.com/CatchTheTornado/text-extract-api.git
GitHub social preview for CatchTheTornado/text-extract-api
Stars
3.2K
3,181
Forks
279
279
Language
Python
License
MIT
Created
Oct 23, 2024
1.9 years old
Last push
Dec 8, 2025

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

datalab-to/lift

Extract structured data from documents quickly and accurately.

PythonApache-2.0updated Jun 19, 2026
GitHub ↗★ 910⑂ 85
OCR & Documents

AaronGIG/pdf2zh-desktop🔥 active

📖 开箱即用的 PDF 学术翻译神器 | Win + Mac 双平台 | 公式排版完美保留 · Zotero 深度联动 · 35 种语言 · 20+ AI 翻译引擎 · 表格/OCR/术语库 · 批量翻译 | 基于 PDFMathTranslate (EMNLP 2025)

Pythonno licenseupdated Sep 18, 2026
GitHub ↗★ 466⑂ 21
OCR & Documents

opendatalab/MinerU🔥 active

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

PythonOtherupdated Sep 24, 2026
GitHub ↗★ 80.6K⑂ 6.7K
OCR & Documents

hiroi-sora/Umi-OCR

OCR software, free and offline. 开源、免费的离线OCR软件。支持截屏/批量导入图片,PDF文档识别,排除水印/页眉页脚,扫描/生成二维码。内置多国语言库。

PythonMITupdated Nov 20, 2025
GitHub ↗★ 47.5K⑂ 4.6K
OCR & Documents

paperless-ngx/paperless-ngx🔥 active

A community-supported supercharged document management system: scan, index and archive all your documents

PythonGPL-3.0updated Sep 24, 2026
GitHub ↗★ 46K⑂ 3.2K
OCR & Documents

ocrmypdf/OCRmyPDF🔥 active

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

PythonMPL-2.0updated Sep 22, 2026
GitHub ↗★ 34.9K⑂ 2.4K

Data from GitHub · snapshot Sep 24, 2026