πŸ† #1,109 overall#162 of 313 in OCR & Documents

datalab-to /lift

Extract structured data from documents quickly and accurately.

$ git clone https://github.com/datalab-to/lift.git
GitHub social preview for datalab-to/lift
Stars
910
910
Forks
85
85
Language
Python
License
Apache-2.0
Created
Jun 3, 2026
0.3 years old
Last push
Jun 19, 2026

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

ocrmypdf/OCRmyPDFπŸ”₯ active

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

PythonMPL-2.0updated Sep 22, 2026
GitHub β†—β˜… 34.9Kβ‘‚ 2.4K
OCR & Documents

pymupdf/PyMuPDFπŸ”₯ active

PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.

PythonAGPL-3.0updated Sep 24, 2026
GitHub β†—β˜… 10.8Kβ‘‚ 804
OCR & Documents

bytedance/Dolphin

The official repo for β€œDolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

PythonOtherupdated Mar 25, 2026
GitHub β†—β˜… 9.1Kβ‘‚ 778
OCR & Documents

axa-group/Parsr

Transforms PDF, Documents and Images into Enriched Structured Data

🧠 Natural Language Processing
JavaScriptApache-2.0updated Mar 20, 2026
GitHub β†—β˜… 6.2Kβ‘‚ 316
OCR & Documents

Sumanth077/Hands-On-AI-EngineeringπŸ”₯ active

A curated collection of practical AI projects implementing OCR systems, RAG, AI agents, and other AI use cases.

Pythonno licenseupdated Sep 24, 2026
GitHub β†—β˜… 3.7Kβ‘‚ 897
OCR & Documents

CatchTheTornado/text-extract-api

Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown

PythonMITupdated Dec 8, 2025
GitHub β†—β˜… 3.2Kβ‘‚ 279

Data from GitHub Β· snapshot Sep 24, 2026