πŸ† #574 overall#94 of 313 in OCR & Documents

WZBSocialScienceCenter /pdftabextract

A set of tools for extracting tables from PDF files helping to do data mining on (OCR-processed) scanned documents.

$ git clone https://github.com/WZBSocialScienceCenter/pdftabextract.git
GitHub social preview for WZBSocialScienceCenter/pdftabextract
Stars
2.3K
2,253
Forks
365
365
Language
Python
License
Apache-2.0
Created
Jul 8, 2016
10.2 years old
Last push
Jun 24, 2022

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

ocrmypdf/OCRmyPDFπŸ”₯ active

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

PythonMPL-2.0updated Sep 22, 2026
GitHub β†—β˜… 34.9Kβ‘‚ 2.4K
OCR & Documents

JaidedAI/EasyOCR

Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.

πŸ” Search & Retrieval
PythonApache-2.0updated Dec 5, 2025
GitHub β†—β˜… 30Kβ‘‚ 3.6K
OCR & Documents

pymupdf/PyMuPDFπŸ”₯ active

PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.

PythonAGPL-3.0updated Sep 24, 2026
GitHub β†—β˜… 10.8Kβ‘‚ 804
OCR & Documents

bytedance/Dolphin

The official repo for β€œDolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

PythonOtherupdated Mar 25, 2026
GitHub β†—β˜… 9.1Kβ‘‚ 778
OCR & Documents

axa-group/Parsr

Transforms PDF, Documents and Images into Enriched Structured Data

🧠 Natural Language Processing
JavaScriptApache-2.0updated Mar 20, 2026
GitHub β†—β˜… 6.2Kβ‘‚ 316
OCR & Documents

clawsoftware/clawPDF

Open Source Virtual (Network) Printer for Windows that allows you to create PDFs, OCR text, and print images, with advanced features usually available only in enterprise solutions.

C#AGPL-3.0updated May 16, 2023
GitHub β†—β˜… 2Kβ‘‚ 224

Data from GitHub Β· snapshot Sep 24, 2026