๐Ÿ† #233 overall#40 of 313 in OCR & Documents

axa-group /Parsr

Transforms PDF, Documents and Images into Enriched Structured Data

$ git clone https://github.com/axa-group/Parsr.git
GitHub social preview for axa-group/Parsr
Stars
6.2K
6,174
Forks
316
316
Language
JavaScript
License
Apache-2.0
Created
Aug 5, 2019
7.1 years old
Last push
Mar 20, 2026

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

eikek/docspell๐Ÿ”ฅ active

Assist in organizing your piles of documents, resulting from scanners, e-mails and other sources with miminal effort.

๐Ÿง  Natural Language Processing
ElmAGPL-3.0updated Sep 22, 2026
GitHub โ†—โ˜… 2.3Kโ‘‚ 184
OCR & Documents

NanoNets/docext

An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)

๐Ÿง  Natural Language Processing
PythonApache-2.0updated Mar 17, 2026
GitHub โ†—โ˜… 2.1Kโ‘‚ 156
OCR & Documents

pd3f/pd3f

๐Ÿญ PDF text extraction pipeline: self-hosted, local-first, Docker-based

๐Ÿค– Language Models
HTMLAGPL-3.0updated Oct 13, 2023
GitHub โ†—โ˜… 336โ‘‚ 38
OCR & Documents

yobix-ai/extractous

Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.

๐Ÿง  Natural Language Processing
RustApache-2.0updated Dec 21, 2024
GitHub โ†—โ˜… 1.8Kโ‘‚ 96
OCR & Documents

robocorp/rpaframework

Collection of open-source libraries and tools for Robotic Process Automation (RPA), designed to be used with both Robot Framework and Python

๐Ÿง  Natural Language Processing
PythonApache-2.0updated Aug 29, 2026
GitHub โ†—โ˜… 1.6Kโ‘‚ 275
OCR & Documents

ocrmypdf/OCRmyPDF๐Ÿ”ฅ active

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

PythonMPL-2.0updated Sep 22, 2026
GitHub โ†—โ˜… 34.9Kโ‘‚ 2.4K

Data from GitHub ยท snapshot Sep 24, 2026