πŸ† #226 overall#37 of 313 in OCR & DocumentsπŸ”₯ active this week

oomol-lab /pdf-craft

PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.

$ git clone https://github.com/oomol-lab/pdf-craft.git
GitHub social preview for oomol-lab/pdf-craft
Stars
6.3K
6,324
Forks
458
458
Language
Python
License
MIT
Created
Feb 12, 2025
1.6 years old
Last push
Sep 23, 2026
πŸ”₯ this week

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

th1nhhdk/local_ai_ocr

An local, offline (after initial setup), portable OCR software that can process images and PDF files, using DeepSeek-OCR-2 AI (running directly on your machine).

PythonApache-2.0updated Jul 5, 2026
GitHub β†—β˜… 876β‘‚ 207
OCR & Documents

vorojar/Folio-OCR

Open-source batch OCR workbench β€” a free, local alternative to ABBYY FineReader. Powered by Ollama + GLM-OCR + PP-DocLayoutV3, ~0.5s/page on RTX 4090. Three-panel editor, layout-aware, PDF/image batch processing, Markdown/Word export. 批量OCRε·₯δ½œε°οΌŒηΊ―ζœ¬εœ°θΏθ‘ŒοΌŒε…θ΄ΉεΉ³ζ›ΏABBYYοΌŒι€‚εˆδΉ¦η±ζ–‡ζ‘£ζ•°ε­—εŒ–γ€‚

PythonMITupdated Jun 18, 2026
GitHub β†—β˜… 474β‘‚ 62
OCR & Documents

lucasrla/remarks

Extract annotations (highlights and scribbles) from PDF, EPUB, and notebooks marked with reMarkable tablets. Export to Markdown, PDF, PNG, SVG

PythonGPL-3.0updated May 26, 2024
GitHub β†—β˜… 399β‘‚ 34
OCR & Documents

run-llama/liteparseπŸ”₯ active

A fast, helpful, and open-source document parser

RustApache-2.0updated Sep 22, 2026
GitHub β†—β˜… 12.6Kβ‘‚ 860
OCR & Documents

pymupdf/PyMuPDFπŸ”₯ active

PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.

PythonAGPL-3.0updated Sep 24, 2026
GitHub β†—β˜… 10.8Kβ‘‚ 804

Data from GitHub Β· snapshot Sep 24, 2026