th1nhhdk/local_ai_ocr
An local, offline (after initial setup), portable OCR software that can process images and PDF files, using DeepSeek-OCR-2 AI (running directly on your machine).
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
$ git clone https://github.com/oomol-lab/pdf-craft.gitAn local, offline (after initial setup), portable OCR software that can process images and PDF files, using DeepSeek-OCR-2 AI (running directly on your machine).
Open-source batch OCR workbench β a free, local alternative to ABBYY FineReader. Powered by Ollama + GLM-OCR + PP-DocLayoutV3, ~0.5s/page on RTX 4090. Three-panel editor, layout-aware, PDF/image batch processing, Markdown/Word export. ζΉιOCRε·₯δ½ε°οΌηΊ―ζ¬ε°θΏθ‘οΌε θ΄ΉεΉ³ζΏABBYYοΌιεδΉ¦η±ζζ‘£ζ°εεγ
Extract annotations (highlights and scribbles) from PDF, EPUB, and notebooks marked with reMarkable tablets. Export to Markdown, PDF, PNG, SVG
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
A fast, helpful, and open-source document parser
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Data from GitHub Β· snapshot Sep 24, 2026