πŸ† #1,605 overall#244 of 313 in OCR & Documents

anig1scur /tocify

Add PDF bookmarks online and make it searchable by OCR

$ git clone https://github.com/anig1scur/tocify.git
GitHub social preview for anig1scur/tocify
Stars
479
479
Forks
20
20
Language
JavaScript
License
GPL-3.0
Created
Jan 23, 2025
1.7 years old
Last push
Jun 27, 2026

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

opendatalab/MinerUπŸ”₯ active

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

PythonOtherupdated Sep 24, 2026
GitHub β†—β˜… 80.6Kβ‘‚ 6.7K
OCR & Documents

ocrmypdf/OCRmyPDFπŸ”₯ active

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

PythonMPL-2.0updated Sep 22, 2026
GitHub β†—β˜… 34.9Kβ‘‚ 2.4K
OCR & Documents

run-llama/liteparseπŸ”₯ active

A fast, helpful, and open-source document parser

RustApache-2.0updated Sep 22, 2026
GitHub β†—β˜… 12.6Kβ‘‚ 860
OCR & Documents

getomni-ai/zerox

OCR & Document Extraction using vision models

TypeScriptMITupdated May 20, 2025
GitHub β†—β˜… 12.3Kβ‘‚ 848
OCR & Documents

pymupdf/PyMuPDFπŸ”₯ active

PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.

PythonAGPL-3.0updated Sep 24, 2026
GitHub β†—β˜… 10.8Kβ‘‚ 804
OCR & Documents

bytedance/Dolphin

The official repo for β€œDolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

PythonOtherupdated Mar 25, 2026
GitHub β†—β˜… 9.1Kβ‘‚ 778

Data from GitHub Β· snapshot Sep 24, 2026