πŸ† #170 overall#27 of 313 in OCR & Documents

bytedance /Dolphin

The official repo for β€œDolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

$ git clone https://github.com/bytedance/Dolphin.git
GitHub social preview for bytedance/Dolphin
Stars
9.1K
9,050
Forks
778
778
Language
Python
License
Other
Created
May 13, 2025
1.4 years old
Last push
Mar 25, 2026

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

opendatalab/MinerUπŸ”₯ active

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

PythonOtherupdated Sep 24, 2026
GitHub β†—β˜… 80.6Kβ‘‚ 6.7K
OCR & Documents

opendatalab/MinerU-Diffusion

[ECCV 2026] A diffusion-based framework for document OCR that replaces autoregressive decoding with block-level parallel diffusion decoding.

PythonMITupdated Jun 18, 2026
GitHub β†—β˜… 619β‘‚ 42
OCR & Documents

DocumindHQ/documind

Open-source platform for extracting structured data from documents using AI.

JavaScriptOtherupdated May 15, 2025
GitHub β†—β˜… 1.5Kβ‘‚ 62
OCR & Documents

ocrmypdf/OCRmyPDFπŸ”₯ active

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

PythonMPL-2.0updated Sep 22, 2026
GitHub β†—β˜… 34.9Kβ‘‚ 2.4K
OCR & Documents

run-llama/liteparseπŸ”₯ active

A fast, helpful, and open-source document parser

RustApache-2.0updated Sep 22, 2026
GitHub β†—β˜… 12.6Kβ‘‚ 860
OCR & Documents

pymupdf/PyMuPDFπŸ”₯ active

PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.

PythonAGPL-3.0updated Sep 24, 2026
GitHub β†—β˜… 10.8Kβ‘‚ 804

Data from GitHub Β· snapshot Sep 24, 2026