🏆 #563 overall#93 of 313 in OCR & Documents🔥 active this week

wxyhgk /retain-pdf

在保留版面、公式与结构的前提下进行 PDF 翻译,适用于科研与技术文档

$ git clone https://github.com/wxyhgk/retain-pdf.git
GitHub social preview for wxyhgk/retain-pdf
Stars
2.3K
2,314
Forks
278
278
Language
Python
License
MIT
Created
Mar 29, 2026
0.5 years old
Last push
Sep 21, 2026
🔥 this week

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

run-llama/liteparse🔥 active

A fast, helpful, and open-source document parser

RustApache-2.0updated Sep 22, 2026
GitHub ↗★ 12.6K⑂ 860
OCR & Documents

SylphxAI/citra🔥 active

Citra — PDF answers with page-level proof. Local-first structured text, tables, OCR, visual evidence, and citations via MCP, CLI, and SDK.

TypeScriptMITupdated Sep 23, 2026
GitHub ↗★ 935⑂ 82
OCR & Documents

AaronGIG/pdf2zh-desktop🔥 active

📖 开箱即用的 PDF 学术翻译神器 | Win + Mac 双平台 | 公式排版完美保留 · Zotero 深度联动 · 35 种语言 · 20+ AI 翻译引擎 · 表格/OCR/术语库 · 批量翻译 | 基于 PDFMathTranslate (EMNLP 2025)

Pythonno licenseupdated Sep 18, 2026
GitHub ↗★ 466⑂ 21
OCR & Documents

opendatalab/MinerU🔥 active

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

PythonOtherupdated Sep 24, 2026
GitHub ↗★ 80.6K⑂ 6.7K
OCR & Documents

ocrmypdf/OCRmyPDF🔥 active

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

PythonMPL-2.0updated Sep 22, 2026
GitHub ↗★ 34.9K⑂ 2.4K
OCR & Documents

getomni-ai/zerox

OCR & Document Extraction using vision models

TypeScriptMITupdated May 20, 2025
GitHub ↗★ 12.3K⑂ 848

Data from GitHub · snapshot Sep 24, 2026