🏆 #1,610 overall#245 of 313 in OCR & Documents

tiantian91091317 /OCR-Corrector

利用语言模型,纠正OCR识别错误

$ git clone https://github.com/tiantian91091317/OCR-Corrector.git
GitHub social preview for tiantian91091317/OCR-Corrector
Stars
477
477
Forks
103
103
Language
Python
License
Apache-2.0
Created
Jul 2, 2019
7.2 years old
Last push
May 22, 2023

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

PaddlePaddle/PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

PythonApache-2.0updated Sep 16, 2026
GitHub ↗★ 90.1K⑂ 11.4K
OCR & Documents

opendatalab/MinerU🔥 active

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

PythonOtherupdated Sep 24, 2026
GitHub ↗★ 80.6K⑂ 6.7K
OCR & Documents

hiroi-sora/Umi-OCR

OCR software, free and offline. 开源、免费的离线OCR软件。支持截屏/批量导入图片,PDF文档识别,排除水印/页眉页脚,扫描/生成二维码。内置多国语言库。

PythonMITupdated Nov 20, 2025
GitHub ↗★ 47.5K⑂ 4.6K
OCR & Documents

paperless-ngx/paperless-ngx🔥 active

A community-supported supercharged document management system: scan, index and archive all your documents

PythonGPL-3.0updated Sep 24, 2026
GitHub ↗★ 46K⑂ 3.2K
OCR & Documents

naptha/tesseract.js

Pure Javascript OCR for more than 100 Languages 📖🎉🖥

JavaScriptApache-2.0updated May 17, 2026
GitHub ↗★ 38.7K⑂ 2.4K

Data from GitHub · snapshot Sep 24, 2026