PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
A supermarket receipt parser written in Python using tesseract OCR
$ git clone https://github.com/ReceiptManager/receipt-parser-legacy.gitTurn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Tesseract Open Source OCR Engine (main repository)
OCR software, free and offline. ๅผๆบใๅ ่ดน็็ฆป็บฟOCR่ฝฏไปถใๆฏๆๆชๅฑ/ๆน้ๅฏผๅ ฅๅพ็๏ผPDFๆๆกฃ่ฏๅซ๏ผๆ้คๆฐดๅฐ/้กต็้กต่๏ผๆซๆ/็ๆไบ็ปด็ ใๅ ็ฝฎๅคๅฝ่ฏญ่จๅบใ
A community-supported supercharged document management system: scan, index and archive all your documents
Pure Javascript OCR for more than 100 Languages ๐๐๐ฅ
Data from GitHub ยท snapshot Sep 24, 2026