Sumanth077/Hands-On-AI-Engineeringπ₯ active
A curated collection of practical AI projects implementing OCR systems, RAG, AI agents, and other AI use cases.
Vision utilities for web interaction agents π
$ git clone https://github.com/reworkd/tarsier.gitA curated collection of practical AI projects implementing OCR systems, RAG, AI agents, and other AI use cases.
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
The official repo for βDolphin: Document Image Parsing via Heterogeneous Anchor Promptingβ, ACL, 2025.
Transforms PDF, Documents and Images into Enriched Structured Data
An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.
Data from GitHub Β· snapshot Sep 24, 2026