datalab-to/lift
Extract structured data from documents quickly and accurately.
Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown
$ git clone https://github.com/CatchTheTornado/text-extract-api.gitExtract structured data from documents quickly and accurately.
📖 开箱即用的 PDF 学术翻译神器 | Win + Mac 双平台 | 公式排版完美保留 · Zotero 深度联动 · 35 种语言 · 20+ AI 翻译引擎 · 表格/OCR/术语库 · 批量翻译 | 基于 PDFMathTranslate (EMNLP 2025)
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
OCR software, free and offline. 开源、免费的离线OCR软件。支持截屏/批量导入图片,PDF文档识别,排除水印/页眉页脚,扫描/生成二维码。内置多国语言库。
A community-supported supercharged document management system: scan, index and archive all your documents
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Data from GitHub · snapshot Sep 24, 2026