๐Ÿ† #541 overall#88 of 313 in OCR & Documents

X-PLUG /mPLUG-DocOwl

mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

$ git clone https://github.com/X-PLUG/mPLUG-DocOwl.git
GitHub social preview for X-PLUG/mPLUG-DocOwl
Stars
2.4K
2,413
Forks
155
155
Language
Python
License
Apache-2.0
Created
Jul 4, 2023
3.2 years old
Last push
May 30, 2025

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

dlgjr/FINAR-VL๐Ÿ”ฅ active

Open-source financial multimodal LLM for document understanding, chart/table reasoning, numerical reasoning, and financial analysis, with SFT, dual-branch RL, MOPD distillation, and public benchmark evaluation.

Pythonno licenseupdated Sep 23, 2026
GitHub โ†—โ˜… 100โ‘‚ 1
OCR & Documents

AlibabaResearch/AdvancedLiterateMachinery

A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.

C++Apache-2.0updated Mar 17, 2026
GitHub โ†—โ˜… 1.8Kโ‘‚ 195
OCR & Documents

HUANGCHIHHUNGLeo/claude-real-video๐Ÿ”ฅ active

Let Claude (or any LLM) actually watch a video โ€” scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.

PythonMITupdated Sep 19, 2026
GitHub โ†—โ˜… 2.2Kโ‘‚ 194
OCR & Documents

OpenBMB/VisRAG

Parsing-free RAG supported by VLMs

PythonApache-2.0updated Aug 9, 2026
GitHub โ†—โ˜… 981โ‘‚ 79
OCR & Documents

wenwenyu/PICK-pytorch

Code for the paper "PICK: Processing Key Information Extraction from Documents using Improved Graph Learning-Convolutional Networks" (ICPR 2020)

PythonMITupdated Jul 25, 2024
GitHub โ†—โ˜… 570โ‘‚ 189

Data from GitHub ยท snapshot Sep 24, 2026