dlgjr/FINAR-VL๐ฅ active
Open-source financial multimodal LLM for document understanding, chart/table reasoning, numerical reasoning, and financial analysis, with SFT, dual-branch RL, MOPD distillation, and public benchmark evaluation.
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
$ git clone https://github.com/X-PLUG/mPLUG-DocOwl.gitOpen-source financial multimodal LLM for document understanding, chart/table reasoning, numerical reasoning, and financial analysis, with SFT, dual-branch RL, MOPD distillation, and public benchmark evaluation.
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.
A Repo For Document AI
Let Claude (or any LLM) actually watch a video โ scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.
Code for the paper "PICK: Processing Key Information Extraction from Documents using Improved Graph Learning-Convolutional Networks" (ICPR 2020)
Data from GitHub ยท snapshot Sep 24, 2026