dlgjr/FINAR-VL๐ฅ active
Open-source financial multimodal LLM for document understanding, chart/table reasoning, numerical reasoning, and financial analysis, with SFT, dual-branch RL, MOPD distillation, and public benchmark evaluation.
Parsing-free RAG supported by VLMs
$ git clone https://github.com/OpenBMB/VisRAG.gitOpen-source financial multimodal LLM for document understanding, chart/table reasoning, numerical reasoning, and financial analysis, with SFT, dual-branch RL, MOPD distillation, and public benchmark evaluation.
A curated collection of practical AI projects implementing OCR systems, RAG, AI agents, and other AI use cases.
A Repo For Document AI
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.
Multi-modal OCR pipeline optimized for ML training (text, figure, math, tables, diagrams)
Data from GitHub ยท snapshot Sep 24, 2026