🏆 #328 overall#140 of 1,075 in Natural Language Processing

esbatmop /MNBVC

MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集。对标chatGPT训练的40T数据。MNBVC数据集不但包括主流文化,也包括各个小众文化甚至火星文的数据。MNBVC数据集包括新闻、作文、小说、书籍、杂志、论文、台词、帖子、wiki、古诗、歌词、商品介绍、笑话、糗事、聊天记录等一切形式的纯文本中文数据。

$ git clone https://github.com/esbatmop/MNBVC.git
GitHub social preview for esbatmop/MNBVC
Stars
4.3K
4,277
Forks
296
296
Language
Other
License
MIT
Created
Dec 31, 2022
3.7 years old
Last push
Sep 13, 2026

Categories

GitHub topics

More in Natural Language Processing

Natural Language Processing

brightmart/nlp_chinese_corpus

大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP

🗂️ Text Classification❓ Question Answering🧬 Embeddings🤖 Language Models
OtherMITupdated Feb 6, 2026
GitHub ↗★ 9.9K⑂ 1.6K
Natural Language Processing

CVI-SZU/Linly

Chinese-LLaMA 1&2、Chinese-Falcon 基础模型;ChatFlow中文对话模型;中文OpenLLaMA模型;NLP预训练/指令微调数据集

🤖 Language Models
Pythonno licenseupdated Apr 14, 2024
GitHub ↗★ 3K⑂ 222
Natural Language Processing

tim5go/zhopenie

Chinese Open Information Extraction (Tree-based Triple Relation Extraction Module)

⛏️ Information Extraction
Pythonno licenseupdated Jun 19, 2017
GitHub ↗★ 116⑂ 26
Natural Language Processing

Morizeyao/GPT2-Chinese

Chinese version of GPT2 training code, using BERT tokenizer.

PythonMITupdated Apr 25, 2024
GitHub ↗★ 7.6K⑂ 1.7K
Natural Language Processing

NLPchina/ansj_seg

ansj分词.ict的真正java实现.分词效果速度都超过开源版的ict. 中文分词,人名识别,词性标注,用户自定义词典

JavaApache-2.0updated Nov 19, 2023
GitHub ↗★ 6.5K⑂ 2.3K
Natural Language Processing

OFA-Sys/Chinese-CLIP

Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.

Jupyter NotebookMITupdated Mar 31, 2026
GitHub ↗★ 6K⑂ 551

Data from GitHub · snapshot Sep 24, 2026