Language Models
brightmart/nlp_chinese_corpus
大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP
🗂️ Text Classification❓ Question Answering🧬 Embeddings🧠 Natural Language Processing
OtherMITupdated Feb 6, 2026
青简 Qingjian:用 Rust 写的拼音输入法,候选词旁多一条正在学的语言的译词
$ git clone https://github.com/qingjian-team/qingjian.git大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP
中文语言理解测评基准 Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard
Chinese-LLaMA 1&2、Chinese-Falcon 基础模型;ChatFlow中文对话模型;中文OpenLLaMA模型;NLP预训练/指令微调数据集
Pre-trained Chinese ELECTRA(中文ELECTRA预训练模型)
基于Pytorch和torchtext的自然语言处理深度学习框架。
Z-Bench 1.0 by 真格基金:一个麻瓜的大语言模型中文测试集。Z-Bench is a LLM prompt dataset for non-technical users, developed by an enthusiastic AI-focused team in Zhenfund.
Data from GitHub · snapshot Sep 24, 2026