Language Models
brightmart/nlp_chinese_corpus
大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP
🗂️ Text Classification❓ Question Answering🧬 Embeddings🧠 Natural Language Processing
OtherMITupdated Feb 6, 2026
언어모델을 학습하기 위한 공개 한국어 instruction dataset들을 모아두었습니다.
$ git clone https://github.com/HeegyuKim/open-korean-instructions.git大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP
中文语言理解测评基准 Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard
Korean BERT pre-trained cased (KoBERT)
Pretrained ELECTRA Model for Korean
DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection and Instruction-Aware Models for Conversational AI
Data from GitHub · snapshot Sep 24, 2026