Language Models
huggingface/tokenizers🔥 active
💥 Fast State-of-the-Art Tokenizers optimized for Research and Production
🧠 Natural Language Processing
RustApache-2.0updated Sep 23, 2026
A Dutch RoBERTa-based language model
$ git clone https://github.com/iPieter/RobBERT.git💥 Fast State-of-the-Art Tokenizers optimized for Research and Production
Korean BERT pre-trained cased (KoBERT)
Revisiting Pre-trained Models for Chinese Natural Language Processing (MacBERT)
大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP
Google AI 2018 BERT pytorch implementation
Chinese-LLaMA 1&2、Chinese-Falcon 基础模型;ChatFlow中文对话模型;中文OpenLLaMA模型;NLP预训练/指令微调数据集
Data from GitHub · snapshot Sep 24, 2026