Language Models
explosion/spacy-transformers
🛸 Use pretrained transformers like BERT, XLNet and GPT-2 in spaCy
🧠 Natural Language Processing
PythonMITupdated Mar 27, 2026
[NAACL'21 & ACL'21] SapBERT: Self-alignment pretraining for BERT & XL-BEL: Cross-Lingual Biomedical Entity Linking.
$ git clone https://github.com/cambridgeltl/sapbert.git🛸 Use pretrained transformers like BERT, XLNet and GPT-2 in spaCy
ExtremeBERT is a toolkit that accelerates the pretraining of customized language models on customized datasets, described in the paper “ExtremeBERT: A Toolkit for Accelerating Pretraining of Customized BERT”.
💥 Fast State-of-the-Art Tokenizers optimized for Research and Production
大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP
Google AI 2018 BERT pytorch implementation
Chinese-LLaMA 1&2、Chinese-Falcon 基础模型;ChatFlow中文对话模型;中文OpenLLaMA模型;NLP预训练/指令微调数据集
Data from GitHub · snapshot Sep 24, 2026