🏆 #2,267 overall#537 of 601 in Language Models

cambridgeltl /sapbert

[NAACL'21 & ACL'21] SapBERT: Self-alignment pretraining for BERT & XL-BEL: Cross-Lingual Biomedical Entity Linking.

$ git clone https://github.com/cambridgeltl/sapbert.git
GitHub social preview for cambridgeltl/sapbert
Stars
234
234
Forks
40
40
Language
Python
License
MIT
Created
Apr 9, 2021
5.5 years old
Last push
Apr 28, 2023

Categories

GitHub topics

More in Language Models

Language Models

explosion/spacy-transformers

🛸 Use pretrained transformers like BERT, XLNet and GPT-2 in spaCy

🧠 Natural Language Processing
PythonMITupdated Mar 27, 2026
GitHub ↗★ 1.4K⑂ 179
Language Models

extreme-bert/extreme-bert

ExtremeBERT is a toolkit that accelerates the pretraining of customized language models on customized datasets, described in the paper “ExtremeBERT: A Toolkit for Accelerating Pretraining of Customized BERT”.

🧠 Natural Language Processing
PythonApache-2.0updated Mar 5, 2023
GitHub ↗★ 267⑂ 14
Language Models

huggingface/tokenizers🔥 active

💥 Fast State-of-the-Art Tokenizers optimized for Research and Production

🧠 Natural Language Processing
RustApache-2.0updated Sep 23, 2026
GitHub ↗★ 11.1K⑂ 1.2K
Language Models

brightmart/nlp_chinese_corpus

大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP

🗂️ Text Classification❓ Question Answering🧬 Embeddings🧠 Natural Language Processing
OtherMITupdated Feb 6, 2026
GitHub ↗★ 9.9K⑂ 1.6K
Language Models

codertimo/BERT-pytorch

Google AI 2018 BERT pytorch implementation

🧠 Natural Language Processing
PythonApache-2.0updated Sep 15, 2023
GitHub ↗★ 6.5K⑂ 1.3K
Language Models

CVI-SZU/Linly

Chinese-LLaMA 1&2、Chinese-Falcon 基础模型;ChatFlow中文对话模型;中文OpenLLaMA模型;NLP预训练/指令微调数据集

🧠 Natural Language Processing
Pythonno licenseupdated Apr 14, 2024
GitHub ↗★ 3K⑂ 222

Data from GitHub · snapshot Sep 24, 2026