huggingface/tokenizersπ₯ active
π₯ Fast State-of-the-Art Tokenizers optimized for Research and Production
BlueBERT, pre-trained on PubMed abstracts and clinical notes (MIMIC-III).
$ git clone https://github.com/ncbi-nlp/bluebert.gitπ₯ Fast State-of-the-Art Tokenizers optimized for Research and Production
πΈ Use pretrained transformers like BERT, XLNet and GPT-2 in spaCy
π Awesome Treasure of Transformers Models for Natural Language processing contains papers, videos, blogs, official repo along with colab Notebooks. π«βοΈ
The prime repository for state-of-the-art Multilingual Question Answering research and development.
a sklearn wrapper for Google's BERT model
ExtremeBERT is a toolkit that accelerates the pretraining of customized language models on customized datasets, described in the paper βExtremeBERT: A Toolkit for Accelerating Pretraining of Customized BERTβ.
Data from GitHub Β· snapshot Sep 24, 2026