huggingface/tokenizersπ₯ active
π₯ Fast State-of-the-Art Tokenizers optimized for Research and Production
πΈ Use pretrained transformers like BERT, XLNet and GPT-2 in spaCy
$ git clone https://github.com/explosion/spacy-transformers.gitπ₯ Fast State-of-the-Art Tokenizers optimized for Research and Production
π Awesome Treasure of Transformers Models for Natural Language processing contains papers, videos, blogs, official repo along with colab Notebooks. π«βοΈ
ExtremeBERT is a toolkit that accelerates the pretraining of customized language models on customized datasets, described in the paper βExtremeBERT: A Toolkit for Accelerating Pretraining of Customized BERTβ.
Interpret text data with LLMs (sklearn compatible).
Transformers4Rec is a flexible and efficient library for sequential and session-based recommendation and works with PyTorch.
The prime repository for state-of-the-art Multilingual Question Answering research and development.
Data from GitHub Β· snapshot Sep 24, 2026