πŸ† #140 overall#50 of 601 in Language ModelsπŸ”₯ active this week

huggingface /tokenizers

πŸ’₯ Fast State-of-the-Art Tokenizers optimized for Research and Production

$ git clone https://github.com/huggingface/tokenizers.git
GitHub social preview for huggingface/tokenizers
Stars
11.1K
11,109
Forks
1.2K
1,207
Language
Rust
License
Apache-2.0
Created
Nov 1, 2019
6.9 years old
Last push
Sep 23, 2026
πŸ”₯ this week

Categories

GitHub topics

More in Language Models

Language Models

explosion/spacy-transformers

πŸ›Έ Use pretrained transformers like BERT, XLNet and GPT-2 in spaCy

🧠 Natural Language Processing
PythonMITupdated Mar 27, 2026
GitHub β†—β˜… 1.4Kβ‘‚ 179
Language Models

ashishpatel26/Treasure-of-Transformers

πŸ’ Awesome Treasure of Transformers Models for Natural Language processing contains papers, videos, blogs, official repo along with colab Notebooks. πŸ›«β˜‘οΈ

πŸŽ™οΈ Speech Recognition🧠 Natural Language Processing
Jupyter NotebookMITupdated Aug 1, 2025
GitHub β†—β˜… 1.2Kβ‘‚ 238
Language Models

SKTBrain/KoBERT

Korean BERT pre-trained cased (KoBERT)

🧠 Natural Language Processing
PythonApache-2.0updated Jun 14, 2025
GitHub β†—β˜… 1.4Kβ‘‚ 376
Language Models

ymcui/MacBERT

Revisiting Pre-trained Models for Chinese Natural Language Processing (MacBERT)

🧠 Natural Language Processing
OtherApache-2.0updated Apr 19, 2026
GitHub β†—β˜… 722β‘‚ 61
Language Models

charles9n/bert-sklearn

a sklearn wrapper for Google's BERT model

🏷️ Named Entity Recognition🧠 Natural Language Processing
Jupyter NotebookApache-2.0updated Oct 26, 2022
GitHub β†—β˜… 302β‘‚ 70
Language Models

extreme-bert/extreme-bert

ExtremeBERT is a toolkit that accelerates the pretraining of customized language models on customized datasets, described in the paper β€œExtremeBERT: A Toolkit for Accelerating Pretraining of Customized BERT”.

🧠 Natural Language Processing
PythonApache-2.0updated Mar 5, 2023
GitHub β†—β˜… 267β‘‚ 14

Data from GitHub Β· snapshot Sep 24, 2026