๐Ÿ† #1,899 overall#60 of 148 in Tokenization & Preprocessing

guillaume-be /rust-tokenizers

Rust-tokenizer offers high-performance tokenizers for modern language models, including WordPiece, Byte-Pair Encoding (BPE) and Unigram (SentencePiece) models

$ git clone https://github.com/guillaume-be/rust-tokenizers.git
GitHub social preview for guillaume-be/rust-tokenizers
Stars
344
344
Forks
34
34
Language
Rust
License
Apache-2.0
Created
Nov 9, 2019
6.9 years old
Last push
Jan 22, 2026

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

sugarme/tokenizer

NLP tokenizers written in Go language

๐Ÿง  Natural Language Processing
GoApache-2.0updated May 26, 2026
GitHub โ†—โ˜… 335โ‘‚ 69
Tokenization & Preprocessing

amazon-far/deltatok

[CVPR 2026 Highlight] A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens

PythonApache-2.0updated Jul 17, 2026
GitHub โ†—โ˜… 261โ‘‚ 7
Tokenization & Preprocessing

frothywater/kanade-tokenizer

Kanade is a single-layer disentangled speech tokenizer that extracts compact tokens suitable for both generative and discriminative modeling.

Pythonno licenseupdated Jul 18, 2026
GitHub โ†—โ˜… 116โ‘‚ 16
Tokenization & Preprocessing

rasbt/LLMs-from-scratch๐Ÿ”ฅ active

Implement a ChatGPT-like LLM in PyTorch from scratch, step by step

๐Ÿค– Language Models๐Ÿง  Natural Language Processing
Jupyter NotebookOtherupdated Sep 22, 2026
GitHub โ†—โ˜… 105.5Kโ‘‚ 16.2K
Tokenization & Preprocessing

explosion/spaCy

๐Ÿ’ซ Industrial-strength Natural Language Processing (NLP) in Python

๐Ÿท๏ธ Named Entity Recognition๐Ÿ—‚๏ธ Text Classification๐Ÿง  Natural Language Processing
PythonMITupdated Aug 24, 2026
GitHub โ†—โ˜… 33.9Kโ‘‚ 4.7K

Data from GitHub ยท snapshot Sep 24, 2026