Tokenization & Preprocessing
sugarme/tokenizer
NLP tokenizers written in Go language
๐ง Natural Language Processing
GoApache-2.0updated May 26, 2026
Rust-tokenizer offers high-performance tokenizers for modern language models, including WordPiece, Byte-Pair Encoding (BPE) and Unigram (SentencePiece) models
$ git clone https://github.com/guillaume-be/rust-tokenizers.gitNLP tokenizers written in Go language
[CVPR 2026 Highlight] A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
Implementation of the GBST block from the Charformer paper, in Pytorch
Kanade is a single-layer disentangled speech tokenizer that extracts compact tokens suitable for both generative and discriminative modeling.
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
๐ซ Industrial-strength Natural Language Processing (NLP) in Python
Data from GitHub ยท snapshot Sep 24, 2026